Machine Learning Engineer
Scale.jobs · Chicago, IL · 1 mo ago
RemoteRemoteEngineeringFull-time
About The Role
The role drives the development, scaling, and maintenance of core machine learning pipelines and real-time model serving infrastructure. This engineer will translate advanced prototypes into reliable, highly performant production systems that serve millions of predictions daily. Working within a collaborative team of data scientists and backend engineers, this role directly impacts system latency, model accuracy, and overall engineering velocity by building robust MLOps patterns.
Key Responsibilities
- Design, implement, and maintain scalable ML training and inference pipelines using Python, PySpark, and orchestration tools like Airflow or Prefect
- Deploy deep learning and classical ML models to production environments using AWS SageMaker, Kubernetes, and Triton Inference Server
- Develop and maintain a centralized feature store to ensure consistency between offline training data and real-time online features
- Implement automated monitoring, logging, and alerting systems to track model drift, performance degradation, and system latency
- Optimize model latency and throughput through techniques such as quantization, compilation (TensorRT/ONNX), and distributed serving
- Collaborate on architectural decisions to integrate ML services seamlessly with microservices-based backend APIs
What We Are Looking For
- 3-6 years of professional experience as a Machine Learning Engineer or Software Engineer with a heavy focus on production ML systems
- Strong proficiency in Python and hands-on experience with ML/DL frameworks such as PyTorch, XGBoost, or TensorFlow
- Solid software engineering fundamentals, including experience with Docker, CI/CD pipelines, and writing unit/integration tests
- Experience with cloud infrastructure (specifically AWS or GCP) and container orchestration using Kubernetes or EKS
- .S. or M.S. in Computer Science, Data Science, or a related quantitative engineering field
- Bonus: Experience with Triton Inference Server, MLflow, Ray, or running instruction fine-tuning for LLMs