Machine Learning Engineer
Scale.jobs · Chicago, IL · 3 days ago
RemoteRemoteEngineeringFull-time
About The Role
This role focus on building, scaling, and maintaining the core machine learning infrastructure and production models that power real-time decisioning and intelligent product features. The engineer will bridge the gap between experimental data science and high-performance software engineering, ensuring models are optimized for latency, throughput, and cost. Working within a collaborative team of data platform engineers and product scientists, this role will own the deployment pipelines, model registry, and serving infrastructure. The focus is on driving architectural decisions that enable rapid, reliable experimentation and seamless production continuous deployment.
Key Responsibilities
- Deploy, monitor, and scale machine learning models in production using Kubernetes, Triton Inference Server, or AWS SageMaker
- Develop and optimize robust feature stores and batch/streaming data pipelines using PySpark, Kafka, and dbt
- Implement automated retraining pipelines and continuous integration/continuous deployment (CI/CD) for ML assets using MLflow and GitHub Actions
- Optimize model performance through quantization, pruning, and hardware-accelerated runtimes like TensorRT or ONNX
- Establish rigorous observability and monitoring frameworks to detect data drift, conceptual drift, and system latency anomalies
- Collaborate with backend engineers to integrate ML APIs seamlessly into core microservices architectures
What We Are Looking For
- 3–6 years of experience as a Machine Learning Engineer or MLOps Engineer in a production cloud environment
- Expert-level Python programming skills along with deep familiarity with frameworks like PyTorch, scikit-learn, and Hugging Face
- Hands-on experience with containerization (Docker, Kubernetes) and orchestrators such as Airflow, Kubeflow, or Prefect
- Solid understanding of relational and non-relational databases, including PostgreSQL, Redis, and vector databases like Pinecone or Milvus
- Strong background in software engineering best practices: unit testing, version control, CI/CD, and system design
- Bonus: Experience with large language model (LLM) serving frameworks (vLLM, TGI) or distributed training frameworks (Ray, DeepSpeed)