MLOps Engineer
Evlo AI · Atlanta, GA · 3 days ago
RemoteRemoteEngineeringFull-time
About The Role
The role owns the infrastructure, orchestration, and scaling of machine learning systems, ensuring models transition smoothly from research notebooks to high-availability production services. The engineering team builds robust MLOps pipelines, monitoring frameworks, and automated deployment tooling to support high-throughput AI workloads.
Key Responsibilities
- Design and implement end-to-end MLOps pipelines using Kubernetes, Docker, Kubeflow, and MLflow for automated model training, validation, and deployment
- Provision and manage scalable cloud infrastructure on AWS or GCP to support intensive distributed model training and low-latency inference workloads
- Build robust data and feature pipelines to ensure low-latency data access and integrity across training and serving environments
- Implement comprehensive model monitoring systems for real-time detection of data drift, concept drift, and performance degradation
- Optimize model inference latency, throughput, and resource utilization through quantization, pruning, and hardware acceleration techniques
- Collaborate with machine learning engineers and data scientists to establish standardized CI/CD pipelines for model artifacts and code
What We Are Looking For
- 3–7 years of experience in MLOps, DevOps, or machine learning engineering with a strong focus on production infrastructure
- Deep expertise in containerization and orchestration tools: Docker, Kubernetes, and Helm
- Hands-on experience with ML lifecycle platforms and model registries such as MLflow, Weights & Biases, Kubeflow, or SageMaker
- Solid understanding of CI/CD principles and automated testing for both software and machine learning models
- Bonus: Experience deploying and scaling Large Language Models (LLMs) and vector databases in production environments