MLOps Engineer
Evlo AI · Washington, DC · 1 wk ago
RemoteRemoteEngineeringFull-time
About the role
The role owns the infrastructure and pipelines that take machine learning models from experimentation to robust, scalable production environments. The team builds and maintains the core MLOps platform, ensuring reliable model serving, automated CI/CD workflows, and continuous monitoring for performance degradation and data drift.
Responsibilities
- Architect, build, and maintain scalable ML infrastructure and deployment pipelines using Kubernetes, Docker, and Terraform
- Implement automated CI/CD workflows for machine learning applications, integrating unit testing, model validation, and security scans
- Optimize model serving architectures for low latency and high throughput using Triton Inference Server or Ray Serve
- Configure automated monitoring systems to detect data drift, concept drift, and resource bottlenecks across production models
- Collaborate with data scientists and ML engineers to streamline feature stores and artifact registries
- Write clean, infrastructure-as-code configurations and maintain comprehensive documentation for deployment processes
Requirements
- 3–6 years of experience in MLOps, DevOps, or machine learning engineering with a strong focus on production infrastructure
- Demonstrated experience deploying and managing ML models in public cloud environments (AWS, GCP, or Azure)
- Strong proficiency in containerization and orchestration tools: Docker, Kubernetes, and Helm
- Hands-on experience with feature stores, model registries, and experiment tracking tools such as MLflow or Weights & Biases
- Solid software engineering fundamentals in Python, including scripting, API development, and automation
- Bonus: Experience with LLM serving infrastructure, vLLM, or distributed training frameworks like DeepSpeed