MLOps Engineer
Evlo AI · Seattle, WA · 5 days ago
HybridFull-time
About The Role
The role owns the infrastructure and operational pipelines that keep machine learning models and large language models running reliably, scalably, and efficiently in production. The team works closely with machine learning engineers and applied scientists to bridge the gap between experimental notebooks and high-throughput, low-latency production services.
Key Responsibilities
- Design, build, and maintain robust MLOps infrastructure for training, fine-tuning, and deploying machine learning and GenAI models at scale
- Implement automated CI/CD pipelines for ML, ensuring seamless model validation, testing, and deployment using tools like MLflow, Kubeflow, or Argo
- Manage and optimize cloud infrastructure on AWS or GCP, utilizing Kubernetes, Docker, and Terraform for reproducible deployments
- Monitor production models and infrastructure for latency regressions, resource utilization bottlenecks, and data drift
- Collaborate with security and data engineering teams to establish strict governance, data lineage tracking, and model artifact registries
- Write clean, maintainable automation scripts and infrastructure-as-code to support rapid iteration for the engineering organization
What We Are Looking For
- 3–6 years of experience in MLOps, DevOps, or machine learning engineering with a heavy focus on production infrastructure
- Hands-on experience with containerization and orchestration platforms, specifically Docker and Kubernetes
- Familiarity with ML model serving frameworks and inference optimization techniques, such as TensorRT, vLLM, or Triton
- Proficiency in Python and Bash for automation, along with infrastructure-as-code tools like Terraform or CloudFormation
- Strong understanding of cloud platforms (AWS, GCP, or Azure) and CI/CD pipelines (GitHub Actions, GitLab CI)
- Bonus: Experience managing GPU clusters for LLM training/inference, contributing to open-source MLOps tools, or holding a relevant cloud certification