MLOps Engineer
Evlo AI · Chicago, IL · 1 wk ago
RemoteRemoteEngineeringFull-time
About the role
The role owns the infrastructure, orchestration, and continuous deployment pipelines that power machine learning at scale. The focus is on turning experimental AI prototypes into reliable, low-latency, production-ready microservices. The team works closely with applied scientists and data engineers to build robust platforms that handle model tracking, automated retraining, and infrastructure scaling across cloud environments.
Responsibilities
- Architect, build, and maintain end-to-end MLOps pipelines using Kubernetes, Docker, and CI/CD tools like GitHub Actions or ArgoCD
- Deploy and manage model serving infrastructure leveraging Triton, Ray Serve, or vLLM to optimize throughput and inference latency
- Implement comprehensive model monitoring systems to detect data drift, concept drift, and system anomalies in production
- Automate model training and data preprocessing pipelines using orchestration tools such as Kubeflow, Airflow, or Prefect
- Collaborate with security and engineering teams to ensure governance, compliance, and reproducibility across all deployed ML artifacts
- Write clean, resilient infrastructure-as-code using Terraform and participate in on-call rotations for production systems
Requirements
- 3–6 years of experience in MLOps, DevOps, or machine learning engineering with a strong focus on production infrastructure
- Deep expertise in containerization, Kubernetes orchestration, and cloud computing (AWS, GCP, or Azure)
- Hands-on experience with model serving frameworks, inference optimization techniques, and feature stores
- Proficiency in Python and Bash scripting, alongside experience building automated CI/CD pipelines
- Solid understanding of software engineering best practices, monitoring tools (Prometheus, Grafana), and logging infrastructure
- Bonus: Experience managing LLM infrastructure, fine-tuning pipelines, and vector database deployments at scale