Senior ML Ops Engineer (Machine Learning Infrastructure)
Parallel · Los Angeles, CA · 1 mo ago
HybridInformation Technology$150k/yrFull-time
Responsibilities
- Design and implement robust MLOps solutions, including automated pipelines for data management, model training, deployment and monitoring.
- Architect, deploy, and manage scalable ML infrastructure for distributed training and inference.
- Collaborate with ML engineers to gather requirements and develop strategies for data management, model development and deployment.
- Build and operate cloud-based systems (e.g., AWS, GCP) optimized for ML workloads in R&D, and production environments.
- Build scalable ML infrastructure to support continuous integration/deployment, experiment management, and governance of models and datasets.
- Support the automation of model evaluation, selection, and deployment workflows.
Qualifications
- Bachelor’s or higher degree in Computer Science, Machine Learning, or a relevant engineering discipline.
- 5+ years of experience building large-scale, reliable systems; 2+ years focused on ML infrastructure or MLOps.
- Proven experience architecting and deploying production-grade ML pipelines and platforms.
- Strong knowledge of ML lifecycle: data ingestion, model training, evaluation, packaging, and deployment.
- Hands-on experience with MLOps tools (e.g., MLflow, Kubeflow, SageMaker, Airflow, Metaflow, or similar).
- Deep understanding of CI/CD practices applied to ML workflows.
- Proficiency in Python, Git, and system design with solid software engineering fundamentals.
- Experience with cloud platforms (AWS, GCP, or Azure) and designing ML architectures in those environments.