MLOps Engineer
Atomic Machines · Emeryville, CA · 3 days ago
On-siteEngineering$200k/yrFull-time
About The Role
We are seeking an MLOps Engineer to join our AI and Modeling & Simulation org within the Data Engineering and Analytics team.
What You’ll Do
- Build and evolve the MLOps platform and CI/CD: Own the path from experiment to production, including experiment tracking, model registry, packaging, automated training and retraining, deployment, and safe rollout and rollback.
- Operate model serving infrastructure: Build reliable, scalable batch, streaming, and real-time inference for models and digital twins supporting design, process control, scheduling, and inspection.
- Maintain ML data infrastructure: Support feature-store capabilities and a lakehouse foundation using Apache Iceberg on S3, with strong data quality, lineage, versioning, and reproducibility.
- Create paved roads for ML development: Develop standardized tooling and workflows that enable Data and AI engineers to move quickly while maintaining production reliability and reproducibility.
- Close the model feedback loop: Build model observability and human-in-the-loop systems that capture production signals and expert corrections, version them as ground truth, and feed them into evaluation and retraining workflows.
- Drive technical ownership: Identify infrastructure, reliability, and scalability challenges and drive solutions from design through production.
- Collaborate across disciplines: Work with Process, Chemical, Materials, Simulation, Software, Data, and AI engineers to define deployment, serving, and data-collection requirements.
What You’ll Need
- 5+ years of relevant industry experience building production software, infrastructure, data, or machine learning systems.
- Proven experience building and operating machine learning systems in production, with a strong MLOps/DevOps orientation.
- Strong DevOps fundamentals, including CI/CD, containers, Kubernetes, cloud infrastructure, and infrastructure-as-code.
- Proficiency in Python and SQL.
- Hands-on experience with MLflow or similar tooling for experiment tracking, model registry, and model lifecycle management.
- Experience with S3, lakehouse technologies such as Apache Iceberg, and workflow orchestration tools such as Airflow or Dagster.
- Experience building pipelines for multimodal ML data, including images, time-series, structured, and semi-structured data.
- Familiarity with manufacturing systems, sensors, process automation, or other physical-world data systems.
- Strong problem-solving skills, attention to data quality and reliability, and clear technical communication.
- Bachelor's or Master's degree in Computer Science, Data Engineering, Data Science, or a related STEM field, or equivalent practical experience.
Bonus Points For
- Experience with feature stores, human-in-the-loop systems, active learning, or data-labeling infrastructure.
- Robotics or robotic automation experience, including sensors, vision systems, or robotics data.
- Experience operating ML systems in manufacturing or other physical-world environments.
- Experience building internal tools for expert feedback, labeling, model evaluation, or model interaction.
- Experience designing shared ML infrastructure or platforms used across multiple teams or applications.