Jobs · Engineering · California

Staff Engineer - ML Infra / MLOps

AIJobs.ai · Palo Alto, CA · 1 mo ago
Engineering$218k–$285k/yrFull-time

Founded in 2018, Quince challenges the idea that nice things have to cost a lot. Our mission is to make high-quality essentials at low prices, produced fairly and sustainably. We operate a direct-to-consumer (DTC) model that cuts out middlemen and leverages just-in-time manufacturing to minimize waste and maximize value. Quince is a tech company disrupting retail by putting AI, analytics, and automation at the center of everything we do.

Our Team and Success

At Quince, you will be part of a high-performing team redefining quality, value, and sustainability in modern retail. We are a destination for builders, innovators, and operators who challenge the status quo. Our collective ambition is to create an entirely new category and customer experience—one that democratizes luxury and provides high-quality products at radically low prices. If you are motivated by impact, growth, and purpose, you will find a strong sense of belonging at Quince.

About the Role

We are seeking a Staff ML Engineer to join our growing team. The ideal candidate is a deeply technical ML infrastructure engineer who combines hands-on mastery with system-level thinking. You have built and operated production-grade ML systems at scale—from distributed training pipelines and feature stores to high-throughput inference serving—and you take pride in engineering platforms that other engineers love to use. You design for extensibility, observability, and resilience, tackling the hardest problems in optimizing GPU utilization, zero-downtime model deployment, and industrializing AI at scale. You operate with high autonomy, hold yourself to exceptional standards, and elevate the engineers around you through mentorship and technical leadership.

Responsibilities

  • Architect the ML Infrastructure Foundation: Own the end-to-end technical design of Quince’s ML platform—including model training, serving, feature pipelines, and monitoring—ensuring it is modular, scalable, and built for long-term extensibility.
  • Build the “Paved Road” for Production: Design and implement the core developer experience for Quince’s Data Scientists and AI Researchers, enabling them to move from “idea to production” with minimal friction and maximum reliability.
  • Drive Technical Excellence Across the Stack: Set and uphold engineering standards in CI/CD for ML, Infrastructure as Code (IaC), model versioning, experiment tracking, and deployment strategies (blue-green, canary)—and build the tooling that makes those standards the path of least resistance.
  • Own High-Impact System Design Decisions: Lead the technical evaluation and selection of core platform components—from inference runtimes and feature stores to orchestration frameworks—with a clear-eyed view of build vs. buy tradeoffs.
  • Optimize Compute Performance & Cost: Design and implement GPU utilization optimizations, model batching strategies, and cloud cost controls to maximize performance per dollar across training and inference workloads.
  • Ensure Production Scalability & Reliability: Architect ML serving infrastructure that gracefully handles traffic surges, seasonal spikes, and model version transitions, with robust monitoring, alerting, and automated recovery.
  • Mentor and Elevate the Engineering Team: Provide deep technical mentorship to junior and mid-level engineers through design reviews, code reviews, and pairing sessions—raising the collective technical bar without adding process overhead.
  • Champion Operational Excellence: Lead root-cause analyses (RCAs) for production failures and drive systemic, permanent fixes over reactive patches. Model a culture of rigorous on-call discipline and accountability.

Requirements

  • 8+ years of industry experience, with at least 4+ years of focused, hands-on work in ML Infrastructure, MLOps, or large-scale Data Platform engineering.
  • Proven track record of designing and building MLOps platforms that support the full model lifecycle—from data ingestion and distributed training to real-time inference and model governance.
  • Deep expertise in cloud-native infrastructure (preferably AWS), Kubernetes (EKS), Docker, and Infrastructure as Code tools (Terraform/Pulumi).
  • Hands-on mastery of ML frameworks such as PyTorch, TensorFlow, Kubeflow, or SageMaker, with strong opinions on building a cohesive, high-leverage developer experience.
  • Expertise in building Feature Stores and high-throughput data pipelines (Spark, Flink, Kafka), with a strong understanding of training/serving skew and data consistency.
  • Expert-level knowledge of CI/CD for ML, including model versioning, experiment tracking, and deployment strategies such as blue-green and canary rollouts.
  • Demonstrated ability to optimize GPU utilization, implement model batching, and systematically reduce cloud infrastructure costs.
  • Strong operational instincts, with a history of improving reliability through rigorous on-call practices, proactive monitoring, and root-cause analysis.
  • Ability to thrive in a startup environment, handling ambiguity with curiosity, quick learning, and a love for experimentation.

Pay

Pay range: $218,000–$285,000 (base) + bonus and stock. All posted ranges are reflective of base salary and may vary depending on experience level and location.

Similar jobs

Staff Engineer, DevOps

Fresenius Medical CareWaltham, MA· 1 mo ago
Engineering$122k–$205k/yrapply on jobs.freseniusmedicalcare.com