Jobs · Engineering · California

Technical Lead Manager, Machine Learning Runtime & Serving

A16Z GAMES · Mountain View, CA · 6 days ago
Engineering$251k–$310k/yrFull-time

Technical Lead Manager, Machine Learning Runtime & Serving

The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.

  • Guide the technical vision of our core ML infrastructure while actively growing and managing a high-performing team of 6 engineers to deliver Waymo’s next-generation ML ecosystem, encompassing both the in-vehicle inference engine and the cloud-based serving infrastructure for our foundational models.

  • Architect scalable, high-performance ML runtime systems that operate flawlessly across two extreme domains: the highly constrained edge compute environment of autonomous vehicles and our large-scale, offboard data centers.

  • Navigate complex engineering trade-offs, driving feature development that seamlessly balances the strict, real-time latency and memory limits of onboard execution with the high-throughput, highly concurrent demands of fleet-scale cloud serving.

  • Spearhead the strategic transition of core ML workloads to a JAX-native runtime architecture, which includes actively extending and modifying underlying ML compilers and runtimes (e.g., OpenXLA/PjRT, TensorRT).

  • Partner across organizational boundaries with world-class ML researchers in Perception and Planning to deeply analyze system-level workloads and unlock massive performance gains through hardware-aware compute optimizations.

  • Drive systemic performance excellence by designing advanced profiling and benchmarking infrastructure to identify, triage, and eliminate bottlenecks across the entire end-to-end ML software stack.

Qualifications

  • B.S. or M.S. in CS, EE, Deep Learning or a related field.

  • People management experience, with a proven track record of recruiting, mentoring, and guiding high-performing teams of senior engineers.

  • 8+ years of professional software engineering experience architecting, building, and scaling complex ML systems and infrastructure.

  • Strong production programming expertise.

  • Prominent track record of optimizing ML software to maximize the performance of hardware accelerators (e.g., GPUs, TPUs, or custom silicon).

  • Hands-on experience developing distributed backend systems that are low-latency, highly concurrent, and fault-tolerant at scale.

  • Deep expertise in modifying and extending ML software stacks, including compilers, runtimes, or inference engines (e.g., OpenXLA/PjRT, TensorRT, ONNX Runtime, TVM).

  • Proven track record of building and scaling LLM serving systems, leveraging advanced distributed inference and performance optimization techniques.

  • Deep expertise in edge computing and automotive ML deployment, navigating strict power, thermal, and real-time latency constraints to optimize and deploy mission-critical models on resource-constrained embedded hardware.

Benefits

  • Competitive compensation, bonus opportunities, equity, employees provident fund, and lots of other perks and employee discounts.

  • Flexible work arrangements, including the option to work remotely for four weeks per year.

  • Enhanced leave options, including 26 weeks of paid leave for birthing parents and 18 weeks of paid leave for non-birthing parents.

  • Access to Employee Resource Groups (ERGs), personal and professional development opportunities, and time off to volunteer.

  • Cool perks such as access to Google offices, cafes, wellness centers, massages, and at-home fitness and cooking classes.

Pay

The expected base salary range for this full-time position across US locations is $251,000—$310,000 USD.

Schedule

This is a full-time position.

Similar jobs