Technical Lead Manager, Machine Learning Runtime & Serving
Technical Lead Manager, Machine Learning Runtime & Serving
The Waymo Driver powers Waymo’s fully autonomous ride-hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten million rider-only trips, enabled by its experience autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states.
Guide the technical vision of our core ML infrastructure while actively growing and managing a high-performing team of 6 engineers to deliver Waymo’s next-generation ML ecosystem, encompassing both the in-vehicle inference engine and the cloud-based serving infrastructure for our foundational models.
Architect scalable, high-performance ML runtime systems that operate flawlessly across two extreme domains: the highly constrained edge compute environment of autonomous vehicles and our large-scale, offboard data centers.
Navigate complex engineering trade-offs, driving feature development that seamlessly balances the strict, real-time latency and memory limits of onboard execution with the high-throughput, highly concurrent demands of fleet-scale cloud serving.
Spearhead the strategic transition of core ML workloads to a JAX-native runtime architecture, which includes actively extending and modifying underlying ML compilers and runtimes (e.g., OpenXLA/PjRT, TensorRT).
Partner across organizational boundaries with world-class ML researchers in Perception and Planning to deeply analyze system-level workloads and unlock massive performance gains through hardware-aware compute optimizations.
Drive systemic performance excellence by designing advanced profiling and benchmarking infrastructure to identify, triage, and eliminate bottlenecks across the entire end-to-end ML software stack.
Qualifications
B.S. or M.S. in CS, EE, Deep Learning or a related field.
People management experience, with a proven track record of recruiting, mentoring, and guiding high-performing teams of senior engineers.
8+ years of professional software engineering experience architecting, building, and scaling complex ML systems and infrastructure.
Strong production programming expertise.
Prominent track record of optimizing ML software to maximize the performance of hardware accelerators (e.g., GPUs, TPUs, or custom silicon).
Hands-on experience developing distributed backend systems that are low-latency, highly concurrent, and fault-tolerant at scale.
Deep expertise in modifying and extending ML software stacks, including compilers, runtimes, or inference engines (e.g., OpenXLA/PjRT, TensorRT, ONNX Runtime, TVM).
Proven track record of building and scaling LLM serving systems, leveraging advanced distributed inference and performance optimization techniques.
Deep expertise in edge computing and automotive ML deployment, navigating strict power, thermal, and real-time latency constraints to optimize and deploy mission-critical models on resource-constrained embedded hardware.
Benefits
Competitive compensation, bonus opportunities, equity, employees provident fund, and lots of other perks and employee discounts.
Flexible work arrangements, including the option to work remotely for four weeks per year.
Enhanced leave options, including 26 weeks of paid leave for birthing parents and 18 weeks of paid leave for non-birthing parents.
Access to Employee Resource Groups (ERGs), personal and professional development opportunities, and time off to volunteer.
Cool perks such as access to Google offices, cafes, wellness centers, massages, and at-home fitness and cooking classes.
Pay
The expected base salary range for this full-time position across US locations is $251,000—$310,000 USD.
Schedule
This is a full-time position.