Lead ML Inference Engineer, Advertising
Roku · San Jose, CA · 1 mo ago
HybridEngineering$247k–$486k/yrFull-time
About the role
In this role, you will architect, design, and lead the development of a SOTA Inference platform that can handle Advertising-level low latencies, scale, throughput, and availability with optimizations that span across hardware, software, and models.
Responsibilities
- Lead the design and development of a SOTA Inference platform
- Oversee the development of monitoring, observability, and other tooling to ensure system and model performance, reliability, and scalability of online inference services
- Identify and resolve system inefficiencies, performance bottlenecks, and reliability issues, ensuring optimized end-to-end performance
- Stay at the forefront of advancements in inference frameworks, ML hardware acceleration, and distributed systems, and incorporate innovations where and when they are impactful
Requirements
- M.S. or above in CS, ECE, or a related field
- 10+ years of experience in developing and deploying large-scale, distributed systems, with at least 5 years in a leadership or technical lead role
- Strong programming skills in high-performance languages
- Deep understanding of inference frameworks and ML system deployment
- Proven experience optimizing performance for large-scale machine learning systems, including a deep knowledge of SOTA model optimizations, hardware-software co-design, GPU acceleration, and HPC techniques
- Excellent communication and collaboration skills
- Experience leading teams working on high-throughput, low-latency ML serving systems
- Experience collaborating with and leading global, cross-functional teams
Qualifications
- Contributions to open-source ML or systems projects
Skills
- Experience with inference frameworks and ML system deployment
- Knowledge of SOTA model optimizations, hardware-software co-design, GPU acceleration, and HPC techniques
- Ability to optimize performance for large-scale machine learning systems
- Experience leading teams working on high-throughput, low-latency ML serving systems
- Experience collaborating with and leading global, cross-functional teams
- Contributions to open-source ML or systems projects
Benefits
- Health insurance
- Equity awards
- Life insurance
- Disability benefits
- Parental leave
- Wellness benefits
- Paid time off