Lead Machine Learning Inference Engineer, Advertising
About the role
In this role, you will architect, design, and lead the development of a SOTA Inference platform that can handle Advertising-level low latencies, scale, throughput, and availability with optimizations that span across hardware, software, and models.
We’re looking for a strong technical leader with deep experience in ML serving, high-performance computing, and industry standard frameworks - someone excited to mentor engineers, innovate at scale, and shape the future of machine learning at Roku.
For California Only - The estimated annual salary for this position is between $246,500 - $486,100 annually. Compensation packages are based on factors unique to each candidate, including but not limited to skill set, certifications, and specific geographical location.
This role is eligible for health insurance, equity awards, life insurance, disability benefits, parental leave, wellness benefits, and paid time off.
What you’ll be doing
- Lead the design and development of a SOTA Inference platform
- Oversee the development of monitoring, observability, and other tooling to ensure system and model performance, reliability, and scalability of online inference services
- Identify and resolve system inefficiencies, performance bottlenecks, and reliability issues, ensuring optimized end-to-end performance
- Stay at the forefront of advancements in inference frameworks, ML hardware acceleration, and distributed systems, and incorporate innovations where and when they are impactful
Requirements
- M.S. or above in CS, ECE, or a related field
- 10+ years of experience in developing and deploying large-scale, distributed systems, with at least 5 years in a leadership or technical lead role
- Strong programming skills in high-performance languages
- Deep understanding of inference frameworks and ML system deployment
- Proven experience optimizing performance for large-scale machine learning systems, including a deep knowledge of SOTA model optimizations, hardware-software co-design, GPU acceleration, and HPC techniques
- Excellent communication and collaboration skills
- Experience leading teams working on high-throughput, low-latency ML serving systems
- Experience collaborating with and leading global, cross-functional teams
- Contributions to open-source ML or systems projects
Qualifications
- Experience with inference frameworks such as TensorFlow, PyTorch, etc.
- Knowledge of distributed systems, cloud platforms like AWS, GCP, etc.
- Experience with ML hardware acceleration technologies like GPUs, TPUs, etc.
- Experience with large-scale data processing and storage systems like Hadoop, Spark, etc.
- Experience with performance tuning and optimization techniques
Skills
- Strong analytical and problem-solving skills
- Ability to communicate effectively with cross-functional teams
- Experience with version control systems like Git
- Experience with CI/CD pipelines
Benefits
- Health insurance
- Equity awards
- Life insurance
- Disability benefits
- Parental leave
- Wellness benefits
- Paid time off
Pay
The estimated annual salary for this position is between $246,500 - $486,100 annually.
Schedule
Roku fosters an inclusive and collaborative environment where teams generally work in the office Monday through Thursday. Fridays are generally flexible for remote work, except for employees whose specific roles or assigned office location require five days' a week attendance.