Jobs · Engineering · Texas

Lead ML Inference Engineer, Advertising

Roku · Austin, TX · 1 mo ago
HybridEngineeringFull-time

About the role

In this role, you will architect, design, and lead the development of a SOTA Inference platform that can handle Advertising-level low latencies, scale, throughput, and availability with optimizations that span across hardware, software, and models. We’re looking for a strong technical leader with deep experience in ML serving, high-performance computing, and industry standard frameworks - someone excited to mentor engineers, innovate at scale, and shape the future of machine learning at Roku.

Responsibilities

  • Lead the design and development of a SOTA Inference platform
  • Oversee the development of monitoring, observability, and other tooling to ensure system and model performance, reliability, and scalability of online inference services
  • Identify and resolve system inefficiencies, performance bottlenecks, and reliability issues, ensuring optimized end-to-end performance
  • Stay at the forefront of advancements in inference frameworks, ML hardware acceleration, and distributed systems, and incorporate innovations where and when they are impactful

Requirements

  • M.S. or above in CS, ECE, or a related field
  • 10+ years of experience in developing and deploying large-scale, distributed systems, with at least 5 years in a leadership or technical lead role
  • Strong programming skills in high-performance languages
  • Deep understanding of inference frameworks and ML system deployment
  • Proven experience optimizing performance for large-scale machine learning systems, including a deep knowledge of SOTA model optimizations, hardware-software co-design, GPU acceleration, and HPC techniques
  • Excellent communication and collaboration skills
  • Experience leading teams working on high-throughput, low-latency ML serving systems
  • Experience collaborating with and leading global, cross-functional teams

Qualifications

  • Contributions to open-source ML or systems projects

Skills

  • Experience in designing and implementing scalable machine learning systems
  • Knowledge of inference frameworks such as TensorFlow, PyTorch, or ONNX
  • Experience with distributed systems and cloud platforms like Kubernetes, AWS, or Google Cloud
  • Understanding of hardware acceleration techniques and their impact on inference performance
  • Experience with monitoring and observability tools for machine learning systems

Benefits

  • Global access to mental health and financial wellness support and resources
  • Statutory and voluntary benefits including healthcare (medical, dental, and vision), life, accident, disability, commuter, and retirement options (401(k)/pension)

Pay

  • Competitive salary based on experience and qualifications

Schedule

  • Hybrid work approach with office days on Monday through Thursday and flexible Fridays for remote work

Similar jobs