Senior Research Engineer - Enterprise Products
Thomas To · King County, WA · Today
EngineeringFull-time
About the role
We are looking for a Senior Research Engineer passionate about Generative AI inference. Join NVIDIA in changing how people infuse AI into products and services. Our team develops optimized inferencing technologies to support growing generative AI needs, contributing to the entire machine learning lifecycle—from conceptualization and applied research to engineering and deployment. Collaborate with research teams, engineers, and the open-source community.
Responsibilities
- Design and evaluate routing policies for LLM traffic to optimize the use of mixture-of-model systems.
- Build and run agentic benchmarks (e.g., Terminal-Bench) to measure algorithm quality, and convert results into calibration data and routing profiles.
- Contribute to open-source repositories: design docs, code reviews, documentation, and community engagement.
- Collaborate with engineering teams across NVIDIA to ensure seamless integration with the accelerated serving stack.
Requirements
- Bachelor’s or Master’s degree in Computer Science or equivalent experience.
- 8+ years of industry experience in Deep Learning frameworks (PyTorch or TensorFlow).
- Experience designing or running LLM evaluations/benchmarks—ideally agentic ones—and drawing statistically sound conclusions.
- Understanding of modern techniques in Machine Learning, Deep Neural Networks, Natural Language Processing, or Speech Recognition.
- Empirical research mindset: forming hypotheses, running calibrations, and iterating on results.
- Strong communication and interpersonal skills, with the ability to work in a dynamic, distributed team.
- Strong computer science fundamentals: algorithms, data structures, computational complexity, parallel and distributed computing, and system software.
Skills
- History of mentoring junior engineers and interns (a plus).
- Desire to constantly grow and learn new things.
- Experience architecting or developing large-scale distributed systems for deep learning (stands out).
- Agentic benchmark creation and publications (stands out).
- Knowledge of CPU and/or GPU architecture (stands out).
- GPU programming (CUDA) (stands out).
Pay
Base salary range: $192,000–$304,750 USD (Level 4) or $224,000–$356,500 USD (Level 5), determined by location, experience, and comparable roles. Eligible for equity and benefits.