Inference Performance Engineer, AI Inference Configuration Optimization
NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position provides an opportunity to employ optimization knowledge in an autonomous optimization framework used by AI agents to run benchmark, profile, and tune processes.
About the role
If you enjoy extracting maximum performance from GPUs and scaling your skills beyond individual efforts, this role is a great fit. You will distill performance instincts into reusable skills, workflows, and evidence-backed methodologies that AI agents can execute autonomously.
Responsibilities
- Distill performance instincts into reusable skills, workflows, and evidence-backed methodologies for autonomous AI agents.
- Review agent-generated experiments, validate findings, and curate best-known configurations.
- Improve AI inference workload performance by exploring configuration options, parallelism techniques, batching, KV cache handling, quantization, and speculative decoding settings to increase throughput-per-GPU and user interactivity.
- Measure and optimize aggregated and disaggregated serving architectures across TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA's latest GPU platforms.
- Profile workloads using Nsight Systems, kernel traces, and internal analysis tools; use roofline and speed-of-light analysis to identify credible headroom and drive fixes from hypothesis to measured wins.
- Land improvements upstream: serving framework patches, optimized kernels, and deployment recipes that advance the public Pareto frontier while maintaining strict model correctness.
- Collaborate with TensorRT-LLM, SGLang, vLLM, kernel, benchmarking, and GPU architecture teams to convert profiling insights into delivered performance improvements.
Requirements
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Math, or a related field, or equivalent experience.
- 3+ years of relevant engineering experience.
- Extensive knowledge of AI model execution efficiency and optimization, covering continuous batching, throughput-latency tradeoffs, KV cache and memory limitations, parallel processing techniques, MoE serving, quantization, and meeting serving SLAs.
- Hands-on experience benchmarking and profiling GPU workloads using tools such as Nsight Systems, Nsight Compute, CUPTI, or PyTorch profiler, and interpreting kernel-level performance data.
- Strong Python engineering skills and ability to navigate and modify large C++/CUDA serving codebases.
- Rigorous experimental methodology with controlled single-variable comparisons, reproducible benchmarks, and evidence-backed optimization decisions.
- Strong written and verbal communication skills to explain performance tradeoffs clearly to both humans and documentation for autonomous systems.
Skills
Ways to stand out from the crowd:
- Direct contributions to TensorRT-LLM, vLLM, SGLang, FlashInfer, Dynamo, or comparable inference frameworks.
- Experience with disaggregated serving, wide expert-parallel MoE inference, KV cache transfer, or NCCL/NIXL/NVSHMEM communication at multi-node scale.
- CUDA kernel authorship or optimization experience on Hopper/Blackwell architectures, focusing on Tensor Cores, TMA, and warp specialization.
- Proven results on public inference benchmarks such as MLPerf Inference or SemiAnalysis InferenceX.
- Experience building or operating agentic AI workflows to automate engineering tasks.
Pay
Your base salary will be determined based on location, experience, and the pay of employees in similar positions. The base salary range is:
- 124,000 USD - 195,500 USD for Level 2
- 152,000 USD - 241,500 USD for Level 3
You will also be eligible for equity and benefits.
Benefits
NVIDIA offers highly competitive salaries and a comprehensive benefits package. Learn more about what we can offer to you and your family at www.nvidiabenefits.com.