Senior Inference Engineer, GPU Kernel Optimization
NVIDIA · New York, NY · 1 mo ago
EngineeringFull-time
About the role
The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. These systems include GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic kernel optimization.
Responsibilities
- Measure competing kernel implementations at real-silicon fidelity across the full configuration space that production LLM deployments demand.
- Connect performance evidence to model-level serving economics, surfacing high-value optimization opportunities, and producing optimization policies for production inference deployments.
- Apply AI-driven analysis to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements.
- Collaborate closely with compiler, kernel, hardware, and framework teams to deliver upstream improvements and production-grade performance gains.
Requirements
- Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 6+ years of relevant industry experience.
- Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.
- Strong Python and C++ skills with proven software engineering fundamentals.
- Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
- Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.
- Workings knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.
Qualifications
- Deep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).
- Track record shipping agentic systems end-to-end — tool invent, multi-agent orchestration, and silicon-verified validation — within a performance engineering or kernel optimization context.
- Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).
Skills
- Experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM.
- Knowledge of GPU kernel optimization techniques including CUDA, CUTLASS, Triton, or equivalent.
- Ability to analyze and optimize GPU kernels at the assembly level.
- Proficiency in Python and C++.
- Experience with GPU profiling tools such as CUPTI, NSYS, and NCU.
- Understanding of model-level serving economics and optimization policies.
Benefits
- NVIDIA offers highly competitive salaries and a comprehensive benefits package.
- As you plan your future, see what we can offer to you and your family here.
Pay
Base salary range: 184,000 USD - 287,500 USD.
Schedule
Not specified.
Application Instructions
Applications for this job will be accepted at least until July 31, 2026.