Senior Inference Engineer, GPU Kernel Optimization
NVIDIA AI · Santa Clara, NY · 4 wk ago
EngineeringFull-time
About the role
The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. These systems include GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic kernel optimization.
Responsibilities
- Measure competing kernel implementations at real-silicon fidelity across the full configuration space that production LLM deployments demand.
- Connect performance evidence to model-level serving economics, surfacing high-value optimization opportunities, and producing optimization policies for production inference deployments.
- Apply AI-driven analysis to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements.
- Collaborate closely with compiler, kernel, hardware, and framework teams to deliver upstream improvements and production-grade performance gains.
Requirements
- Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 6+ years of relevant industry experience.
- Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.
- Strong Python and C++ skills with proven software engineering fundamentals.
- Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
- Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.
- Work experience with GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.
Qualifications
- Deep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).
- Track record shipping agentic systems end-to-end — tool invent, multi-agent orchestration, and silicon-verified validation — within a performance engineering or kernel optimization context.
- Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).
Skills
- Expertise in LLM inference frameworks such as TRT-LLM, SGLang, or vLLM.
- Proficiency in GPU kernel optimization techniques including CUDA, CUTLASS, Triton, or equivalent.
- Strong background in computer science and engineering with a focus on performance engineering and kernel optimization.
Benefits
- Comprehensive benefits package including health insurance, retirement plans, and paid time off.
- Competitive salary range of $184,000 - $287,500 based on location, experience, and comparable market rates.
- Equity compensation for selected candidates.
Pay
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is $184,000 - $287,500.
Schedule
Full-time position.