Jobs · Engineering · New York

Senior Inference Engineer, GPU Kernel Optimization

NVIDIA AI · Santa Clara, NY · 4 wk ago
EngineeringFull-time

About the role

The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. These systems include GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic kernel optimization.

Responsibilities

  • Measure competing kernel implementations at real-silicon fidelity across the full configuration space that production LLM deployments demand.
  • Connect performance evidence to model-level serving economics, surfacing high-value optimization opportunities, and producing optimization policies for production inference deployments.
  • Apply AI-driven analysis to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements.
  • Collaborate closely with compiler, kernel, hardware, and framework teams to deliver upstream improvements and production-grade performance gains.

Requirements

  • Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 6+ years of relevant industry experience.
  • Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.
  • Strong Python and C++ skills with proven software engineering fundamentals.
  • Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
  • Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.
  • Work experience with GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.

Qualifications

  • Deep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).
  • Track record shipping agentic systems end-to-end — tool invent, multi-agent orchestration, and silicon-verified validation — within a performance engineering or kernel optimization context.
  • Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).

Skills

  • Expertise in LLM inference frameworks such as TRT-LLM, SGLang, or vLLM.
  • Proficiency in GPU kernel optimization techniques including CUDA, CUTLASS, Triton, or equivalent.
  • Strong background in computer science and engineering with a focus on performance engineering and kernel optimization.

Benefits

  • Comprehensive benefits package including health insurance, retirement plans, and paid time off.
  • Competitive salary range of $184,000 - $287,500 based on location, experience, and comparable market rates.
  • Equity compensation for selected candidates.

Pay

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is $184,000 - $287,500.

Schedule

Full-time position.

Similar jobs