Jobs · Engineering · New York

Senior Inference Engineer, GPU Kernel Optimization

NVIDIA · New York, NY · 1 mo ago
EngineeringFull-time

About the role

The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. These systems include GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic kernel optimization.

Responsibilities

  • Measure competing kernel implementations at real-silicon fidelity across the full configuration space that production LLM deployments demand.
  • Connect performance evidence to model-level serving economics, surfacing high-value optimization opportunities, and producing optimization policies for production inference deployments.
  • Apply AI-driven analysis to diagnose performance gaps, explore optimization opportunities across the kernel ecosystem, and validate findings with rigorous silicon measurements.
  • Collaborate closely with compiler, kernel, hardware, and framework teams to deliver upstream improvements and production-grade performance gains.

Requirements

  • Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 6+ years of relevant industry experience.
  • Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.
  • Strong Python and C++ skills with proven software engineering fundamentals.
  • Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
  • Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.
  • Workings knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.

Qualifications

  • Deep knowledge of SASS/PTX-level kernel analysis, compiler middle-end optimization, or GPU code generation pipelines (LLVM, MLIR, ptxas, or similar).
  • Track record shipping agentic systems end-to-end — tool invent, multi-agent orchestration, and silicon-verified validation — within a performance engineering or kernel optimization context.
  • Active contributions to open-source LLM inference or GPU kernel libraries (FlashInfer, Triton, CUTLASS, or similar).

Skills

  • Experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM.
  • Knowledge of GPU kernel optimization techniques including CUDA, CUTLASS, Triton, or equivalent.
  • Ability to analyze and optimize GPU kernels at the assembly level.
  • Proficiency in Python and C++.
  • Experience with GPU profiling tools such as CUPTI, NSYS, and NCU.
  • Understanding of model-level serving economics and optimization policies.

Benefits

  • NVIDIA offers highly competitive salaries and a comprehensive benefits package.
  • As you plan your future, see what we can offer to you and your family here.

Pay

Base salary range: 184,000 USD - 287,500 USD.

Schedule

Not specified.

Application Instructions

Applications for this job will be accepted at least until July 31, 2026.

Similar jobs