Jobs · Business Development · California

Manager, Large Language Model Inference

NVIDIA AI · Santa Clara, CA · 1 mo ago
HybridBusiness DevelopmentFull-time

About the role

NVIDIA is seeking a highly skilled and driven Engineering Manager to take the lead in developing the next generation of LLM/VLM/VLA inference software technologies. This is a high-impact, hands-on leadership role at the intersection of deep technical expertise and world-class management. You will architect and guide a brilliant team of engineers who are building the core LLM inference runtime. Your work will be highly collaborative, interfacing directly with NVIDIA Researchers, GPU Architects, and other teams across the company to ensure we ship production-grade, lightning-fast software that sets the global standard for AI performance.

Responsibilities

  • Lead and grow a team responsible for specialized kernel development, runtime optimizations, and frameworks for LLM inference.
  • Drive the design, development, and delivery of production inference software, targeting NVIDIA's next-generation enterprise and edge hardware platforms.
  • Integrate cutting-edge technologies developed at NVIDIA and offer an intuitive developer experience for LLM deployment.
  • Lead software development execution, with responsibility for project planning, milestone delivery, and cross-functional coordination.

Requirements

  • MS, PhD, or equivalent experience in Computer Science, Computer Engineering, AI, or a related technical field.
  • 7+ overall years of overall software engineering experience, including 3+ years of technical leadership experience.
  • Proven ability to lead and scale high-performing engineering teams, especially across distributed and cross-functional groups.
  • Strong background in C++ or Python, with expertise in software design and delivering production-quality software libraries.
  • Demonstrated expertise in large language models (LLM) and/or vision language models (VLM).

Preferred qualifications

  • Deep understanding of GPU architecture, CUDA programming, and system-level performance tuning.
  • Background in LLM inference or working with frameworks such as TensorRT-LLM, vLLM, or SGLang.
  • Passion for building scalable, user-friendly APIs and enabling developers in the AI ecosystem.
  • Have a proven track record of growing and managing a team that encourages idea sharing, empowers team members, and provides opportunities for professional growth.

Pay

Base salary range is 184,000 USD - 287,500 USD for Level 2, and 224,000 USD - 356,500 USD for Level 3, with eligibility for equity and benefits.

Similar jobs

Large Language Model Specialist

Bright Vision TechnologiesSammamish, WA· 2 wk ago
Business Development$100k–$150k/yrapply on brightvisiontechnologies.applytojob.com

Large Language Model Specialist

Bright Vision TechnologiesAhwatukee, AZ· 1 mo ago
Business Development$100k–$150k/yrapply on brightvisiontechnologies.applytojob.com

Large Language Model Specialist

Bright Vision TechnologiesBellevue, WA· 1 mo ago
Business Development$100k–$150k/yrapply on brightvisiontechnologies.applytojob.com