Jobs · Engineering

Deep Learning Software Engineer, Inference - New College Grad 2026

NVIDIA · Washington, United States · 2 days ago
RemoteRemoteEngineeringFull-time

About the role

NVIDIA is seeking a Software Engineer specializing in Deep Learning Inference to join a dynamic team focused on advancing AI technologies. The ideal candidate will contribute to the development and optimization of GPU-accelerated software that powers cutting-edge AI applications.

Responsibilities

  • Design, build, and optimize GPU-accelerated software for efficient large-scale model serving and inference.
  • Contribute features and code to NVIDIA’s inference libraries, including vLLM and SGLang, and software solutions like FlashInfer and LLM.
  • Work with cross-functional teams to develop innovative solutions for frameworks, NVIDIA libraries, and inference optimization.
  • Identify and drive performance improvements for state-of-the-art Language Model (LLM) and Generative AI models across various NVIDIA accelerators, from datacenter GPUs to edge SoCs.
  • Implement and optimize model serving pipelines using open-source tools and plugins such as CUTLASS, OAI Triton, NCCL, and CUDA kernels.
  • Perform performance modeling, profiling, debugging, and code optimization.

Requirements

  • Pursuing or recently completed a Master's or PhD in Computer Engineering, Computer Science, EECS, AI, or related field, or equivalent experience.
  • Software development experience.
  • Excellent C/C++ programming and software design skills.
  • Agile software development skills are helpful, and Python experience is a plus.
  • Prior experience with training, deploying, or optimizing the inference of DL models in production.
  • Prior background with performance modeling, profiling, debugging, and code optimization, or architectural knowledge of CPU and GPU.
  • Experience with GPU programming (CUDA, OAI TRITON, or CUTLASS).

Qualifications

  • Experience with deep learning software projects, such as PyTorch, vLLM, and SGLang.
  • Contributions to advancements in the field of deep learning.
  • Experience with Multi-GPU communications (NCCL, NVSHMEM).

Benefits

Base salary range: $124,000 - $195,500 for Level 2, and $152,000 - $241,500 for Level 3. Eligible for equity and benefits.

Pay

Base salary will be determined based on location, experience, and comparable positions.

Schedule

Full-time position.

How to Apply

Applications for this job will be accepted at least until July 26, 2026.

Similar jobs