Senior Deep Learning Software Engineer, Inference
NVIDIA · California, United States · 3 wk ago
RemoteRemoteEngineeringFull-time
NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference to help design, build, and optimize the GPU-accelerated software that powers today’s most sophisticated AI applications. Our team develops and maintains high-performance open-source frameworks at the forefront of efficient large-scale model serving and inference.
Responsibilities
- Performance optimization, analysis, and tuning of DL models in domains like LLM, Multimodal, and Generative AI.
- Scale performance of DL models across different architectures and types of NVIDIA accelerators.
- Contribute features and code to NVIDIA’s inference libraries, including vLLM, SGLang, FlashInfer, and LLM software solutions.
- Work with cross-collaborative teams across frameworks, NVIDIA libraries, and inference optimization solutions.
Requirements
- Masters or PhD (or equivalent experience) in Computer Engineering, Computer Science, EECS, or AI.
- 5+ years of relevant software development experience.
- Excellent C/C++ programming and software design skills; Agile and Python experience is a plus.
- Prior experience with training, deploying, or optimizing DL model inference in production is a plus.
- Background in performance modeling, profiling, debug, code optimization, or architectural knowledge of CPU/GPU is a plus.
- GPU programming experience (CUDA, OAI Triton, or CUTLASS) is a plus.
Skills
- Contribute to deep learning software projects (e.g., PyTorch, vLLM, SGLang) to drive advancements in the field.
- Experience with Multi-GPU communications (NCCL, NVSHMEM).
Benefits
NVIDIA offers highly competitive salaries and a comprehensive benefits package, widely considered among the best in the technology industry.
Pay
Base salary ranges: $152,000–$241,500 USD (Level 3) or $184,000–$287,500 USD (Level 4), plus equity and benefits.
JR2003655