Senior Performance Engineer - DGX Cloud
NVIDIA · Santa Clara, CA · 1 wk ago
EngineeringFull-time
About the role
Joining NVIDIA's DGX Cloud AI Efficiency Team means advancing the performance, efficiency, and resiliency of large-scale AI workloads. We help AI researchers and platform teams understand end-to-end behavior across GPUs, networking, storage, and software stacks.
Responsibilities
- Analyze end-to-end performance of large-scale AI workloads across compute, network, storage, and software stacks.
- Design and execute rigorous performance studies to establish baselines, diagnose regressions, and quantify bottlenecks.
- Define performance and efficiency evaluation methodologies, benchmarks, and success metrics for AI workloads.
- Use profiling, observability, and data analysis to turn performance measurements into actionable optimization plans.
- Partner with deep learning engineers, platform teams, and GPU architects to validate and deliver performance improvements.
- Communicate performance findings, tradeoffs, and recommendations clearly to influence system and software design decisions.
Requirements
- BS or higher degree in computer science, computer engineering, or a related field (or equivalent experience).
- 12+ years of experience in strong programming skills in C++ and Python, with the ability to build reliable analysis and automation workflows.
- Solid foundation in operating systems, computer architecture, and distributed systems.
- Experience with performance engineering, benchmarking, profiling, and optimization of complex software or systems.
- Ability to communicate technical findings, prioritize high-impact work, and build alignment across teams.
Qualifications
- Experience analyzing large-scale AI clusters or distributed training and inference workloads.
- Experience with CUDA, GPU computing systems, and GPU performance analysis.
- Hands-on experience with deep learning frameworks such as PyTorch or JAX/XLA.
- Deep understanding of system-level performance analysis, workload characterization, and optimization.
Skills
- Strong programming skills in C++ and Python.
- Experience with performance engineering, benchmarking, profiling, and optimization.
- Understanding of operating systems, computer architecture, and distributed systems.
- Experience with CUDA, GPU computing systems, and GPU performance analysis.
- Hands-on experience with deep learning frameworks such as PyTorch or JAX/XLA.
- Deep understanding of system-level performance analysis, workload characterization, and optimization.
Benefits
- NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization.
- The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services.
- We enable amazing creativity and discovery, and power what were once science fiction inventions, from artificial intelligence to autonomous cars.
Pay
- Base salary range: $224,000 - $356,500 for Level 5, and $272,000 - $431,250 for Level 6.
Schedule
- Full-time position.