Senior Production Engineer - DGX Cloud
About the role
We are seeking a Senior Production Engineer to join our DGX Cloud team. This role involves scaling up NVIDIA's AI infrastructure and contributing to the development of leading infrastructure solutions for various AI-based applications.
Responsibilities
- Implementing monitoring and health management capabilities for large-scale GPU clusters.
- Working with teams across NVIDIA to ensure production AI clusters run reliably and consistently.
- Evaluating system failures and improving services based on a defined incident management process.
- Managing and automating large-scale distributed systems independent of cloud providers.
- Advanced hands-on experience and deep understanding of cluster management systems like Kubernetes, Slurm, and Bright Cluster Manager.
- Ensuring reliable and performant AI infrastructure through operational excellence.
Requirements
- Direct experience in a Production Engineering/DevOps/SRE role within a highly technical organization with demonstrable impact.
- Highly motivated with strong communication skills, able to work with multi-functional teams and coordinate effectively across organizational boundaries and geographies.
- 8+ years of experience in similar roles and experience on large-scale production systems.
- Experience with Production Engineering/DevOps/SRE principles, tools, and techniques.
- A BS in Computer Science, Engineering, Physics, Mathematics, or a comparable degree or equivalent experience.
- Technical knowledge including a systems programming language (e.g., Go, Python) and a solid understanding of data structures and algorithms.
Qualifications
- Significant experience with site reliability principles and techniques including reliability assessments, incident management processes, production system observability, monitoring and alerting, automated deployments, and toil elimination.
- Significant contributions to the codebase, with a strong focus on software engineering.
- Out-of-the-box thinking and the ability to provide new ideas with strong execution bias.
- Constantly challenging, improving, and evolving for the better.
Skills
- Strong technical competency in managing and automating large-scale distributed systems independent of cloud providers.
- Advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, Slurm, Bright Cluster Manager).
- Proven operational excellence in maintaining reliable and performant AI infrastructure.
Benefits
- Competitive base salary ranging from $168,000 to $270,250 for Level 4, and $208,000 to $333,500 for Level 5.
- Eligibility for equity and benefits.
Pay
- Base salary determined based on location, experience, and the pay of employees in similar positions.
Schedule
- Full-time position.
Contact Information
Applications for this job will be accepted at least until July 10, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.