Senior Solutions Architect - AI Infrastructure
About the role
NVIDIA is building the world’s most groundbreaking and innovative accelerated computing platforms for AI and HPC. As a Senior Solutions Architect on the NVIDIA Cloud Partners team, you will focus on GPU, NVLink, and infrastructure design. You will assist with designs and architectures for large next-generation GPU-based clusters, enabling advanced AI supercomputers and enterprise AI infrastructure. This role bridges NVIDIA’s GPU and NVLink technology with engineering and field teams, supporting customers with demanding requirements. You will influence how leading AI companies, cloud providers, hyperscalers, research institutions, and enterprises build their infrastructure.
Responsibilities
- Partner with NVIDIA Cloud Partners in GPU cluster design and networking, conveying architecture and optimal processes for next-generation architectures.
- Guide NVIDIA Cloud Partners in cluster design, balancing design principles and situational limitations to create high-performance, supportable GPU clusters.
- Work closely with NVIDIA Cloud Partners to ensure successful first deployments of new products, including network architectures and topologies.
- Provide feedback on cluster design and workflows from customer/field perspectives to engineering teams.
- Perform hands-on debugging of cluster design, configuration, and performance issues, leveraging internal engineering expertise.
- Support NPI customer deployments with new GPU and networking architectures.
Requirements
- BS, MS, or PhD in Computer Science, Electrical Engineering, Computer Engineering, Physics, or related field (or equivalent experience).
- 8+ years of experience in cluster design, validation, and issue resolution, specifically on GPU and HPC clusters.
- Proven expertise in designing large-scale distributed systems, AI clusters, or HPC infrastructure.
- Ability to translate sophisticated engineering concepts into customer-ready documentation, diagrams, and reference material.
- Expertise in driving customer/partner issues to resolution with product and engineering teams.
- Ability to handle multi-functional communications across customer, product, support, and engineering teams.
Skills
- Experience leading large-scale AI Factory or HPC cluster bring-ups or builds.
- Hands-on experience with NVIDIA products, including GPUs, NVLink, and NVIDIA Networking, particularly debugging deployment issues.
- Knowledge of NCCL, MPI, IMEX, NMX, and collectives in distributed training as they pertain to cluster designs.
- External customer-facing skill set and background.
- Effective time management and ability to balance multiple tasks and customers while debugging and solving problems creatively.
Pay
The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.