Senior System Architect, Enterprise Reference Architectures
NVIDIA · Austin, TX · 3 wk ago
EngineeringFull-time
About the role
The Senior System Architect will define, design, and validate enterprise AI factory reference architectures. This role bridges product strategy, system architecture, customer requirements, and hands-on infrastructure validation.
Responsibilities
- Define and drive full-stack enterprise AI factory baseline architectures across compute, networking, storage, virtualization, orchestration, security, observability, and NVIDIA AI software.
- Architect reference builds for end-to-end software systems and platforms based on scalable, portable, and resilient builds and drive evaluation of different platforms.
- Develop and validate scalable cluster designs for enterprise AI/ML systems, including on-premises and hybrid-cloud deployments.
- Evaluate tradeoffs across performance, scalability, resiliency, manageability, security, power, cooling, cost, TCO, and operational complexity.
- Build and detail comprehensive API strategies, middleware integrations, and orchestration workflows to ensure seamless communication across distributed enterprise software.
- Integrate with non-cloud technologies and third-party vendor products/services.
Requirements
- Bachelor's degree or equivalent experience.
- 12+ years of software or infrastructure architecture experience.
- 3+ years of experience with micro-services architectures.
- Experience with system architecture, performance, and networking specifically developing and deploying distributed solutions.
- Proven experience designing, building, or maintaining complex AI, HPC, or cloud-native infrastructure from rack-scale systems through full data center deployments.
- Proven expertise in crafting and deploying complete software platforms spanning edge, on-premises, and cloud environments.
- Strong understanding of GPU-accelerated systems, high-density servers, PCIe, NVLink, CPU/GPU platform design, rack integration, power, cooling, mechanical constraints, and data center deployment realities.
- Deep networking knowledge across Ethernet, InfiniBand, RDMA, RoCE, routing, switching, congestion control, network segmentation, and high-performance cluster fabrics.
- Deep proficiency in modern integration patterns, event-driven architectures (e.g., Kafka, RabbitMQ), and container orchestration (e.g., Kubernetes, Docker).
- Strong grasp of cloud-native systems with emphasis on high availability, scalability, and security in compute environments.
Qualifications
- Ability to work optimally with NVIDIA and its partners to deliver prescriptive, validated, and scalable enterprise AI factory architectures that reduce deployment complexity and accelerate time to value.
- Ability to connect product vision, customer feedback, and deep system architecture to develop practical designs that enterprises can deploy, operate, and expand in production.
Skills
- Ability to work optimally with NVIDIA and its partners to deliver prescriptive, validated, and scalable enterprise AI factory architectures that reduce deployment complexity and accelerate time to value.
- Ability to connect product vision, customer feedback, and deep system architecture to develop practical designs that enterprises can deploy, operate, and expand in production.
Benefits
- NVIDIA is widely considered one of the technology world’s most desirable employers.
- We have some of the world's most forward-thinking and hardworking people on our team.
Pay
Base salary range: $208,000 - $327,750 for Level 5, and $240,000 - $379,500 for Level 6.
Schedule
Applications for this job will be accepted at least until June 2, 2026.