Senior Solutions Architect, AI Infrastructure Enterprise ISVs
NVIDIA AI · Santa Clara, CA · 6 days ago
EngineeringFull-time
About the role
NVIDIA is seeking outstanding AI Solutions Architects to assist and support customers that are building solutions with our newest AI and accelerated computing technologies.
Responsibilities
- Partner with ISVs on discovery, architecture reviews, technical deep dives, POCs, benchmarks, demos, and production deployment guidance
- Advises on the design, build-out, and optimization of accelerated AI infrastructure, including large-scale clusters
- Supports infrastructure design across compute, networking, storage, containers, observability, security, power, and data center operations
- Drives adoption of systems monitoring, telemetry, and management tools to improve cluster utilization, reliability, performance and workload insight
- Builds repeatable reference architectures, deployment guides, sizing guidance, benchmark reports, technical playbooks, demos and whitepapers
- Travel up to 20% customer meetings may be required
Requirements
- BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience)
- 8+ years of hands-on experience in AI infrastructure, accelerated computing, distributed systems, cloud infrastructure, high-performance computing, or machine learning platforms
- In-depth knowledge of AI cluster orchestration, scheduling, automation and CI/CD deployment pipelines
- Understanding of data center networking technologies such as InfiniBand, Ethernet, RDMA, network configuration or performance tuning
- Familiarity with infrastructure requirements for AI workloads, including distributed training, inference serving, model deployment, storage performance, and cluster reliability
- Excellent presentation, communication, problem-solving, documentation, and collaboration skills
Qualifications
- Experience architecting AI factories, large GPU clusters, multi-node training environments, production inference platforms
- Experience deploying LLM training, fine-tuning, RAG, and inference workflows on large-scale AI infrastructure
- Experience evaluating cluster performance using benchmarks such as MLPerf, HPL, or workload-specific performance tests
- Applications and systems-level knowledge of OpenMPI, NCCL, distributed training frameworks, and GPU communication patterns
- Experience delivering technical training, workshops, whitepapers, blogs, or mentoring engineers, researchers, and customers on AI/HPC infrastructure
Skills
- Strong experience designing, deploying, and operating accelerated computing infrastructure at scale
- In-depth knowledge of AI cluster orchestration, scheduling, automation and CI/CD deployment pipelines
- Understanding of data center networking technologies such as InfiniBand, Ethernet, RDMA, network configuration or performance tuning
- Familiarity with infrastructure requirements for AI workloads, including distributed training, inference serving, model deployment, storage performance, and cluster reliability
- Excellent presentation, communication, problem-solving, documentation, and collaboration skills
Benefits
- Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
- The base salary range is 184,000 USD - 287,500 USD.
- You will also be eligible for equity and benefits.
Pay
- Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
- The base salary range is 184,000 USD - 287,500 USD.
Schedule
- Full-time