Jobs · Engineering · California

Senior Solutions Architect, AI Infrastructure Enterprise ISVs

NVIDIA AI · Santa Clara, CA · 6 days ago
EngineeringFull-time

About the role

NVIDIA is seeking outstanding AI Solutions Architects to assist and support customers that are building solutions with our newest AI and accelerated computing technologies.

Responsibilities

  • Partner with ISVs on discovery, architecture reviews, technical deep dives, POCs, benchmarks, demos, and production deployment guidance
  • Advises on the design, build-out, and optimization of accelerated AI infrastructure, including large-scale clusters
  • Supports infrastructure design across compute, networking, storage, containers, observability, security, power, and data center operations
  • Drives adoption of systems monitoring, telemetry, and management tools to improve cluster utilization, reliability, performance and workload insight
  • Builds repeatable reference architectures, deployment guides, sizing guidance, benchmark reports, technical playbooks, demos and whitepapers
  • Travel up to 20% customer meetings may be required

Requirements

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience)
  • 8+ years of hands-on experience in AI infrastructure, accelerated computing, distributed systems, cloud infrastructure, high-performance computing, or machine learning platforms
  • In-depth knowledge of AI cluster orchestration, scheduling, automation and CI/CD deployment pipelines
  • Understanding of data center networking technologies such as InfiniBand, Ethernet, RDMA, network configuration or performance tuning
  • Familiarity with infrastructure requirements for AI workloads, including distributed training, inference serving, model deployment, storage performance, and cluster reliability
  • Excellent presentation, communication, problem-solving, documentation, and collaboration skills

Qualifications

  • Experience architecting AI factories, large GPU clusters, multi-node training environments, production inference platforms
  • Experience deploying LLM training, fine-tuning, RAG, and inference workflows on large-scale AI infrastructure
  • Experience evaluating cluster performance using benchmarks such as MLPerf, HPL, or workload-specific performance tests
  • Applications and systems-level knowledge of OpenMPI, NCCL, distributed training frameworks, and GPU communication patterns
  • Experience delivering technical training, workshops, whitepapers, blogs, or mentoring engineers, researchers, and customers on AI/HPC infrastructure

Skills

  • Strong experience designing, deploying, and operating accelerated computing infrastructure at scale
  • In-depth knowledge of AI cluster orchestration, scheduling, automation and CI/CD deployment pipelines
  • Understanding of data center networking technologies such as InfiniBand, Ethernet, RDMA, network configuration or performance tuning
  • Familiarity with infrastructure requirements for AI workloads, including distributed training, inference serving, model deployment, storage performance, and cluster reliability
  • Excellent presentation, communication, problem-solving, documentation, and collaboration skills

Benefits

  • Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
  • The base salary range is 184,000 USD - 287,500 USD.
  • You will also be eligible for equity and benefits.

Pay

  • Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
  • The base salary range is 184,000 USD - 287,500 USD.

Schedule

  • Full-time

Similar jobs