Jobs · Sales · California

Senior Solutions Architect, AI Factory Observability and Visualization - NVIS

NVIDIA · Santa Clara, CA · 1 wk ago
SalesFull-time

Skills

  • Experience with AI factory or large-scale AI infrastructure build, deployment, or operations
  • Experience with HPC systems engineering, SRE, or systems analysis for GPU-accelerated environments
  • Experience building automation and data pipelines that feed dashboards and reporting at scale
  • Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions

Responsibilities

  • Run AI factory validation tools, microbenchmarks, and workloads provided by the team, and interpret results to assess system health and performance
  • Gain a comprehensive understanding of the system from start to finish, including network topology, interconnects, and compute
  • Establish what "healthy" represents across the stack — the metrics, logs, and signals that confirm a system is functioning well, and the thresholds that show it isn't
  • Build and extend the telemetry surface across hardware, fabric, and workload, crafting how data is collected, transformed, stored, and surfaced
  • Serve as the observability expert, investigating gaps in visibility to ensure it reflects true system behavior
  • Develop automation (Python, Shell) for collecting, transforming, and presenting system and network data
  • Recommend improvements to system visibility, data sources, and reporting that give teams clearer insight
  • Collaborate with hardware, software, networking, datacenter, and product groups to ready HPC systems and AI factories for customer deployment, contributing documentation and readiness materials throughout the process

Requirements

  • Bachelor's degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field
  • 6+ years of experience managing Linux-based systems in HPC, distributed systems, or large AI/ML settings
  • Hands-on experience with the architecture of multi-GPU and/or multi-node clusters, including networking and interconnects
  • Solid grasp of how HPC and AI factory systems fit together end to end, from network fabric through compute
  • Proficiency with Python and Shell/Bash for scripting, automation, and tooling
  • Practical experience working with observability systems (e.g., Prometheus, Grafana, Loki, or similar), including building custom exporters or collectors, setting up alerts, and handling metric cardinality and retention on a large scale
  • Experience transforming metrics, logs, and traces into clear, actionable insight for complex distributed environments
  • Familiarity with GPU and fabric telemetry (e.g., DCGM, NVLink, InfiniBand/Ethernet fabric counters) and using it to diagnose performance regressions
  • Strong communication skills and the ability to work effectively with cross-functional teams

Qualifications

  • Experience with AI factory or large-scale AI infrastructure build, deployment, or operations
  • Experience with HPC systems engineering, SRE, or systems analysis for GPU-accelerated environments
  • Experience building automation and data pipelines that feed dashboards and reporting at scale
  • Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions

Pay

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

Schedule

The role is remote.

Similar jobs