Jobs · Engineering

Principal Software Engineer, Rack-Scale System Software — CSP Engagements

NVIDIA · California, United States · 2 wk ago
RemoteRemoteEngineeringFull-time

About the role

Join our CSP Engagements team as the technical focal point for rack-scale system software and firmware. In this role, you will collaborate with NVIDIA's cross-functional rack-scale system SW/FW engineering teams, providing dedicated CSP-facing technical leadership. Your focus will be on system-level software that manages, monitors, and recovers the rack as a whole, including fabric management, GPU/NVSwitch error handling and recovery, health telemetry APIs, firmware update orchestration, and software-driven serviceability.

Responsibilities

  • Drive rack-scale SW/FW architecture alignment across CSP engagements, including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot-plug support, component isolation, firmware-driven recovery), and multi-component firmware orchestration.
  • Lead technical work streams with CSP engineering teams on rack-scale system software, ensuring deep understanding of fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW-controlled recovery operations.
  • Capture and synthesize CSP engineering feedback on rack-scale system software, including health monitoring APIs, SW-driven serviceability workflows, firmware update orchestration, and error recovery behavior. Champion this feedback into NVIDIA's architecture decisions.
  • Collaborate with multi-functional teams to ensure customer operational requirements are reflected in system software and firmware development.
  • Identify cross-CSP patterns in rack-scale SW/FW issues, error handling behavior, and system configuration practices. Drive documentation, tooling, and test strategy improvements as a result.
  • Collaborate with execution teams on left-shift strategy, ensuring customer-side SW/FW integration work is identified early and completed ahead of hardware availability.
  • Make critical technical decisions on rack-scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams.

Requirements

  • 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering.
  • BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience).
  • Deep understanding of rack-scale system software challenges: multi-component coordination, error propagation, health monitoring, and serviceability/reliability.
  • Experience with fabric management software, cluster management, or system-level orchestration frameworks.
  • Familiarity with firmware architectures and update lifecycle management (multi-component update sequencing, rollback, recovery).
  • Understanding of error handling and recovery design patterns in distributed systems, including fault isolation, retry policies, and graceful degradation.
  • Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet-level observability.
  • Customer obsession with a genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience.
  • Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority.
  • Strong communication skills, with the ability to translate complex system software architecture into actionable mentorship for customer engineering teams.

Skills

  • Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus.
  • Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software.
  • Background in system software for large-scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms).
  • Experience crafting error handling and recovery frameworks for multi-component systems (hundreds or thousands of coordinating devices).
  • Familiarity with GPU or accelerator fleet operations, including driver lifecycle, firmware rollout strategies, and health-based scheduling.
  • Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale.

Pay

The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits.

Similar jobs