Jobs · Engineering · California

Distinguished Engineer – Data Center System Software Architect

NVIDIA · Santa Clara, CA · 1 mo ago
EngineeringFull-time

About the role

NVIDIA is seeking a technical architect to lead the architecture of their data center systems, including system software components like firmware, kernel drivers, operating systems, and user mode drivers. This role involves working closely with major customers, cloud service providers, and internal component leads to develop and deploy next-generation data center products.

Responsibilities

  • Serve as the primary technical point of contact for major customers, leading technological discussions, defining Key Performance Indicators (KPIs), gathering requirements, and addressing complex technical queries.
  • Lead technical innovation and strategic collaborations with major hyperscalers to architect next-generation data center products.
  • Align NVIDIA's roadmap with major customers' requirements through direct engagement.
  • Develop and drive adoption of new technologies and protocols.
  • Make critical technical decisions in ambiguous situations, mitigating risks through left-shift strategies.

Requirements

  • Deep expertise in scalable and performant server system architecture, focusing on SW/HW interfaces.
  • Extensive experience with complex system software for accelerators (GPUs, DPUs, FPGAs).
  • Mastery of system firmware (SBIOS, OpenBMC), embedded systems, and Linux kernel internals.
  • Proficiency in Out-of-Bound and In-Bound management architectures, device management protocols (e.g., MCTP, PLDM, SPDM, RDE) and system management protocols (Redfish, IPMI).
  • Extensive knowledge of networking technologies and protocols, including TCP/IP, Ethernet, InfiniBand, as well as advanced switching and routing concepts.
  • Experience collaborating with platform security experts to define tradeoffs between security and ease of use.
  • Demonstrated success in leading complex, cross-functional projects to completion, showcasing the ability to influence and achieve results without direct authority in large-scale, collaborative environments.
  • Knowledge of cloud and cluster level deployment and management systems.
  • Familiarity with NVIDIA HPC programming models and libraries (CUDA, cuDNN, DOCA).
  • Knowledge of enterprise storage architectures and distributed parallel processing paradigms.

Qualifications

  • BS or MS degree in Computer Science, Electrical Engineering or related field (or equivalent experience).
  • 20+ years of experience in system architecture and design.

Skills

  • Knowledge of cloud and cluster level deployment and management systems.
  • Participation and contributions in standards bodies such as OCP and DMTF.
  • Familiarity with NVIDIA HPC programming models and libraries (CUDA, cuDNN, DOCA).
  • Knowledge of enterprise storage architectures and distributed parallel processing paradigms.

Benefits

NVIDIA offers competitive compensation, including a base salary range of $320,000 - $488,750, along with equity and comprehensive benefits. Applications for this position are open until July 4, 2026.

Pay

Base salary range: $320,000 - $488,750

Schedule

Not specified

Benefits

Not specified

Skills

  • Knowledge of cloud and cluster level deployment and management systems.
  • Participation and contributions in standards bodies such as OCP and DMTF.
  • Familiarity with NVIDIA HPC programming models and libraries (CUDA, cuDNN, DOCA).
  • Knowledge of enterprise storage architectures and distributed parallel processing paradigms.

Benefits

Not specified

Similar jobs