Jobs · Engineering · California

Senior Software Engineer - Manufacturing and Factory

Thomas To · Santa Clara, CA · 5 days ago
EngineeringFull-time

About the role

NVIDIA is tapping into the unlimited potential of AI to define the next era of computing—an era in which our GPU acts as the brains of computers, robots, and self-driving cars. We’re searching for a highly motivated, technical leader to design, drive, and operationalize rack-scale factory and deployment flows for next-generation data center products. The ideal candidate will combine deep systems expertise, decisive technical leadership, and a passion for building reliable, debuggable, and scalable manufacturing and deployment solutions.

Responsibilities

  • Lead and drive rack-scale/L11 flows for factory and initial data center deployment.
  • Design and implement end-to-end factory workflows, including firmware flashing sequences, security provisioning, and deployment of software mitigations.
  • Collaborate with data center architects, ODMs, and OEMs to define factory and data center requirements that ensure efficient and reliable production ramp.
  • Champion reliability, debuggability, and optimization in firmware, diagnostic, and deployment tool design.
  • Drive pre-silicon readiness for factory & manufacturing workflows for rack-scale products using NVIDIA's industry-leading simulation & emulation technology.
  • Mentor architects and engineering teams to grow them into future leaders.
  • Make key technical decisions even when faced with ambiguity.

Requirements

  • BS or MS degree in Computer Engineering, Computer Science, or related degree or equivalent experience.
  • 8+ years in system architecture and design.
  • Deep experience in designing architecture for scalable and performant server systems, particularly at the SW/HW interface.
  • Strong understanding of networking technology & protocols (e.g., Ethernet, InfiniBand).
  • Previous experience working with complex system software for accelerators such as GPUs, DPUs, or FPGAs.
  • Expertise in out-of-band and in-band management architectures.
  • Knowledge of system management protocols such as Redfish and IPMI.
  • Demonstrable experience in implementing left-shift strategy to de-risk program execution.
  • Excellent written and verbal communication skills.

Skills

  • Knowledge of large-scale cloud and cluster-level deployment and management systems.
  • Demonstrated track record of leading data center products across the entire lifecycle, spanning inception, pre-silicon development, post-silicon bring-up, manufacturing, and deployment.

Pay

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.

Similar jobs