Jobs · Engineering · California

Senior Systems Software Engineer, Data Center Platform Enablement

Thomas To · Santa Clara, CA · 5 days ago
EngineeringFull-time

About the role

NVIDIA data center systems have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA Networking, NVIDIA Data Center CPUs, and a fully optimized NVIDIA AI and HPC software stack. We are seeking an excellent Senior System Engineer to work on bring up, integration, validation, and troubleshooting for compute tray platforms of GPU Racks — ensuring servers are fully functional and validated as per requirement before mass deployment in data centers.

Responsibilities

  • Work with a global team of architects and developers on NVIDIA AI server system (CPU, GPU, Memory, NIC, PCIe, NVMe SSD, Cooling, etc.) designs.
  • Work closely with hardware teams to influence hardware system design, reviewing schematics and board design to ensure speed-of-light execution.
  • Use simulation and emulation to left-shift early design and initiate new concepts and innovations.
  • Hands-on debug and support of early server prototypes bring-up and power-on test debug in datacenter system labs.
  • Factory and Manufacturing support: Support manufacturing flows, firmware updates, and diagnostic procedures. Ensure BOM change signoff and process optimization.
  • Issue Resolution & Customer Support: Drive root cause analysis and resolution of bring-up failures. Collaborate with partners, ODMs, and customers for technical support.
  • Work with design architects to develop proper diagnostic test tools and automation for qualifying the early system software and firmware stack.

Requirements

  • Strong experience in computer system and firmware architecture and design for server products.
  • Experience with CPLD or FPGA design and RTL.
  • Solid experience in end-to-end delivery of high-end enterprise server products from definition to customer deployment.
  • Solid understanding of low-level interfaces between SBIOS, BMC, and OS like I2C/SPI/PCIe/JTAG etc.
  • System boot and initialization, PCIe enumeration, high-speed I/O at the platform level for enterprise systems.
  • Experience working closely with HW teams, ODMs, and vendors to introduce and support server platforms.
  • Experience with C/C++ development, bash/python for scripting, and solid hands-on debugging skills in embedded Linux operating environments.
  • Experience using simulation such as QEMU and emulation tools to perform validation of early design work.
  • Experience accelerating design and debug by using AI tools.
  • Excellent written and oral communication skills, good work ethics, high sense of teamwork, and commitment to quality.
  • Self-starter who loves to find creative solutions to exciting problems.
  • Bachelor’s Degree or higher in Electrical Engineering or Computer Science (or equivalent experience), and 8+ years of experience with demonstrated strong ability as an individual contributor.

Skills

  • Experience with early design and power-on debug.
  • Proven leadership in tackling challenging issues and working in a fast-paced environment.

Pay

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.

Similar jobs