Manager, System Software Engineering - Factory
NVIDIA · Santa Clara, CA · 1 mo ago
EngineeringFull-time
What You’ll Be Doing
- Build, lead, mentor, and grow a global factory engineering team spanning the US and Taiwan — operating a follow-the-sun coverage model with on-site presence at ODM/CM partner factories.
- Own hiring, career development, calibration, and succession planning.
- Define Factory readiness scope and workflows for rack scale products coordinating multi-functionally with product management, technical architects and program management.
- Deliver those workflows through the validation matrix, ensuring delivered firmware and software is of the highest quality.
- Solutions must scale and be resilient.
- Own technical leadership for how firmware, software, and diagnostics releases reach factories building rack-scale systems.
- These systems include tightly coupled compute and switch trays.
- Build the end-to-end infrastructure and workflows that ensure every release arrives with efficient quality.
- Left-shift release quality: partner with all matrixed organizations — developers, SWQA, and product engineering — in a fast-moving environment with end-to-end CI/CD so that no bug is first found at a factory site.
- Enforce well-placed quality gates at every product landmark, publish and track indicators at a regular cadence, and report release progress to collaborators and executives.
- Own the factory escalation path: triage SLAs, 24×7 coverage, failure root-cause and deflection, and bonepile burn-down — minimizing line-down time through NPI ramps and mass production.
- Shape the team's roadmap and drive innovation with a strong focus on automation and AI-assisted validation and triage — automating station readiness, firmware-update flows, and log triage so senior engineering time shifts from setup to analysis.
- Continuously analyze factory processes, systems, and workflows to identify improvement and optimization opportunities; remove bottlenecks, document and publish standard operating procedures (SOPs), and ensure the team performs in the most efficient and transparent way against measurable targets.
What We Need To See
- 10+ overall years in the software industry with specialization in system software and/or firmware development.
- 3+ years of engineering management or technical leadership experience, including building and leading geographically distributed teams.
- BS, MS, or PhD in CS, CE, EE, or a related technical field — or equivalent experience.
- Proven track record of shipping scalable server products through factory ramps — from NPI bring-up to mass production — collaborating with hardware, firmware, manufacturing, diagnostics, and QA teams.
- Experience working with ODM/OEM partners to deliver quality servers and solutions for large-scale data centers.
- A self-starter who loves finding creative solutions to complicated problems, with excellent written and oral communication skills — including executive-level reporting — strong work ethic, and dedication to teamwork.
- Flexibility to work and communicate effectively across teams, partners, and time zones.
- Experience leading bring-up for sophisticated rack-scale compute architectures like GB200/GB300 NVL72.
- Familiarity with manufacturing test flows (L6/L10/L11/L12 stations), factory test coverage, and MES integration.
- Hands-on experience with x86/ARM system architecture and coding (C/C++, Python).
- Experience with SCM (Git, Perforce) and project management tools (Jira).
- Track record of integrating AI/LLM tooling into engineering workflows — for triage, validation, log analysis, or test generation.
- Experience standing up follow-the-sun support organizations with measurable response SLAs.