Technical Program Manager, Cloud Infrastructure
NVIDIA · Seattle, WA · Today
HybridFull-time
What You'll Be Doing
As a DGX Cloud Technical Program Manager, you'll be a key partner to our Engineering, Infrastructure, Software teams and their leadership, driving critical programs related to AI capacity enablement and management.
- Develop and mature foundational capabilities and processes for DGX Cloud, spanning critical areas such as cluster/capacity bring-up including CPU, storage, networking and compute requirements to support GPUs.
- Work in close coordination with storage engineering and network engineering teams to define and communicate requirements to CSP (Cloud Service Providers) and NCP’s (NVIDIA Cloud Providers).
- Drive alignment and a POR for capacity blocks based on workload needs.
- Drive early engagement with CSP (Cloud Service Providers) and NCP’s (NVIDIA Cloud Providers) to understand their managed storage, network solutions and influence alignment with NVIDIA Cloud roadmap.
- Gather technical requirements, develop comprehensive roadmaps, establish clear milestones, and ensure adherence to our Product Lifecycle (PLC) process.
- Manage ongoing capacity operations and the engineering engagement with CSP (Cloud Service Providers) and NCP’s (NVIDIA Cloud Provider) partners, collaborating closely with engineering leads.
- Focus on availability, maintenance and other critical performance indicators.
- Partner closely within NVIDIA to understand workload requirements and related hardware and infrastructure needs. This includes speeds and feeds to optimize infrastructure readiness with cloud vendors and NVIDIA Cloud Providers.
- Leverage Jira and other program management platforms to instill rigor and structure in the management of engineering deliverables.
- Identify and drive opportunities to onboard the adoption of third-party and in-house cloud infrastructure solutions for deployments, support, security, compliance and observability across DGX Cloud.
- Establish key performance indicators (KPIs) and quantitatively demonstrate the value and impact delivered by your programs.
- Proactively identify, resolve, and mitigate risks and issues that could affect scope, schedule, and quality across all program aspects.
- Encourage a culture of continuous improvement, consistently seeing opportunities for process improvements within our cloud infrastructure operations.
What We Need To See
- 10+ years of technical program management experience.
- You have driven the planning and execution of large-scale cloud infrastructure programs with outside organizations.
- You focus strongly on software engineering projects within a matrixed organization.
- Extensive hands-on experience in cloud infrastructure, preferably gained from working at a major Cloud Service Provider (CSP).
- Domain knowledge in the bring-up and end to end operations of compute, storage and GPU (including common failure points at the HW and SW levels).
- Expert-level proficiency with Jira, Smartsheet, or similar program management tools, with the ability to expertly guide engineering teams on their use of the tools.
- Outstanding strategic and tactical thinking abilities, coupled with a strong capacity to build consensus and drive program success.
- Comfort and efficiency in growing within ambiguous environments.
- Possess excellent communication and technical presentation skills, particularly for executive audiences.
- BS or MS in Electrical Engineering or Computer Science, or equivalent experience.
Ways To Stand Out From The Crowd
- In depth knowledge of NVIDIA GPU products, including deployment and bring-up.
- Working knowledge of various cloud technologies (Kubernetes, API integration, Terraform, etc).
- A highly enthusiastic, energetic, responsive, and passionate individual with a keen eye for identifying process improvement opportunities.
- Much significant experience with productivity tools and process automation is a major plus.
- Deep familiarity with cloud-native product / services environments and familiarity with AI, ML infrastructure, and cloud/services.
Pay & Benefits
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 258,750 USD for Level 4, and 200,000 USD - 322,000 USD for Level 5. You will also be eligible for equity and benefits.
Application Information
Applications for this job will be accepted at least until July 26, 2026.