Technical Project Manager, GPU Infrastructure Deployment
Position Summary
We are seeking an experienced Technical Project Manager (TPM), GPU Infrastructure Deployment to drive the planning, coordination, and execution of large-scale GPU cluster infrastructure deployment programs. This role partners across multiple engineering teams to manage project schedules, track milestones, resolve blockers, and communicate status across technical teams, data center operations, engineering, and external vendors. The TPM operates at the intersection of program leadership and hands-on technical deployment, ensuring that end-to-end GPU cluster buildouts — spanning compute, high-speed networking fabrics, high-performance storage, power, cooling, and automation — are delivered on time, within budget, and to operational readiness standards. This role also coordinates mechanical, electrical, and plumbing (MEP) readiness with facilities engineering and data center operations to ensure power capacity, cooling distribution, and physical infrastructure are in place ahead of rack-and-stack activities.
The ideal candidate brings strong technical project management expertise in hyperscale or cloud infrastructure environments, combined with a working understanding of AI/GPU deployment lifecycles and cross-functional program execution.
Responsibilities
- AI & GPU Cluster Deployment & Delivery: Lead end-to-end project management for large-scale AI/GPU cluster deployments, including multi-rack GPU compute platforms such as NVIDIA DGX and similar technologies. Manage deployment programs spanning compute, high-performance storage, networking, power, cooling, and automation. Coordinate procurement, rack-and-stack sequencing, cabling schedules, network deployment timelines, burn-in testing, cluster validation, and operational handoff. Partner with engineering teams and managers to establish consistent project tracking, milestone reporting, and status updates across teams and systems such as Jira and Confluence. Coordinate deployment activities and dependencies to ensure projects meet schedule, budget, quality, and operational readiness requirements. Contribute to the development of repeatable deployment methodologies and scalable operational standards.
- Networking & GPU Fabric Coordination: Coordinate deployment activities involving high-performance GPU networking fabrics, including InfiniBand and Ethernet-based architectures. Support project execution involving InfiniBand NDR/XDR, Spectrum-X Ethernet fabrics, spine-leaf topologies, RDMA, and RoCE. Work closely with network engineering teams to track topology implementation, fabric configuration, validation, and deployment dependencies. Coordinate project activities related to Subnet Manager, UFM, network validation, and operational handoff. Identify networking dependencies and blockers that could affect GPU cluster deployment schedules and drive cross-functional resolution.
- Storage & Data Infrastructure: Coordinate with storage engineering teams on the deployment and integration of high-performance storage environments supporting AI/GPU workloads. Support deployment programs involving VAST Data or similar high-performance storage platforms and distributed storage architectures. Track storage integration, validation, and readiness milestones as part of the overall GPU cluster deployment schedule. Coordinate storage dependencies with compute, networking, provisioning, and operational teams.
- Infrastructure Provisioning & Automation: Coordinate deployment activities involving automated infrastructure provisioning and lifecycle management. Support project execution involving Canonical MaaS or similar automated provisioning systems. Partner with technical teams using Infrastructure-as-Code, configuration management, and automation technologies. Track hardware integration, firmware management, provisioning, validation, and operational readiness activities. Support coordination of infrastructure automation and diagnostics involving technologies such as Ansible, Python, Shell, and SQL.
- Data Center & MEP Readiness: Coordinate with facilities engineering and data center operations to align MEP readiness with GPU infrastructure deployment schedules. Track power distribution, cooling capacity, floor layout, containment, and other physical infrastructure dependencies ahead of rack-and-stack activities. Coordinate readiness for high-density infrastructure, including UPS and PDU capacity, liquid and air cooling, CDU deployment, facility water loops, and heat rejection systems. Identify power, cooling, or physical infrastructure dependencies that could impact deployment timelines and drive timely resolution across teams.
- Program & Stakeholder Management: Manage complex, cross-functional technical deployment programs involving distributed engineering teams, data center operations, facilities teams, vendors, and external partners. Communicate project status, risks, milestones, dependencies, and required decisions to stakeholders at all levels, including executive leadership. Prepare and present regular program reviews, steering committee updates, and ad-hoc project analyses. Maintain centralized project documentation and ensure consistent and accurate project data. Track deployment KPIs, including schedule variance, budget adherence, and quality metrics. Drive continuous improvement through data-driven retrospectives and lessons learned. Contribute to improving and documenting project management best practices and repeatable deployment processes.
Qualifications
- Bachelor's degree in Engineering, Computer Science, Information Technology, or a related field, or equivalent professional experience
- 5+ years of technical project management experience in HPC/AI infrastructure environments
- Proven experience managing complex, cross-functional technical projects with distributed teams
- Strong understanding of data center and AI/GPU infrastructure
- Working knowledge of GPU server architectures and NVIDIA platforms, including DGX and HGX
- Familiarity with L10/L11/L12 infrastructure integration
- Working knowledge of high-performance networking technologies, including InfiniBand (NDR/XDR), Spectrum-X Ethernet fabrics, spine-leaf topologies, RDMA, and RoCE
- Understanding of high-performance storage platforms such as VAST Data and distributed storage architectures
- Familiarity with Canonical MaaS or similar automated provisioning systems
- Understanding of liquid and air cooling environments, including CDU deployment, facility water loops, and heat rejection systems
- Understanding of UPS, PDU, and power distribution within high-density data center environments
- Working knowledge of AI/GPU infrastructure deployment lifecycles, including hardware integration, fabric configuration (Subnet Manager, UFM), NCCL validation, firmware management, and operational handoff
- Experience with project management tools and methodologies, including scheduling and resource planning in fast-paced, remote environments
- Excellent communication, presentation, and stakeholder management skills
- Strong problem-solving and risk-management capabilities
Preferred Qualifications
- Experience with hyperscale cloud providers, large-scale data center deployments, or AI infrastructure programs
- Familiarity with Infrastructure-as-Code and configuration management tools, including Ansible
- Familiarity with Python, Shell, and SQL for infrastructure automation and diagnostics
- Experience with DCIM, BMS, and observability platforms supporting liquid-cooled environments
- Familiarity with NVIDIA/Mellanox networking platforms
- Knowledge of NVLink and NVSwitch technologies
- Experience coordinating with MEP engineering teams on electrical and cooling infrastructure scheduling for large-scale deployments
- Experience with procurement processes in infrastructure delivery environments
Key Competencies
- Technical project and program management
- AI/GPU infrastructure deployment
- Cross-functional project execution
- Data center infrastructure coordination
- Schedule and milestone management
- Risk identification and mitigation
- Vendor and stakeholder management
- Technical communication and executive reporting
- Process improvement and deployment standardization
- Strong problem-solving and organizational skills
Pay
$155,000–$175,000 USD
Schedule & Location
Remote – U.S. (Washington, D.C., Phoenix, AZ, Memphis, TN, Columbus, OH, or nearby regions). Full-Time. Travel requirements: Up to 25%.