Jobs · Engineering · Texas

Technical Program Manager- AI Cluster Validation

AMD · Austin, TX · 2 wk ago
On-siteEngineeringFull-time

About the Role

We are seeking a Technical Program Manager to lead execution of AI cluster engineering programs with deep focus on GPU platforms, rack-level solutions, and AI Cluster validation. This role is responsible for driving end-to-end delivery from GPU + server integration through rack bring-up, scale testing, failure analysis, and system debug closure, ensuring platform readiness for hyperscale and enterprise AI deployments. This role operates at the intersection of hardware, firmware, networking, and scale-test execution, and requires strong technical depth combined with disciplined program execution.

About the Person

You are a hands-on TPM who thrives in complex, fast-moving ecosystems, and can connect deep technical details to crisp program plans, executive reporting, and customer outcomes. You are comfortable driving execution in bring-up and EVT/DVT/PVT working closely with engineers to root-cause issues, unblock debug, and make data-driven tradeoffs to keep programs moving. You bring urgency, ownership, and clarity to ambiguous problem spaces and can communicate effectively from lab floor to executive review.

Key Responsibilities

  • Program Leadership & Execution
    • Define, plan, and drive program plans for AI infrastructure systems validation and readiness, including server integration, rack bring-up, and cluster-scale deployment readiness.
    • Create and maintain core PM artifacts: schedules, dependency maps, resource forecasts, risk/issue logs, and program dashboards/status reports.
    • Identify and drive mitigation plans for issues/risks, including cross-team escalations and corrective actions across multiple engineering areas.
    • Drive regular execution reviews with engineering teams and provide concise, data-driven updates to senior leadership.
  • GPU & Platform Execution
    • Own program execution for GPU-based AI platforms, spanning system bring-up, qualification, scale readiness, and deployment validation across server, rack, and cluster levels.
    • Drive alignment across GPU, CPU, firmware, BIOS/BMC, and system teams to ensure readiness for scale testing and customer workloads.
    • Track platform issues, and debug dependencies; ensure risks are clearly documented, owned, and mitigated.
  • AI Rack / Cluster Validation
    • Own program planning and execution for multi-node and multi-rack scale testing, including test strategy, scheduling, coverage tracking, and readiness gates.
    • Lead end-to-end delivery of rack-level AI solutions, including compute trays, switch trays, cabling, power, cooling, and management infrastructure.
    • Ensure rack bring-up plans are executable, resourced, and gated with clear entry/exit criteria across EVT, DVT, and scale phases.
    • Drive coordination across lab operations, infrastructure, and engineering teams to unblock rack access, power, networking, and test readiness.
    • Partner with scale, performance, and automation teams to ensure workloads, stress tests, and regressions plans are ready before hardware arrives.
  • Debug, Failure Analysis & Risk Management
    • Act as the execution lead for platform debug, coordinating across engineering teams to ensure fast triage, root-cause analysis, and resolution of system-level issues.
    • Track high-impact failures (GPU, HSIO, FW, rack, network) through debug forums ensuring clear ownership and closure plans.
    • Balance debug depth vs. program timelines, escalating tradeoffs when needed and ensuring leadership has a clear view of risk and impact.

Required Qualifications

  • Experience leading complex hardware or AI infrastructure programs with ownership across bring-up, validation, and deployment phases.
  • Strong technical understanding of GPU-based AI systems, rack architectures, and datacenter infrastructure.
  • Proven ability to manage ambiguity, drive debug execution, and lead cross-functional teams without direct authority.
  • Strong written and verbal communication skills, including executive-level status reporting.
  • Proficiency with program management and execution tools (Jira, Confluence, dashboards, Excel/PowerPoint).

Preferred Qualifications

  • Hands-on experience with GPU cluster scale testing, system stress, or performance validation.
  • Familiarity with rack-level bring-up, power/cooling constraints, networking, and failure modes at scale.
  • Experience working through hardware/firmware debug cycles in pre-production or customer-facing environments.

Academic Credentials

Bachelor's or master's degree in systems, EE, CS, or related engineering discipline.

Location: Austin, TX

This role is not eligible for visa sponsorship.

Similar jobs