Jobs · Project Management

Technical Program Manager - Provider Management

Andromeda · San Francisco, CA · Today
RemoteRemoteProject ManagementFull-time

About the role

This role owns our provider relationships on the execution side. You'll run the programs that make a provider work well inside Andromeda. This includes onboarding new providers and sites, capacity rollouts, hardware and fabric quality issues, and major incidents.

Responsibilities

  • Manage new provider and site onboarding, capacity expansions, remediation of recurring hardware or network issues, and provider relationships.
  • Act as incident commander during major provider-side incidents: assemble the right people from our SRE team, the provider, and the affected customer, run the response, own communication throughout, and drive the post-incident remediation program afterward.
  • Keep a real plan for each program; milestones, owners, dependencies, risks.
  • Maintain visibility to everyone involved, including the provider.
  • Hold providers to their commitments: structured check-ins, tracked action items, and escalation to provider leadership when things slip, backed by what's in the contract.
  • Catch capacity and quality risk early, before it becomes a customer-facing problem.
  • Turn what you learn into playbooks and provider-facing standards, so onboarding the next site is faster and cleaner than the last.

Requirements

  • Several years running technical programs in infrastructure.
  • TPM or technical project management with genuine execution ownership, ideally involving external vendors, partners, or suppliers you didn't control.
  • Incident management experience: you've commanded or run point on production incidents involving multiple organizations, and you're comfortable with the off-hours reality that comes with that.
  • Enough technical depth to hold your own with a provider's data-center engineers and our SREs on GPUs, networking, and storage.
  • Solid fundamentals: planning, risk tracking, status communication, and keeping four or five programs on the rails at once.
  • Calm, direct communication. You'd rather deliver bad news early than good news late, and both providers and internal teams trust you because of it.
  • Comfort with ambiguity. The process you'll follow mostly doesn't exist yet; you'll write a lot of it.

Qualifications

  • Direct experience working with or inside neocloud, colocation, or data-center providers.
  • Vendor or supplier management background — SLAs, commitments, escalation frameworks.
  • Familiarity with the stack: NVIDIA data-center GPUs, InfiniBand/RoCE, Slurm or Kubernetes.
  • Experience supporting AI research labs or other large-scale GPU customers on the consuming side.

Skills

  • Strong communication skills.
  • Ability to manage complex projects and relationships.
  • Experience with incident management and problem-solving.
  • Technical knowledge of data-center infrastructure, particularly NVIDIA GPUs, InfiniBand, and storage.

Benefits

  • High-growth environment: Get in early at a company at the center of the AI infrastructure boom.
  • Ownership: First TPM hire for the solutions engineering team, you’ll get to build this function from the ground up.
  • Competitive compensation: + meaningful equity.
  • Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO.

Pay

TBD

Schedule

Full-time

Similar jobs