Technical Program Manager - Provider Management
Andromeda · San Francisco, CA · Today
RemoteRemoteProject ManagementFull-time
About the role
This role owns our provider relationships on the execution side. You'll run the programs that make a provider work well inside Andromeda. This includes onboarding new providers and sites, capacity rollouts, hardware and fabric quality issues, and major incidents.
Responsibilities
- Manage new provider and site onboarding, capacity expansions, remediation of recurring hardware or network issues, and provider relationships.
- Act as incident commander during major provider-side incidents: assemble the right people from our SRE team, the provider, and the affected customer, run the response, own communication throughout, and drive the post-incident remediation program afterward.
- Keep a real plan for each program; milestones, owners, dependencies, risks.
- Maintain visibility to everyone involved, including the provider.
- Hold providers to their commitments: structured check-ins, tracked action items, and escalation to provider leadership when things slip, backed by what's in the contract.
- Catch capacity and quality risk early, before it becomes a customer-facing problem.
- Turn what you learn into playbooks and provider-facing standards, so onboarding the next site is faster and cleaner than the last.
Requirements
- Several years running technical programs in infrastructure.
- TPM or technical project management with genuine execution ownership, ideally involving external vendors, partners, or suppliers you didn't control.
- Incident management experience: you've commanded or run point on production incidents involving multiple organizations, and you're comfortable with the off-hours reality that comes with that.
- Enough technical depth to hold your own with a provider's data-center engineers and our SREs on GPUs, networking, and storage.
- Solid fundamentals: planning, risk tracking, status communication, and keeping four or five programs on the rails at once.
- Calm, direct communication. You'd rather deliver bad news early than good news late, and both providers and internal teams trust you because of it.
- Comfort with ambiguity. The process you'll follow mostly doesn't exist yet; you'll write a lot of it.
Qualifications
- Direct experience working with or inside neocloud, colocation, or data-center providers.
- Vendor or supplier management background — SLAs, commitments, escalation frameworks.
- Familiarity with the stack: NVIDIA data-center GPUs, InfiniBand/RoCE, Slurm or Kubernetes.
- Experience supporting AI research labs or other large-scale GPU customers on the consuming side.
Skills
- Strong communication skills.
- Ability to manage complex projects and relationships.
- Experience with incident management and problem-solving.
- Technical knowledge of data-center infrastructure, particularly NVIDIA GPUs, InfiniBand, and storage.
Benefits
- High-growth environment: Get in early at a company at the center of the AI infrastructure boom.
- Ownership: First TPM hire for the solutions engineering team, you’ll get to build this function from the ground up.
- Competitive compensation: + meaningful equity.
- Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO.
Pay
TBD
Schedule
Full-time