Jobs · Engineering · California

Staff Software Engineer, GPU Infrastructure Lifecycle Management

Together AI · San Francisco, CA · 5 days ago
Engineering$240k–$280k/yrFull-time

Responsibilities

  • Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA — as explicit, versioned states and transitions.
  • Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call — no ticket, no human in the loop.
  • Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically.
  • Owning reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service.
  • Partner with the inference/ML platform team: understand the cluster shapes they need — topology, interconnect, scheduling constraints — and encode them as first-class abstractions in the platform.
  • Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code — this is a product, not a collection of Ansible playbooks.

Requirements

  • Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution.
  • Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines).
  • Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts.
  • A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.

Nice to Have

  • Exposure to bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure.
  • Experience with GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE).
  • Prior work at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
  • Systems programming in Rust or Go.

Qualifications

Core qualifications (all levels):

  • Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution.
  • Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines).
  • Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts.
  • A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.

Skills

Core skills:

  • Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution.
  • Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines).
  • Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts.
  • A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.

Benefits

Competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $240,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Pay

The US base salary range for this full-time position is: $240,000 - $280,000 + equity + benefits.

Schedule

Full-time position.

Similar jobs