Production Engineering Lead, Compute
How We Operate
Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.
Velocity. We drive everything forward as fast as possible.
First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.
Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.
Role Scope
Lead the compute production engineering team keeping tens of thousands of GPUs serving customers.
Own fleet availability for compute: define the SLOs, build the tooling, and move the number.
Build automation for the node lifecycle, provisioning, health checks, remediation, return to service, without human touch.
Set the on-call and escalation model that keeps response sharp without burning the team out.
What We're Looking For
- You've led SRE or production engineering teams running large fleets.
- You've moved an availability number and can explain exactly how.
- You've shipped automated remediation that retired a runbook.
- You hire and grow strong engineers, and they say so.
- Bonus: GPU or HPC fleets. Kubernetes or Slurm. Hardware failure analytics. Customer-facing reliability.
Compensation Range
$260K - $361K