Jobs · Management · New York

Production Engineer, Compute Team Lead

Fluidstack · New York, NY · 4 wk ago
On-siteManagement$225k–$351k/yrFull-time

About the company

We exist to make humanity more free. Technology has historically given people more time for the things they want to do, instead of things they have to do. Powerful AI will be the biggest lever for human choice we've ever built—but only if models are aligned with what humanity actually wants. We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them—with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.

We hire people who care deeply about this problem space and operate with extreme ownership, velocity, and first principles thinking. The frontier of AI is the most interesting problem of our time, and we push it forward with high intensity.

About the team

The Data Center Operations Team is tackling key problems at an unprecedented scale:

  • Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.
  • Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
  • Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.

Responsibilities

  • Lead the compute production engineering team keeping tens of thousands of GPUs serving customers.
  • Own fleet availability for compute: define the SLOs, build the tooling, and move the number.
  • Build automation for the node lifecycle, including provisioning, health checks, remediation, and return to service, without human touch.
  • Set the on-call and escalation model that keeps response sharp without burning the team out.

Requirements

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.

  • You've led SRE or production engineering teams running large fleets.
  • You've moved an availability number and can explain exactly how.
  • You've shipped automated remediation that retired a runbook.
  • You hire and grow strong engineers, and they say so.
  • Bonus: GPU or HPC fleets. Kubernetes or Slurm. Hardware failure analytics. Customer-facing reliability.

Pay

Compensation Range: $225K - $351K

Similar jobs