Jobs · Engineering · New York

Customer Reliability Engineer

Fluidstack · New York, NY · 1 mo ago
On-siteEngineering$208k–$269k/yrFull-time

How We Operate

Extreme ownership. Full autonomy. Own things end to end often taking on scope outside your core role without being asked to get things done.

Velocity. We drive everything forward as fast as possible.

First principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.

Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.

Role Scope

Own reliability for named customer workloads: their clusters, their SLAs, their escalations.

Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.

Run customer-facing incident communication with technical depth and no spin.

Turn recurring customer pain into engineering fixes with the production teams.

What We're Looking For

  • You've supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • You debug distributed systems methodically across layers you don't own.
  • You've written incident updates customers trusted more after reading.
  • You push internal teams to fix causes, not symptoms, and follow up until they do.
  • Bonus: GPU training workloads. InfiniBand or RoCE. Slurm or Kubernetes. NCCL debugging.

Compensation Range

$208K - $269K

Similar jobs