Customer Reliability Engineer
Fluidstack · Seattle, WA · 2 days ago
On-siteInformation Technology$208k–$269k/yrFull-time
About the role
The Data Center Operations Team at Fluidstack is responsible for operating at the scale of a nation, managing power usage equivalent to some countries, and ensuring named customer workloads meet their Service Level Agreements (SLAs).
Responsibilities
- Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
- Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
- Run customer-facing incident communication with technical depth and no spin.
- Turn recurring customer pain into engineering fixes with the production teams.
Requirements
- You have supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
- You debug distributed systems methodically across layers you don't own.
- You have written incident updates customers trusted more after reading.
- You push internal teams to fix causes, not symptoms, and follow up until they do.
- Bonus: Experience with GPU training workloads, InfiniBand or RoCE, Slurm or Kubernetes, and NCCL debugging.
Qualifications
- Commitment to pay equity and transparency.
- Equal opportunity employer.
- Prioritizes diversity and inclusion.
Skills
- Experience with large-scale compute environments.
- Strong debugging skills across distributed systems.
- Effective communication and incident management.
- Problem-solving and root cause analysis abilities.
Benefits
- Competitive compensation range: $208K - $269K.
- Comprehensive benefits package.
- Flexible work schedule.
- Professional development opportunities.
Pay
Compensation Range: $208K - $269K.
Schedule
Flexible work schedule to accommodate global operations.