Jobs · Information Technology · Washington

Customer Reliability Engineer

Fluidstack · Seattle, WA · 2 days ago
On-siteInformation Technology$208k–$269k/yrFull-time

About the role

The Data Center Operations Team at Fluidstack is responsible for operating at the scale of a nation, managing power usage equivalent to some countries, and ensuring named customer workloads meet their Service Level Agreements (SLAs).

Responsibilities

  • Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
  • Run customer-facing incident communication with technical depth and no spin.
  • Turn recurring customer pain into engineering fixes with the production teams.

Requirements

  • You have supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • You debug distributed systems methodically across layers you don't own.
  • You have written incident updates customers trusted more after reading.
  • You push internal teams to fix causes, not symptoms, and follow up until they do.
  • Bonus: Experience with GPU training workloads, InfiniBand or RoCE, Slurm or Kubernetes, and NCCL debugging.

Qualifications

  • Commitment to pay equity and transparency.
  • Equal opportunity employer.
  • Prioritizes diversity and inclusion.

Skills

  • Experience with large-scale compute environments.
  • Strong debugging skills across distributed systems.
  • Effective communication and incident management.
  • Problem-solving and root cause analysis abilities.

Benefits

  • Competitive compensation range: $208K - $269K.
  • Comprehensive benefits package.
  • Flexible work schedule.
  • Professional development opportunities.

Pay

Compensation Range: $208K - $269K.

Schedule

Flexible work schedule to accommodate global operations.

Similar jobs

Customer Engineer

Column N.A.San Francisco, CA· 3 days ago
OTHR$125k–$160k/yrapply on jobs.ashbyhq.com