Jobs · Engineering · California

Member of Technical Staff, Cluster Administration

Inferact · San Francisco, CA · 1 wk ago
On-siteEngineering$200k–$400k/yrFull-time

About the Role

We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock.

Responsibilities

  • Take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across the systems our engineers rely on every day.
  • Work closely with engineering leadership and infrastructure owners to standardize how we provision, operate, debug, and scale compute across providers.
  • Own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.
  • Automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.

Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.
  • Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.
  • Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.
  • Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.
  • Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.

Preferred Qualifications

  • Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments.
  • Experience improving cluster utilization, reducing idle or unavailable GPU capacity, and debugging scheduling or resource contention issues.
  • Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems.
  • Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.
  • Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene.
  • Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations for engineering-heavy teams.

Bonus Points

  • Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team.
  • Built monitoring, alerting, runbooks, health checks, or remediation workflows that materially reduced operational toil or incident resolution time.
  • Operated Kubernetes clusters for ML or GPU workloads at meaningful scale.
  • Standardized provisioning, diagnostics, monitoring, and operating patterns across multiple compute providers.
  • Carried real operational responsibility for infrastructure used by many engineers or researchers.

Logistics

Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Visa sponsorship: We sponsor visas on a case-by-case basis.

Benefits

  • Generous health, dental, and vision benefits.
  • 401(k) company match.

Pay

The expected annual salary range for this position is $200,000 - $400,000 USD + equity.

Similar jobs