Member of Technical Staff, Cluster Administration
Inferact · San Francisco, CA · 1 wk ago
On-siteEngineering$200k–$400k/yrFull-time
About the Role
We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock.
Responsibilities
- Take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across the systems our engineers rely on every day.
- Work closely with engineering leadership and infrastructure owners to standardize how we provision, operate, debug, and scale compute across providers.
- Own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.
- Automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.
Qualifications
- Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.
- Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.
- Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.
- Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.
- Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.
Preferred Qualifications
- Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments.
- Experience improving cluster utilization, reducing idle or unavailable GPU capacity, and debugging scheduling or resource contention issues.
- Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems.
- Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.
- Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene.
- Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations for engineering-heavy teams.
Bonus Points
- Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team.
- Built monitoring, alerting, runbooks, health checks, or remediation workflows that materially reduced operational toil or incident resolution time.
- Operated Kubernetes clusters for ML or GPU workloads at meaningful scale.
- Standardized provisioning, diagnostics, monitoring, and operating patterns across multiple compute providers.
- Carried real operational responsibility for infrastructure used by many engineers or researchers.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits
- Generous health, dental, and vision benefits.
- 401(k) company match.
Pay
The expected annual salary range for this position is $200,000 - $400,000 USD + equity.