Compute Deployment Engineer
Fluidstack · San Francisco, CA · 4 days ago
On-siteInformation Technology$197k–$227k/yrFull-time
About the role
The Infrastructure Team at Fluidstack is responsible for bringing gigawatts of accelerators from first power-on to production, qualifying racks at scale, driving qualification through the base-management Kubernetes platform and provisioning stack, and triaging hardware failures. They also partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.
Responsibilities
- Own compute turn-up from facility availability to ready-for-service
- Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level
- Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services)
- Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates
- Deploy turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks
- Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity
Requirements
- Brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production
- Worked deep in Linux and out-of-band management: BMC, IPMI, and Redfish are daily tools for you, not occasional lookups
- Automated hardware workflows in Python or Go rather than clicking through them, and the second time you do anything by hand you turn it into software
- Worked physically in data halls, racking, cabling, and swapping components, and you're just as effective acting as remote hands or directing them
- Tried to isolate the fault to a component before reaching for a fix
- Travelled for turn-up windows when a new data hall came online
Qualifications
- Bonus: Kubernetes-based bare-metal provisioning
- Accelerator platform bringup (NVIDIA, AMD, or custom)
- Burn-in and stress harness design
- DCIM and inventory tooling
Skills
- Linux and out-of-band management expertise (BMC, IPMI, Redfish)
- Python or Go programming skills for automating hardware workflows
- Physical data center experience, including racking, cabling, and troubleshooting
- Experience with Kubernetes-based bare-metal provisioning and DCIM systems
Benefits
- Pay equity and transparency
- Equal Employment Opportunity Employer
- Committed to accommodating qualified applicants with arrest and conviction records
Pay
$197K - $227K
Schedule
Not specified