Software Engineer, Cloud Infrastructure
Scope of Work
The Production Engineering Team works on a variety of exciting problems, including:
- Make tens of thousands of GPUs legible in real time: build the observability platform that turns raw telemetry into signal, from site-level health down to individual device and link.
- Build the control plane every team at Fluidstack depends on: replace one-off tooling with a stable, versioned API surface that covers unified machine management, actual state inspection, and distributed command execution.
- Make the system's view of itself always match reality: integrate fleet state as a machine-readable source of truth across provisioning, operations, and customer-facing platforms, so every new site and GPU generation lands cleanly from day zero.
What We're Looking For
- You treat toil as a bug. If something requires a human to do it twice, you build the thing that makes it not require a human.
- You design APIs that age well. You've felt the pain of a leaky abstraction at scale and you don't repeat it.
- You move toward ambiguity, not away from it. You walk into the fog, build the map, and explain it to everyone else.
- You learn at a steep slope. You reach real competence in an unfamiliar domain fast. We value this over existing expertise.
- You carry a pager without flinching. You run the incident, write the postmortem, fix the systemic cause, and move on.
- You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks, and you drive Claude Code, Cursor, or similar every day.
- You've shipped production services that other teams depend on at scale, and you're comfortable in any language using AI coding tools.
Compensation and Benefits
- Competitive total compensation package (salary + equity).
- Retirement or pension plan, in line with local norms.
- Health, dental, and vision insurance.
- Generous PTO policy, in line with local norms.
About Fluidstack
We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it. We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI. We hire people who care deeply about this problem space. If that is you, please apply!