Software Engineer, Cloud Infrastructure
About the Role
We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants.
The Production Engineering Team is working on exciting problems such as making tens of thousands of GPUs legible in real time, building the control plane every team at Fluidstack depends on, and making the system's view of itself always match reality.
Responsibilities
Own the observability platform, build and operate the data pipelines, decoration and correlation engine, and healthcheck framework that make the fleet legible — from site down to device and link.
Define and build the API surface for infrastructure, design the contracts between production infrastructure and every tool that touches it.
Build the production control plane, unified machine management, actual state inspection, distributed command execution — and the Kubernetes-based infrastructure that underpins it all.
Own fleet state as source of truth, SLOs, site lifecycle state, and integration with internal infrastructure management and customer-facing operations platforms.
Land new hardware into the platform cleanly, ZTP, DHCP, DNS, artifacts — every new XPU generation and site integration goes through IaaS before production.
Requirements
- Treat toil as a bug, if something requires a human to do it twice, you build the thing that makes it not require a human.
- Design APIs that age well, you've felt the pain of a leaky abstraction at scale and you don't repeat it.
- Move toward ambiguity, not away from it, you walk into the fog, build the map, and explain it to everyone else.
- Learn at a steep slope, you reach real competence in an unfamiliar domain fast.
- Carry a pager without flinching, you run the incident, write the postmortem, fix the systemic cause, and move on.
- Be fluent with AI tooling, LLM APIs, MCP servers, and agentic frameworks, and you drive Claude Code, Cursor, or similar every day.
- Have shipped production services that other teams depend on at scale, and you're comfortable in any language using AI coding tools.
Skills
- Distributed systems and data pipeline engineering.
- Time-series observability stacks (Prometheus, Thanos, VictoriaMetrics).
- API design and versioning at scale.
- Workflow and orchestration engines (Temporal, Cadence).
- BMC/Redfish or hardware telemetry.
- Go, Python, and Postgres.
Benefits
- Competitive total compensation package (salary + equity).
- Retirement or pension plan, in line with local norms.
- Health, dental, and vision insurance.
- Generous PTO policy, in line with local norms.
Pay
Compensation Range: $208K - $269K