Staff Cloud Infrastructure Engineer (up to $250k)
About the role
This company is building a lightweight runtime that turns AI agents, workflows, and backend services into durable processes. Engineers can focus on business logic, not failure mechanics. The runtime ships as a single Rust binary with a custom storage layer and low-latency orchestration, already running critical financial workflows for Fortune 500 enterprises and Tier 1 banks. Founded by creators of Apache Flink and leaders from Meta's messaging infrastructure, it's backed by leading VCs and angels.
You'll take staff-level ownership across the entire cloud infrastructure surface. This means setting the design direction for a managed multi-tenant SaaS offering, as well as for the infrastructure running inside customer cloud accounts for BYOC and on-prem installations. You'll define fleet-wide reliability and observability, building the automation to scale across deployment methods. This isn't just operating existing infrastructure; it's architecting and building the foundational systems for a product that spans multiple deployment models, with significant design authority in a small, high-impact team.
Responsibilities
- Design and implement infrastructure for a multi-tenant SaaS offering, covering control plane, networking, storage, and observability.
- Architect and build infrastructure for customer-managed deployments (BYOC, on-prem), ensuring consistency and scalability.
- Define and implement fleet-wide reliability and observability standards, including SLOs, metrics, tracing, logging, and runbooks.
- Develop automation to scale infrastructure across diverse deployment methods and cloud providers.
- Participate in the cloud on-call rotation, providing critical timezone coverage for a global user base.
Requirements
- Operated production SaaS or platform infrastructure at staff level, with strong opinions on multi-tenant systems formed from real failure modes.
- Deep working knowledge of at least one major cloud provider's architecture; experience with multiple is a plus.
- Designed and run large-scale Kubernetes-based stateful workloads, balancing continuous delivery with safety.
- Fluency in a systems language (Rust, Go, or C++); willingness to learn Rust on the job.
- Hands-on at staff level: comfortable taking a problem from architectural design through to production operations.
- US-based, able to join a cloud on-call rotation for timezone coverage.