Jobs · Information Technology

Staff Cloud Infrastructure Engineer (up to $250k)

Dex · United States · Yesterday
RemoteRemoteInformation TechnologyFull-time

About the role

This company is building a lightweight runtime that turns AI agents, workflows, and backend services into durable processes. Engineers can focus on business logic, not failure mechanics. The runtime ships as a single Rust binary with a custom storage layer and low-latency orchestration, already running critical financial workflows for Fortune 500 enterprises and Tier 1 banks. Founded by creators of Apache Flink and leaders from Meta's messaging infrastructure, it's backed by leading VCs and angels.

You'll take staff-level ownership across the entire cloud infrastructure surface. This means setting the design direction for a managed multi-tenant SaaS offering, as well as for the infrastructure running inside customer cloud accounts for BYOC and on-prem installations. You'll define fleet-wide reliability and observability, building the automation to scale across deployment methods. This isn't just operating existing infrastructure; it's architecting and building the foundational systems for a product that spans multiple deployment models, with significant design authority in a small, high-impact team.

Responsibilities

  • Design and implement infrastructure for a multi-tenant SaaS offering, covering control plane, networking, storage, and observability.
  • Architect and build infrastructure for customer-managed deployments (BYOC, on-prem), ensuring consistency and scalability.
  • Define and implement fleet-wide reliability and observability standards, including SLOs, metrics, tracing, logging, and runbooks.
  • Develop automation to scale infrastructure across diverse deployment methods and cloud providers.
  • Participate in the cloud on-call rotation, providing critical timezone coverage for a global user base.

Requirements

  • Operated production SaaS or platform infrastructure at staff level, with strong opinions on multi-tenant systems formed from real failure modes.
  • Deep working knowledge of at least one major cloud provider's architecture; experience with multiple is a plus.
  • Designed and run large-scale Kubernetes-based stateful workloads, balancing continuous delivery with safety.
  • Fluency in a systems language (Rust, Go, or C++); willingness to learn Rust on the job.
  • Hands-on at staff level: comfortable taking a problem from architectural design through to production operations.
  • US-based, able to join a cloud on-call rotation for timezone coverage.

Similar jobs