Infrastructure Engineer, DevOps
Zof AI · San Francisco, CA · 2 wk ago
Information TechnologyFull-time
San Francisco, CA · Full-time · Mid to Senior · On-site
About This Role
Zof AI is seeking a DevOps Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply. This role owns the execution layer of our control plane: Kubernetes and container orchestration, CI/CD pipelines, hard isolation for untrusted code, observability, and the cost controls that keep large agent fleets affordable. If you have worked as a Site Reliability Engineer, Platform Engineer, Cloud Engineer, or Infrastructure Engineer, this is that discipline at Zof AI. The ideal candidate has operated production infrastructure at scale and treats security, reliability, and cost per agent run as constraints they personally own.
Responsibilities
- Design and operate the sandboxed environments where agents reproduce defects and validate fixes.
- Own Kubernetes, container, and compute infrastructure end to end.
- Build CI/CD pipelines that let engineers ship safely many times a day.
- Harden isolation boundaries so untrusted customer code stays inside its sandbox.
- Instrument fleet health with metrics, logs, tracing, and alerting that catch failures early.
- Drive down cost per agent run through scheduling, autoscaling, and capacity work.
- Automate provisioning, deployment, rollback, and environment management.
- Partner with engineers to make infrastructure fast and safe to build on.
Requirements
- Experience running production infrastructure on a major cloud platform.
- Working knowledge of Kubernetes, containers, and orchestration.
- Experience building CI/CD pipelines, deployment automation, and infrastructure as code.
- Familiarity with observability, monitoring, and on-call practice.
- Judgment about security, reliability, and cost trade-offs.
- Daily use of AI tools to automate operational and engineering work.
- Clear written and verbal communication.
- Comfort operating in a fast-moving environment.
Nice to have
- Experience with sandboxing or multi-tenant isolation tooling such as gVisor or Firecracker.
- Experience with Terraform, Pulumi, or similar infrastructure as code tooling.
- Experience running large batch or job-based workloads cost efficiently.
- Experience in early-stage infrastructure or platform teams.
What we provide
- MacBook Pro
- Premium AI development tools: Cursor Ultra, Claude Code Ultra, OpenAI Codex Max or equivalent advanced AI tooling
- Access to a high-performance AI product environment
- Close collaboration with leadership, engineering, and customers
- Opportunity to work in the San Francisco AI ecosystem
- Wellness and productivity support where applicable
- Competitive startup environment with high ownership and direct product impact