Jobs · Engineering

Staff Platform Engineer

Ecommerce Guide · United States · 2 days ago
RemoteRemoteEngineering$190k–$230k/yrFull-time

About the Role

We are seeking an experienced Staff Platform Engineer to help build and operate Acorn, Postscript's next-generation CDP and messaging platform: ~52 Rust/Python services on EKS, built to run at 500K+ events/sec for 20K+ merchants. What makes this role different is that the platform is agent-operated. Deploys, migrations, drift reconciliation, and infra standup are automated and encoded as self-verifying runbooks that AI agents execute. We are not hiring someone to run those procedures by hand. We are hiring the person who designs the system, writes the guardrails that let agents execute it safely, and owns the failures no runbook covers yet.

Responsibilities

  • System & Platform Design: Own infra topology, scaling and blast-radius boundaries, the build/deploy graph, database ownership boundaries, and cost ceilings — the decisions agents execute but can't make.
  • Guardrail Engineering: Build the validation gates, idempotency checks, drift detectors, and runbooks that keep both humans and agents from doing damage. Your deliverable is the safety system, not the deploy.
  • Escalation Tier: Diagnose the novel failures — consumer lag, DLQs, materialized-view chains, consent ordering. Write the runbook the first time; agents handle it after.
  • Reliability, Capacity & Cost: Own the 500K events/sec target: load testing, Karpenter/HPA tuning, Database sizing, SLOs and alerting.
  • Security & Secrets Boundary: Own IAM/IRSA, secret rotation, the auth chain, and what the agent/MCP surface is allowed to do.
  • Agent-Fleet Operability: Keep skills, runbooks, and agent context in sync with reality. A stale runbook is a wrong actor, executed confidently.
  • Incident Leadership: Lead incidents and capture every fix as a new runbook plus regression test.

Requirements

  • 5+ years operating production Kubernetes and cloud infrastructure (AWS) at scale.
  • Strong Terraform and GitOps/Kustomize practice, with a bias for making operations idempotent, automated, and self-verifying.
  • Real incident-response experience on distributed-systems failures — databases, streaming, consumer lag, data integrity.
  • Fluency with data systems like Postgres, columnar stores, and log/streaming buses, understanding their behavior under load.
  • Comfort working alongside AI agents — writing guardrails, runbooks, and checks for safe automation, and knowing when human intervention is necessary.
  • Instincts to prioritize automation, checks, and elimination of manual steps over documentation.

Qualifications

Experience with agent-based automation, distributed systems, and cloud infrastructure at scale. Strong problem-solving skills and a proactive approach to system design and reliability.

Skills

  • Kubernetes, AWS, Terraform, GitOps, Kustomize
  • Incident response, distributed systems, data integrity
  • Postgres, columnar stores, streaming buses
  • AI agent integration, automation, guardrail engineering

Pay

Salary range of USD $190,000 to $230,000 base plus significant equity (we do not have geo-based salaries).

Benefits

  • High growth startup with opportunities for direct impact and career growth.
  • Work from home (or wherever).
  • Competitive compensation and equity.
  • Flexible paid time off.

Similar jobs

Staff Platform Engineer

CheckfrontNAMER· 2 mo ago
RemoteEngineering$145k–$195k/yrapply on checkfront.applytojob.com

Staff Platform Engineer

CU DirectIrvine, CA· 1 mo ago
Information Technology$148k–$185k/yrapply on origence.clearcompany.com

Staff Platform Engineer

CorientAustin, Texas Metropolitan Area· 1 mo ago
Engineering$220/hrapply on ci.wd3.myworkdayjobs.com