Member of Technical Staff, Infrastructure
About Sycamore
Sycamore is building the trusted agent operating system for the enterprise. Our platform helps companies build, deploy, and orchestrate AI agents that take on real operational work, with the security and control large organizations need. We are a small, engineering-led team working directly with Fortune 500 enterprises. We have raised $65M from Coatue and Lightspeed, along with other investors and industry leaders.
Where you could focus
This is software engineering for consequential distributed systems, not cloud administration: control-plane services and Kubernetes controllers, identity and network boundaries, tenant-isolated data, and failures that cross application, cluster, database, and customer-network layers. Infrastructure is the foundation beneath Product and Core AI, so an error here can affect every application and customer. The team covers three areas. They share an on-call rotation and a great many failures, so nobody works in only one.
- The platform and control plane. Kubernetes-based serving, the sandboxes agents execute in, deployment and rollback, networking and identity, and the operators that reconcile all of it. This is where blast radius is decided.
- The data platform. This is database platform engineering, not analytics: there is no warehouse, no dbt, and no pipeline waiting for an owner. What we have is a growing fleet of PostgreSQL databases carrying enterprise customer data under a genuine isolation requirement, hundreds of migrations across many trees and authors, a connection budget that is a real ceiling, retention obligations measured in years, and a hybrid full-text and vector retrieval path in the hot path of every agent conversation. The problems are fleet-shaped: a migration is easy, and applying it safely across every tenant database with per-tenant failure isolation and no downtime is not.
- Reliability. How the platform behaves over time: service level objectives and error budgets that mean something, an observability estate managed as code, promotion gates that decide whether a release reaches production, capacity and cold-start performance, and incident response from page through postmortem to the structural fix. We have a large monitor fleet and an enforced metric taxonomy, but few service level objectives, no error budgets, no burn-rate alerting, and a promotion gate that warns rather than blocks. Building that practice is the work, not a side project within it.
What you will do
- Build and operate Kubernetes-based serving and sandbox platforms for agents, internal services, and customer applications.
- Develop infrastructure-as-code, Helm charts, provisioning services, admission controls, and operators that reconcile environments safely, along with the change controls, drift detection, release gates, and fail-safe rollout procedures around them.
- Own the database fleet: schema topology and migration safety, provisioning and credential rotation, pooling and connection economics, isolation guarantees, retention and partitioning, and the search and embedding path.
- Define service level indicators and objectives, establish error budgets, and own the observability estate as code: monitors, dashboards, tracing, metric taxonomy, burn-rate alerting, and the policy that keeps it from decaying.
- Lead incident response from detection through containment, root cause, postmortem, and the structural fix.
- Support hosted and customer-controlled deployment patterns while keeping the operational model consistent and auditable.
- Travel to customer sites when working alongside a customer's engineering or security team will materially improve a deployment or incident outcome.
The environment you will work in
Our current infrastructure environment includes GCP, Kubernetes and GKE, Knative, Envoy, Cloudflare, Pulumi, Helm, BuildKit, PostgreSQL and Cloud SQL, object storage, container registries, workload identity, Python and Go control-plane services, Kubernetes operators, GitHub Actions, and production observability. Alert triage is substantially agent-driven, with automated investigation, ticket filing, and a decay process that nominates monitors nobody has acted on for deletion. We also support customer-controlled deployment requirements and maintain more than one application-serving path during platform migrations.
This is context, not a checklist. We do not require previous experience with every cloud product or tool. We care about deep systems fundamentals, the ability to learn unfamiliar infrastructure, and evidence that you have personally operated and recovered important multi-tenant systems.
What we are looking for
- 5-12 years of software or infrastructure engineering experience. We will make exceptions for exceptional people in either direction.
- Experience building production control-plane software, operators, platform services, or developer infrastructure, not only configuring vendor products.
- Experience with infrastructure-as-code and delivery systems, including state management, drift detection, reproducible builds, staged rollout, and rollback.
- Evidence of owning consequential production systems through incidents, capacity changes, recovery exercises, and reliability improvements.
- A security-minded approach to tenant isolation, credentials, access control, software supply chains, and customer data.
- The ability to write maintainable production software and automated tests for infrastructure behavior.
- AI-native. You use coding agents and modern models as a force multiplier while still understanding and owning every critical operational decision.
- Clear communication and high EQ. You can work with a customer's CTO, security team, and infrastructure engineers without losing technical depth.
- Comfort with startup ambiguity, broad ownership, and occasional customer travel.
We do not expect all of the following from one person, and any of it is a strong signal for a particular area: deep PostgreSQL expertise, including connection management, query planning, online schema change, and multi-tenant isolation as an enforced guarantee rather than an application convention; Kubernetes performance and capacity depth; or incident command experience and postmortems that produce structural change rather than an action item nobody does. We care more about what you personally built, operated, and recovered than a particular school or vendor certification.
Interview process
- A 30-minute introductory conversation.
- Two 60-minute technical interviews, one focused on systems design and one on coding.
- A take-home assignment where you build and present a real solution using the tools you would use on the job.
Why join
- Build the secure execution foundation for agents doing real work in large enterprises.
- Solve systems problems spanning Kubernetes, control planes, identity, networking, data, delivery, and recovery.
- Turn difficult customer deployment constraints into a platform that scales across environments.
- Work on infrastructure where software design and operational judgment matter equally.
- Join early enough to define how the engineering team operates and grows.
- Receive competitive cash compensation and meaningful equity in the company you are helping build.