Jobs · Engineering · California

Member of Technical Staff, Core AI

Sycamore · Palo Alto, CA · 1 wk ago
Engineering$65/hrFull-time

About Sycamore

Sycamore is building the trusted agent operating system for the enterprise. Our platform helps companies build, deploy, and orchestrate AI agents that take on real operational work, with the security and control large organizations need. We are a small, engineering-led team working directly with Fortune 500 enterprises. We have raised $65M from Coatue and Lightspeed, along with other investors and industry leaders.

Where you could focus

Core AI owns the horizontal runtime, intelligence, and improvement capabilities that Product and Infrastructure both depend on, along with the Sycamore Forge experiences that make them usable. Engineers here own end-to-end slices, from cloud service and API design through the React interface. The team covers four related areas, and you do not have to pick one to apply. Tell us what you have built and we will work out the fit together.

  • The agent runtime. Multi-turn sessions, model routing, tool execution, memory, durable workflows, and the APIs that expose them. The hard part is correctness across long horizons: surviving retries, provider interruptions, partial failure, and context growth without losing the thread.
  • Harnesses and environments. An agent that writes and runs code needs somewhere to do it. That environment has to start fast, isolate genuinely untrusted execution, reach only the services it legitimately needs, and carry credentials it can use but never read. The same area owns the verification layer that decides whether what an agent produced actually works rather than only appears to.
  • Evaluation. Agent quality is genuinely hard to measure. A change that looks better on a handful of examples often is not, and a judge model can be confidently wrong in the same direction as the system it grades. This is where the offline suites, replay corpora, and statistical discipline that gate a release come from.
  • Self-improving systems. Every agent run produces evidence, and almost all of it is currently thrown away. This area turns it into improvement: structured trajectories at fleet scale, failure clusters surfaced from production rather than guessed at, and proposed changes that are versioned, measured, staged, and reversible.

What you will do

  • Develop the learning data plane around agents: structured trajectories, feedback and outcome signals, offline datasets, lineage, privacy controls, and reliable links between an agent version and its behavior.
  • Create evaluation systems for task completion, tool use, long-horizon behavior, safety, latency, and cost, and own the credibility of the numbers they produce.
  • Design experiment and versioning systems for comparing changes through replay, shadow traffic, canaries, or controlled rollouts, with clear promotion and rollback criteria.
  • Build durable orchestration for long-running tasks, checkpoints, approvals, handoffs, and human-in-the-loop interactions, with typed tool interfaces, protocol-based execution, and memory retrieval that enforces tenant, user, and project visibility boundaries.
  • Own the execution environments agents run in, including isolation, startup performance, resource limits, and cost, and the policy layer that gates self-modification so a proposed change is staged and tested rather than applied silently.
  • Publish reliable APIs, event-driven interfaces, and reusable libraries that work across agent categories and enterprise deployments.

The environment you will work in

Our current Core AI environment includes Python cloud services; React and TypeScript product surfaces in Sycamore Forge; asynchronous and streaming systems; typed APIs and data models; relational and vector data; durable workflows; protocol-based tool execution; multiple model providers; and cloud-native deployment. We use coding agents, automated tests, traces, evaluations, browser automation, offline replay, cost and latency signals, and production feedback as part of everyday engineering. We are building toward a governed collect, learn, evaluate, and apply loop rather than a single monolithic training system.

This is context, not a checklist. We do not require previous experience with every language, framework, model provider, cloud platform, database, or infrastructure tool in our stack. Comparable experience building distributed runtimes, experimentation platforms, retrieval or recommendation systems, workflow engines, developer platforms, or production AI systems is highly relevant.

What we are looking for

  • 5-12 years of software engineering experience. We will make exceptions for exceptional people in either direction.
  • Strong backend and distributed-systems fundamentals, including typed API design, asynchronous workflows, persistence, reliability, and production debugging.
  • Real depth in at least one of the areas above, and genuine interest in working across the others.
  • An empirical approach to AI quality: you can form a hypothesis, design a useful evaluation, interpret noisy evidence, and distinguish a real improvement from movement in a proxy metric.
  • The ability to reason about retries, idempotency, partial failure, long-running state, concurrency, latency, cost, experiment design, and safe rollout.
  • A security-minded approach to multi-tenant systems, identity, authorization, credentials, tool execution, privacy, and auditability.
  • Product judgment. You can find a durable abstraction behind a real requirement without generalizing too early or freezing customer-specific behavior into the platform.
  • AI-native. You use coding agents and modern models as a force multiplier while still owning architecture, correctness, evidence, and operational outcomes.
  • Comfort building product-facing software. You can work in React and TypeScript when a capability needs a great interface, and reason about streaming state and accessibility.
  • Comfort with startup ambiguity, fast feedback loops, and broad ownership.

What strongly signals particular expertise

None of the following is required, and any of them is a strong signal for a particular area:

  • Production Kubernetes work deep enough to have written controllers rather than only configured them
  • A real understanding of isolation boundaries and what each one does and does not contain
  • Comfort with statistics, including the ways an experiment can mislead you
  • Trajectory or event data at scale, including the joins, lineage, and privacy handling that make it usable
  • Reinforcement learning, post-training, or large-scale experimentation

We care more about the systems you personally built, measured, and operated than a particular company, school, language, or model vendor.

Interview process

  • A 30-minute introductory conversation
  • Two 60-minute technical interviews, one focused on systems design and one on coding
  • A take-home assignment where you build and present a real solution using the tools you would use on the job

Why join

  • Build the cloud services that help enterprise agents learn from production experience.
  • Ship Core AI capabilities end to end in Sycamore Forge, from cloud service and API design through the React experience customers use.
  • Turn production outcomes into governed improvements used across customers.
  • Shape how feedback-driven learning, agent evaluation, and safe self-improvement work in a high-trust enterprise setting.
  • Join early enough to shape the Core AI architecture and engineering team.
  • Receive competitive cash compensation and meaningful equity in the company you are helping build.

Similar jobs