Jobs · Engineering

Head of AI Engineering at AIOS — Remote, $200-$400k/yr + equity

AIOS · San Francisco, CA · 3 wk ago
RemoteRemoteEngineering$200k–$400k/yrFull-time

About the role

We are building a world-leading Applied AI team. As Head of AI Engineering at AIOS, your fundamental role is to build the AIOS Agent SDK and make it the foundation for world-class agents across the company. You are not joining to discover our first AI use case or build another chatbot. Our customer-facing agent gathers context from across our product and customer history, retrieves the right knowledge, reasons through multi-step cases, and decides when to act, respond, escalate, or stand down. Our agentic workflows autonomously generate >$100k of revenue per day.

This existing harness will become the nucleus of the AIOS Agent SDK. You'll separate its reusable foundations from its customer-support logic and turn them into a strongly opinionated internal platform. We're also building an AI clinical decision-support system. This helps clinicians evaluate patient eligibility, contraindications, dosing, and risk. It will be the second major system built on the SDK and, over time, a foundation for increasingly autonomous clinical decisions. Once the SDK has proven itself through these two tools, it will become the default foundation for new agents across AIOS.

You'll be the DRI for agent architecture, model strategy, evals, AI reliability, technical safety, provider relationships, and the shared runtime. You'll make these decisions autonomously. You'll ensure we use the best model for each job based on measured quality, reliability, speed, and cost.

This is a technical leadership role. You'll lead by example as you grow the team. You'll ensure:

  • The AIOS Agent SDK exists and is running the show in production, and it is in exceptionally safe technical hands
  • Product engineers can build excellent agents without recreating context, tool, safety, eval, and observability infrastructure
  • Our agents become more capable without becoming less predictable
  • Major changes are supported by trustworthy evidence across quality, reliability, safety, latency, and cost
  • Production failures continuously strengthen our evals, architecture, and models
  • Our engineers actively seek your judgment and trust the direction you set
  • AIOS is clearly an industry leader in applied AI for production healthcare systems

This is a full-time, fully remote role. You can work async in the timezone of your choice, provided you're regularly available until midday Pacific Time for collaboration. This is a senior role. You'll report directly to the VP of Engineering. You'll also work closely with the Head of Product, UK Clinical Lead, and CEO.

Responsibilities

  • Agent SDK: Turn Jesse's (customer support tool) existing harness into the strongly opinionated internal platform powering Jesse, Aegis (clinical support tool), and future AIOS agents. Own its architecture, reusable primitives, supported extension points, developer experience, and integration with our existing infrastructure.
  • Jesse & Aegis: Become the senior technical owner of Jesse and work closely with the engineers and Clinical Product team building Aegis. Improve both systems while extracting the shared foundations they need across context, retrieval, memory, orchestration, tools, state, and escalation.
  • Evals & Experimentation: Build trustworthy benchmarks using deterministic checks, simulations, model-based graders, human judgment, and production outcomes. Establish the path from offline evaluation to controlled production experiments so major changes ship with evidence.
  • Production Learning Loop: Turn traces, poor resolutions, escalations, incidents, tool failures, and successful outcomes into better evals, stronger architecture, improved models, and permanent platform capabilities.
  • Safety & Compliance: Make consequential agent actions safe through authorization, validation, idempotency, auditability, recovery, and human handoff. Encode compliance, privacy, security, and regional requirements into the platform wherever possible.
  • Models & Economics: Own model selection, routing, fallbacks, caching, and our ~$200k monthly model spend. When the evidence supports it, lead the data preparation, fine-tuning, evaluation, and AIOS-controlled deployment of specialized open-weight models.
  • Reliability: Own the shared runtime in production, including tracing, observability, testing, provider resilience, capacity, and incident response. Be the senior engineering DRI when an AI system behaves unsafely, quality regresses, or the platform fails.
  • Technical Leadership: Set AIOS's AI architecture and strategy in close partnership with the VP of Engineering. Make the final call on major technical decisions, guide engineers across product pods, and remain hands-on by writing production code and personally building the most important foundations.
  • Build the Team: Inherit one engineer and build the Applied AI team to approximately five exceptional people during your first year. Own our technical relationships with leading model providers and represent AIOS externally where doing so strengthens our work.

Requirements

  • Experience: 8+ years of software engineering experience and remain an active production contributor
  • Education: At least a bachelor's degree in Computer Science, Machine Learning, or a closely related technical field
  • Production Agents: Personally built and shipped an exceptional agentic system used by real customers that reasoned across multiple steps, used tools, changed state, and operated under real production constraints
  • Agent Architecture: Deep reasoning about harnesses, orchestration, context construction, retrieval, memory, state, tool design, structured workflows, and error recovery
  • Evals: Built or meaningfully owned evaluation systems for probabilistic products including dataset construction, evaluator design, simulations, regression detection, noisy metrics, and the relationship between offline performance and production outcomes
  • Software Engineering: Strong systems-engineering fundamentals including APIs, distributed systems, concurrency, queues, databases, observability, failure modes, and production reliability
  • Consequential Actions: Strong judgment around authorization, validation, idempotency, state transitions, auditability, recovery, and escalation to let an agent act safely
  • Model Judgement: Understand the capabilities and limitations of current frontier and open-weight models; know when the model is the problem and when the real problem is context, tools, data, orchestration, or evaluation
  • Open-Weight Models: Enough technical depth to lead fine-tuning and AIOS-controlled deployment of open-weight models when the evidence supports doing so
  • Leadership: Successfully led and managed a small technical engineering team; set clear direction, raised the quality bar, developed strong engineers, and addressed underperformance
  • Technical Authority: Strong engineers trust your judgment; can make difficult decisions, explain trade-offs clearly, and push back without hesitation when a proposed approach is unsound
  • Communication: Can explain difficult technical ideas to engineers, product leaders, clinicians, and executives without flattening the important details
  • Independence: Create clarity in ambiguous environments and make high-quality decisions without hand-holding
  • Builder: Still write production code; lead from inside the work rather than managing it from a distance
  • Ownership: When quality drops, costs spike, tools fail, or providers degrade, take responsibility for reaching the outcome rather than identifying whose component was technically at fault

Preferred qualifications

  • Built runtimes, SDKs, harnesses, tool layers, evaluation platforms, or shared AI infrastructure used by other engineers
  • Built high-volume customer-service, commerce, or transactional agents operating across complex, multi-step customer journeys
  • Worked on healthcare, financial, insurance, or other systems where correctness, traceability, and careful rollout matter
  • Fine-tuned, distilled, evaluated, or deployed an open-weight model for a specific production workflow
  • Built durable memory, context compression, personalization, or agents operating across sessions and extended periods
  • Worked on voice agents, streaming systems, or other latency-sensitive AI experiences
  • Worked directly with frontier model providers on evaluations, technical issues, capacity, pricing, or early access
  • Strong nose for exceptional AI engineers and know how to create an environment in which they do their best work
  • Can translate relevant research into reliable production systems without confusing novelty with progress
  • Can move from debugging a production trace, to redesigning an eval, to reviewing an agent abstraction, to handling a provider incident

Similar jobs