Staff AI Architect
About the role
Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you’ll find your place here.
Responsibilities
- Own the AI reference architecture. Define the patterns and standards for how DigitalOcean builds agentic systems—orchestration, tool use, capability boundaries, memory and state, retrieval, evaluation, and observability—and make the governed path the easiest one to take.
- Standards enforced through tooling and paved paths, not approval boards.
- Build alongside the team. Design, prototype, and code the hardest and most ambiguous components, and publish reference implementations others build on.
- Architect the internal AI platform. Shape the shared services every agent depends on: model access and routing, agent runtimes, evaluation harnesses, durable orchestration for long-running stateful workflows that pause and resume across days, and the developer experience that makes all of it self-service.
- Design the tool and capability layer. Define how agents discover and invoke capabilities, built on open standards such as the Model Context Protocol: versioned, schema-defined tools with clear ownership, and evaluation gates before anything becomes available for reuse.
- Establish the integration, identity, and authorization patterns that let agents work safely against systems of record—never replacing a system's own authorization, only narrowing it.
- Re-architect business processes to be AI-native. Sit with the people doing the work—finance, recruiting, sales operations, support, IT—to understand a workflow before designing for it, then partner with functional leaders to find the highest-leverage opportunities and ship them.
- Own the unit economics. Design for cost per completed task, not per token: model selection and routing, context management, caching, batching, and cost attribution teams can act on.
- Multiply the team. Set the technical bar through architecture reviews and high-leverage code, mentor engineers across the US and India, and be the escalation point when a design decision crosses team boundaries.
- Partner across the company. Serve as the architectural counterpart to our platform engineering, security, identity, data, and program management teams—and as technical advisor to business owners evaluating AI capabilities in their own platforms.
Requirements
Substantial experience as a software, solution, or enterprise architect (typically 10+ years), several of them owning architecture above a single project or team.
Recent experience building LLM and agentic systems that ran in production: agent orchestration (LangGraph, CrewAI, AutoGen, Semantic Kernel, OpenAI Agents SDK, or equivalents), Model Context Protocol (MCP) tooling, retrieval and vector stores, and LLMOps discipline—evaluation-first development, prompt and agent versioning, regression testing, observability for non-deterministic outputs, cost attribution.
Depth in at least one production language (Python, Go, TypeScript, or Java) and cloud-native infrastructure—Kubernetes, serverless, APIs, event-driven patterns, observability.
Experience architecting on and integrating Workday, Salesforce, NetSuite, Greenhouse, ServiceNow, or similar—their data models, extensibility limits, native agent layers, and permissioning models.
API and event-driven design, workflow and iPaaS platforms (Workato, MuleSoft, Boomi, n8n, or equivalents), and cloud data platforms.
A clear position on securing autonomous systems: agents as first-class principals rather than credential-holders impersonating humans, short-lived machine identity, vault-backed scoped secrets, delegation with preserved provenance, default-deny tool access, prompt-injection defense, and audit trails your security team can use.
Familiarity with the OWASP Agentic AI risk landscape.
Governance Without Bureaucracy: capability boundaries, an autonomy ladder promoted on evidence rather than anecdote, human-in-the-loop escalation, and data governance for prompts and outputs—plus fluency with the NIST AI RMF, ISO/IEC 42001, and the EU AI Act.
Business Translation: you can sit with a finance analyst or a recruiter, understand what they do all day, and turn an ambiguous business problem into a well-bounded AI system.
Bias for Shipping: pragmatic decisions with incomplete information, unblocking engineers rather than gating them, outcomes over outputs.
Distributed Collaboration: effective across time zones, including close partnership with engineering teams in India, and comfortable in a hybrid environment near our Boston/Cambridge community.