AI Engineer
About Tessera Labs: Tessera Labs is a new category of enterprise software—an AI platform that transforms how the world's largest companies operate. Every large enterprise carries decades of accumulated process, data, and code that no longer align with its current business. Changing any of it typically requires years, hundreds of millions of dollars, and armies of consultants, often failing more often than acknowledged. Tessera is a transformation engine: a governed, multi-agent platform that understands an enterprise's process, data, and code as one connected system and changes it in weeks rather than years. We're vendor-agnostic by design, working across SAP, Salesforce, Workday, Oracle, Snowflake, and MuleSoft without being tied to any. The challenges are governance (every action is logged, traceable, and reversible) and generality (the platform must work on landscapes it has never seen). We sell a product, not a service, focusing on making the product successful.
About the role
This role focuses on the systems surrounding the model, owning agents end-to-end: the harness they run in, the tools they call, the context they see, the guardrails around them, and the evals that measure improvement. The core challenge isn’t model access but enabling agents to reason across complex, undocumented enterprise landscapes where failures must be legible, reproducible, and preventable. This is not a prototyping role—everything you build will interact with systems critical to a company’s quarterly operations.
Responsibilities
- Design and ship production agents that perform transformation work: understanding landscapes, planning changes, executing across process/data/code, and proving correctness.
- Build the tool layer: typed, permissioned, well-documented interfaces for agents to read and act across enterprise systems without exceeding human-authorized permissions.
- Improve agent performance through prompting, context construction, tool-use strategy, and decision logic, diagnosing failures to determine the right fix.
- Develop the retrieval layer over enterprise artifacts and metadata: chunking/indexing non-prose content, hybrid search, reranking, grounding, and evaluation.
- Manage context deliberately: handle large enterprise artifacts and accumulated history by deciding what the model sees, compresses, or drops.
- Use classical ML where appropriate (e.g., routing, ranking, classification, anomaly detection) instead of defaulting to LLM calls.
- Build the eval and monitoring layer: task sets from real customer landscapes, regression coverage, and alerting for schema changes, revoked permissions, or model updates.
- Instrument every run to reconstruct model calls, tool invocations, decisions, and human approvals for auditability.
- Diagnose production failures from execution traces to root cause, closing systemic gaps rather than patching prompts.
- Design guardrails, approval gates, and rollback paths to ensure agent safety in live enterprise environments.
- Build generalizable capability—solutions that work across customers, not just one-off fixes.
Representative projects
- Shipping an agent that lets a customer retire a third of their custom estate in a quarter, plus a review surface for human approval in minutes instead of days.
- Designing approval and write semantics for agents in live landscapes to ensure traceable human decisions and no inconsistent partial failures.
- Building a reconciliation agent to align data between merged companies, with evals to distinguish confidence from guesswork.
- Standing up a replay-evaluation pipeline to score candidate agent versions against recorded runs (real system responses, mocked writes) to catch regressions pre-deployment.
- Turning customer-specific patterns into platform primitives that work across multiple customers without manual rewrites.
Requirements
- 3+ years building and operating production software, with recent experience on LLM-powered systems.
- Shipped agentic systems that real users depend on and owned incidents when they failed.
- Fluency in the agent toolkit: tool calling, orchestration, context engineering, RAG, evals, and tracing, with awareness of their limitations.
- Built retrieval systems for messy real-world corpora and measured their effectiveness.
- Background in traditional ML (supervised learning, feature engineering, model evaluation) and ability to use it when it outperforms prompts.
- Treat evaluation as engineering, not an afterthought.
- Think in systems and customer outcomes, not just model metrics.
- Comfort with non-determinism and experience building reliable systems on top of it.
- Strong Python skills and comfort with TypeScript.
- Ship fast and verify before claiming success.
- Interest in large, undocumented enterprise systems—this instinct is core to the role.
Nice to have
- Hands-on experience with enterprise platforms (SAP, Salesforce, Workday, Oracle, Snowflake, MuleSoft, ServiceNow) and their APIs/extension models.
- Experience with knowledge graphs, ontologies, or semantic models over messy real-world systems.
- Experience with code analysis, program transformation, or automated refactoring at scale.
- Familiarity with MCP, sub-agents, or agent skill/plugin architectures.
- Background in distributed systems or workflow engines, where partial failure is the default.
- Early-stage startup or founder experience, where scope adapts to immediate needs.
- Familiarity with enterprise security/compliance: SSO, RBAC, segregation of duties, PII handling, SOC 2, data residency.
Benefits
Compensation Range: $200K - $250K