Research Engineer
Tessera Labs is a new category of enterprise software: an AI platform that changes how the world's largest companies run. Every large enterprise carries decades of accumulated process, data, and code that no longer match the business it has become. Tessera is a transformation engine: a governed, multi-agent platform that understands an enterprise's process, data, and code as one connected system and changes it in weeks rather than years. We're vendor-agnostic by design—SAP, Salesforce, Workday, Oracle, Snowflake, MuleSoft—and tied to none of them. Governance ensures every action is logged, traceable, and reversible, while generality allows the platform to work on landscapes it has never seen.
About the Role
This role focuses on the model itself at scale. Frontier models have never seen most of what we work on—proprietary dialects, customer-specific configurations, and tasks that run forty steps before validation. You’ll build the training, environment, evaluation, and inference machinery that turns hypotheses about agent behavior into measurable results and shipped models. You’ll work closely with Research Scientists, own experiments, and contribute to a verifiable RL setting where transformations are either correct or not, providing a real reward signal.
Responsibilities
- Build and scale the post-training stack: SFT, preference optimization, and reinforcement learning for long-horizon tool use, transformation, and reconciliation over enterprise systems.
- Develop memory and context machinery for long-horizon agents: retention, structuring, retrieval, compaction, and revision, including training models to use it effectively.
- Build the representation layer agents reason over—ontologies and knowledge graphs derived from real enterprise systems—and the pipelines to construct, validate, and maintain them.
- Design and implement data generation and curation pipelines: synthetic landscapes, transformation traces, tool-call trajectories, and curriculum infrastructure for unseen systems.
- Build RL environments: sandboxed landscapes and execution-and-verification harnesses for automatic scoring, backed by real customer traces where synthetic data isn’t sufficient.
- Develop and run offline eval harnesses for long-horizon agentic behavior, ensuring reproducibility and methodological rigor.
- Run experiments end-to-end: design, launch, debug, analyze, and distinguish real effects from noise.
- Optimize training and inference throughput: kernels, parallelism, memory, batching, and serving, with a focus on long-context efficiency.
- Take training results from eval to production: quantization, serving configuration, and rollback paths.
- Establish standards for reproducibility, experiment tracking, and result hygiene.
Representative Projects
- Built a synthetic landscape generator producing enterprise systems with realistic complexity, measuring its impact on held-out customer environments.
- Stood up a distributed RL loop with reward from build-and-regression outcomes, identifying and fixing observation leaks.
- Post-trained a mid-size open-weight model to match frontier-model accuracy on core transformation tasks at a fraction of the serving cost.
- Halved end-to-end latency on multi-agent runs via prefix KV-cache reuse and speculative decoding.
- Rebuilt the eval harness to ensure six-month-old results could be reproduced from a commit hash.
Requirements
- Significant experience training, fine-tuning, or post-training language models, with demonstrable results you owned.
- RL tuning experience (RLHF, RLAIF, RLVR, GRPO-family, or agentic RL) is nearly a requirement.
- Deep thinking about memory and context for long-running agents, whether through architecture, retrieval, or training.
- Strong software engineering fundamentals—experimental code should be maintainable and runnable by others.
- Fluency in Python and PyTorch (or JAX), with comfort debugging distributed training.
- Ability to design, run, and interpret experiments with empirical rigor, distinguishing effects from noise or bugs.
- Experience with GPU infrastructure at scale and understanding of time/memory trade-offs.
- Desire to see research deployed in production, treating it as a constraint rather than a tax.
- Clear written communication—decisions are made from documents.
Qualifications
- Experience building RL environments, execution sandboxes, or verifiable-reward task suites.
- Experience with long-context modeling: extension, efficient attention, position methods, or rigorous evaluation.
- Experience with knowledge graphs, ontologies, or semantic layers over structured enterprise data.
- Experience with agent memory systems (episodic or otherwise) in production, not just theory.
- Experience with code models: repository-scale context, program synthesis, automated repair, or transpilation.
- Contributions to open-source ML systems (vLLM, SGLang, PyTorch, Triton, DeepSpeed, Ray, Megatron, TRL, or similar).
- Track record of publications, technical reports, or open-source releases.
- Advanced degree in CS, ML, math, physics, or related quantitative field, or equivalent industry research experience.
Benefits
Compensation Range: $200K - $300K