Research Engineer
Tessera Labs is a new category of enterprise software: an AI platform that changes how the world's largest companies run. Every large enterprise carries decades of accumulated process, data, and code that no longer match the business it has become. Tessera is a transformation engine: a governed, multi-agent platform that understands an enterprise's process, data, and code as one connected system and changes it in weeks rather than years. We're vendor-agnostic by design and sell a product, not a service. We raised a $60M Series A led by Andreessen Horowitz.
About the role
This role focuses on building the model at the core of our platform, specifically the post-training, environments, and evaluation machinery that enable models to operate effectively on unseen enterprise systems. You will work on scaling reinforcement learning (RL) for long-horizon tool use, transformation, and reconciliation, with a strong emphasis on verifiable reward signals. The role involves close collaboration with Research Scientists, owning experiments, and ensuring results translate into production-ready models.
Responsibilities
- Build and scale the post-training stack: SFT, preference optimization, and reinforcement learning for long-horizon tool use, transformation, and reconciliation over enterprise systems.
- Develop memory and context machinery for long-horizon agents, including retention, structuring, retrieval, compaction, and revision strategies.
- Design and implement the representation layer for agents, including ontologies and knowledge graphs derived from real enterprise systems.
- Create data generation and curation pipelines, such as synthetic landscapes, transformation traces, and tool-call trajectories.
- Build RL environments: sandboxed landscapes and execution-and-verification harnesses for automated scoring, backed by real customer data.
- Develop and maintain offline evaluation harnesses for long-horizon agentic behavior, ensuring reproducibility and rigor.
- Run experiments end-to-end, from design to analysis, and distinguish real effects from noise.
- Optimize training and inference throughput, including kernels, parallelism, memory, batching, and serving configurations.
- Take training results from evaluation to production, including quantization, serving, and rollback strategies.
- Establish standards for reproducibility, experiment tracking, and result hygiene.
Representative projects
- Building a synthetic landscape generator to produce realistic enterprise systems and measuring its impact on model performance.
- Debugging a distributed RL loop where the environment leaked target state into observations, inflating scores.
- Post-training a mid-size open-weight model to match frontier-model accuracy at a fraction of the serving cost.
- Reducing end-to-end latency in multi-agent runs through KV-cache reuse and speculative decoding.
- Rebuilding the evaluation harness to ensure reproducibility of six-month-old results.
Requirements
- Significant experience training, fine-tuning, or post-training language models, with demonstrable results.
- RL tuning experience (e.g., RLHF, RLAIF, RLVR, GRPO-family, or agentic RL) is highly preferred.
- Experience with memory and context for long-running agents, including architecture, retrieval, or training.
- Strong software engineering fundamentals; code written for experiments should be maintainable and reproducible.
- Proficiency in Python and PyTorch (or JAX), with the ability to debug distributed training.
- Ability to design, run, and interpret experiments with empirical rigor.
- Experience with GPU infrastructure at scale and understanding of performance bottlenecks.
- Desire to see research deployed in production, treating it as a constraint rather than a tax.
- Clear written communication skills; decisions are made from documents.
Qualifications
Strong candidates may also have:
- Experience building RL environments, execution sandboxes, or verifiable-reward task suites.
- Experience with long-context modeling, including efficient attention, position methods, or evaluation.
- Experience with knowledge graphs, ontologies, or semantic layers over structured enterprise data.
- Experience with agent memory systems in production (episodic or otherwise).
- Experience with code models, including repository-scale context, program synthesis, or transpilation.
- Contributions to open-source ML systems (e.g., vLLM, SGLang, PyTorch, Triton, DeepSpeed, Ray, Megatron, TRL).
- A track record of publications, technical reports, or open-source releases.
- An advanced degree in CS, ML, math, physics, or a related quantitative field, or equivalent industry research experience.
Pay
Compensation range: $200K - $300K.