Jobs · OTHR · California

Research Scientist

Find Data Science Jobs · San Jose, CA · 6 days ago
OTHR$60/hrFull-time

Tessera Labs is a new category of enterprise software: an AI platform that changes how the world's largest companies run. Every large enterprise carries the same weight—decades of accumulated process, data, and code that no longer match the business it has become. Tessera is a transformation engine: a governed, multi-agent platform that understands an enterprise's process, data, and code as one connected system and changes it in weeks rather than years. We're vendor-agnostic by design—SAP, Salesforce, Workday, Oracle, Snowflake, MuleSoft—and tied to none of them. Governance and generality make this hard: every action is logged, traceable, and reversible, and the platform must work on landscapes it has never seen.

About the role

We're looking for a Research Scientist to set and pursue a research agenda for reliable long-horizon agents operating inside real enterprises. Frontier labs optimize for general capability, but very little rigorous work exists on what it takes for an agent to reason across systems with nineteen years of undocumented decisions, plan a change across forty coupled steps, recover when step twelve reveals the model of the world was wrong, and be right often enough that a CFO signs the go-live. We have the landscapes, traces, and customers to study it.

Two properties make this an unusually good research setting. First, much of the task space is verifiable—a transformation either produces a system that builds, passes regression, and behaves equivalently, or it doesn't. Second, the parts that aren't verifiable are where the interesting work is: designing reward and evaluation across that boundary is the central research question here. You'll invent methods rather than only apply them, work with Research Engineers to run at scale, and hear from a product team within weeks whether you were right. We post-train open-weight models on rented clusters and buy more compute when a result justifies it.

Responsibilities

  • Set and pursue a research agenda on reliable long-horizon agentic behavior in real enterprise environments—decide which questions matter and defend the choice.
  • Invent and validate methods for post-training agents on transformation work: reward design where verification is partial, delayed, or contested; RL formulations for long-horizon planning and tool use; curriculum and data strategy.
  • Define how an agent remembers: memory architecture for runs spanning forty steps and days of wall-clock time—what persists, how it's structured and retrieved, how it's revised when the world changes, and how a model is trained to use it.
  • Own the question of what to measure: develop evaluation methodology whose scores predict customer-observed correctness and demonstrate where cheap automated proxies fail.
  • Study how multi-agent systems fail—error compounding across long trajectories, planning under partial observability, delegation and verification between agents—and design against it.
  • Work on the verification problem directly: how an agent (or another agent) establishes that a change preserved behavior when no test covers it.
  • Solve how a system represents an enterprise to itself: turning process, data, and code into an ontology or knowledge graph an agent can reason over reliably—and one that stays true as the underlying systems change.
  • Investigate what post-training compute and data quantity buy us across model scales we can actually afford, and where the returns bend.
  • Turn findings into things that ship, with Research Engineering and product.
  • Publish papers, technical reports, and open-source artifacts, and represent Tessera's research externally.
  • Raise the team's research bar: review experiment designs, mentor engineers moving into research, and ask whether the result is real.

Representative projects

  • Designing a reward formulation for transformation work that doesn't collapse into reward hacking when the agent discovers it can pass the regression suite by removing the branch the tests don't reach.
  • Developing an evaluation methodology whose scores track expert-reviewed correctness on changes no automated test can verify, and showing where the cheap proxies disagree.
  • Characterizing error compounding across a forty-step transformation plan and proposing a verification scheme that measurably arrests it.
  • Designing a memory architecture for multi-day agent runs and showing, with evidence, that it beats stuffing the context window.
  • Showing that grounding an agent in a learned ontology of a customer's landscape beats retrieval over raw artifacts—or finding that it doesn't, and saving us a year.
  • Running a post-training study across three or four open-weight model sizes and publishing what it says about where domain data quality beats parameter count.
  • Releasing an enterprise agent environment and task suite as an open benchmark, and being honest in the paper about where it doesn't transfer.

Qualifications

  • An MS or PhD in CS, ML, statistics, math, physics, or a related field—or research experience of comparable depth without the credential.
  • A track record of original research in ML: publications at NeurIPS/ICML/ICLR/ACL-tier venues, influential open-source work, or results inside a lab that clearly moved a frontier system.
  • Deep expertise in at least one of: post-training and RL for LLMs, agent memory and long-context reasoning, knowledge representation and structured reasoning, or evaluation methodology. Depth in RL tuning is the strongest signal for this role.
  • Hands-on: you write the code, run the experiments, and read the logs. This is not a role that directs research from a distance.
  • Exceptional research judgment—you pick the question well and kill your own ideas quickly when the evidence says to.
  • Hold rigor and urgency at once: we need results that are true and results that arrive.
  • Write clearly and enjoy explaining a technical argument to people who aren't researchers.

Skills

  • Experience owning a research direction end to end at a frontier lab or a strong academic group.
  • Published work on agents, tool use, RL for LLMs, code generation or repair, reasoning, or evaluation.
  • Experience with verifiable-reward RL, or with the failure modes of reward models where verification is incomplete.
  • Experience with knowledge representation, ontologies, or neurosymbolic approaches to structured domains.
  • Experience with training runs at meaningful scale, including the ones that failed for three weeks.
  • A history of mentoring researchers and engineers into better work.
  • Curiosity about enterprises as a research domain: the problems here are strange and specific, and they reward people who find that interesting.

Similar jobs

Research Scientist

Women In ScienceGreensboro, NC· 3 days ago
OTHRapply on womeninscience.com

Research Scientist

The Henry M. Jackson Foundation for the Advancement of Military MedicineCalifornia, United States· 1 mo ago
OTHR$89k–$140k/yrapply on jobs.hjf.org

Research Scientist

Dignity HealthPhoenix, AZ· 2 mo ago
Healthcare$43.1–$64.11/hrapply on commonspirit.careers

Research Scientist

Corewell HealthGrand Rapids, MI· 1 mo ago
OTHRapply on careers.corewellhealth.org

Research Scientist

North Carolina State UniversityRaleigh, NC· 2 mo ago
OTHRapply on jobs.ncsu.edu

Research Scientist

Original Composites & FiberglassGranville, OH· 2 mo ago
Salesapply on app.jazz.co