Post-Training
About Fleet
Fleet studies how environments produce intelligence. We believe intelligence is an emergent property of environmental pressures: the environment determines what capabilities develop, what behaviors survive, and what "good" looks like. We work with frontier labs on post-training across modalities, building benchmarks that expose where frontier models break, training recipes that close those gaps, and scalable oversight for long-horizon agents. Backed by Sequoia Capital, Menlo Ventures, BCV, and SV Angel.
About the Role
We're looking for a Research Scientist to work on agentic RL post-training: the recipes that turn environment interaction into model capability. Tasks and environments are half the problem; the other half is training on them well. You'll train specialized agent models against Fleet's environments, scale the RL runs behind them, and own the algorithmic work that makes those runs pay off.
Responsibilities
- RL post-training recipes for multi-turn agents across modalities (tool use, browser use, terminal use)
- Denser reward signals: verifiers, process supervision, richer feedback than binary outcomes
- Scaling RL runs: larger models, longer horizons, many environments in a single run
- Training specialized models that demonstrate what Fleet's environments teach
- Algorithmic innovations in agentic RL, published where it makes sense
How We Work
- You carry your own research threads that lead to concrete artifacts: a paper, benchmark release, training recipe, or open-source code
- Direct collaboration with frontier labs
Requirements
- Hands-on experience post-training LLMs with RL, end-to-end
- Publication record at top venues (NeurIPS, ICML, ICLR, or equivalent), or shipped training runs at comparable scale
- High ownership: you manage your own research threads and are measured by what you produce
Location
San Francisco (on-site).
Pay
Highly competitive salary + equity.