Oversight
Fleet AI, Inc. · San Francisco, CA · 3 wk ago
OTHRFull-time
About the role
We're looking for a Research Scientist to work on scalable oversight for agents: detecting misbehavior like reward hacking and deception, understanding why it happens, and keeping verification reliable as models outgrow the humans and automated systems checking them. Because Fleet builds sandboxed environments and training tasks for agents, our task factory will be your controlled laboratory where you can intervene, attribute behavior to its cause, and feed what you learn back into improvements. You'll own research threads that ship as papers, benchmarks, training gyms, models optimized for efficient oversight, or production systems, in direct collaboration with frontier labs.
Responsibilities
- Scalable oversight: methods that keep verification reliable as tasks get longer and models get stronger.
- Monitoring: detecting reward hacking or undesirable behavior patterns in our agent traces.
- Failure mode analysis and attribution: tracing agents' mistakes and misbehaviors upstream to root causes in data, environments, or model weights.
- Turning oversight into training signal: better verifiers, process supervision, and models trained to catch what humans miss.
Requirements
- Publication record at top venues (NeurIPS, ICML, ICLR, or equivalent), or equivalent shipped research, in oversight, monitoring, evaluations, alignment, interpretability, AI control, or RL.
- Strong empirical instincts; comfortable with LLM APIs, harnesses, and large trace datasets.
- High ownership: you run your own research threads and are measured by what you ship.
Location
San Francisco (on-site).
Pay
Highly competitive salary + equity.