Researcher, Alignment CoT Monitorability
About the Team
The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. We study how to measure monitorability, which training mechanisms affect monitorability, and speculative methods to improve monitorability. While we mostly focus on CoT monitorability at the moment, we care more generally about any form of monitorability, auditing methods, and improving alignment. We were the first to show that chain-of-thought monitoring can be a practical additional safety mechanism, and today our monitoring systems are actively used on OpenAI's largest RL training runs to detect misbehavior. The issues we surface are then used to help improve our reward functions, environments, etc (without directly training against a CoT monitor). Our work sits in Alignment and intersects with model training, alignment evaluations, monitoring, and frontier-risk research. We care most about monitorability where the stakes are high, and about preserving useful oversight signals as models become more capable.
About the Role
We're looking for a researcher with strong empirical ML expertise and a deep interest in model behavior, alignment, or interpretability. Direct chain-of-thought interpretability experience is welcome but not required; strong candidates may come from broader interpretability, alignment, model training, or investigative model-behavior work. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees.
Responsibilities
- Design and run empirical studies of chain-of-thought monitorability across frontier reasoning models and training settings.
- Build evaluations that measure whether monitors can reliably predict properties of interest, including high-stakes forms of misbehavior.
- Investigate how pre-training, synthetic data, mid-training, post-training, reinforcement learning, and other interventions improve or degrade monitorability.
- Analyze model behavior and turn observations from monitoring into hypotheses, experiments, and recommendations.
- Translate research findings into practical monitoring and oversight approaches that can inform real training runs.
- Collaborate with researchers and engineers across model training, alignment evaluations, monitoring, and frontier-risk work.
- Produce externally publishable research when results advance the broader science of alignment.
Qualifications
- Strong hands-on experience training, evaluating, or debugging large ML models, especially LLMs.
- Deep curiosity, interest in alignment, and high agency.
- Depth in alignment, interpretability, model behavior, empirical ML, or adjacent research.
- Excitement to investigate chain-of-thought monitorability, monitoring methods, and scalable oversight.
- Ability to turn ambiguous research questions into measurable experiments and follow the evidence when results are subtle or noisy.
- Comfort moving between research ideation and engineering execution.
- Curiosity about multiple approaches to understanding model behavior and openness to different methodological lenses.
- High independence while collaborating closely across research and engineering teams.
- Commitment to making increasingly capable AI systems more monitorable, trustworthy, and safe.
Schedule
Hybrid work model of 3 days in the office per week.