The opportunity
The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. We study how to measure monitorability , which training mechanisms affect monitorability, and speculative methods to improve monitorability.
What you'll do
Design and run empirical studies of chain-of-thought monitorability across: frontier reasoning models and training settings.
Build evaluations that measure whether monitors can reliably predict: properties of interest, including high-stakes forms of misbehavior.
Investigate how pre-training, synthetic data, mid-training, post-training,: reinforcement learning, and other interventions improve or degrade monitorability.
Analyze model behavior and turn observations from monitoring into hypotheses,: experiments, and recommendations.
Translate research findings into practical monitoring and oversight: approaches that can inform real training runs.
Collaborate with researchers and engineers across model training, alignment: evaluations, monitoring, and frontier-risk work.
What they're looking for
- Produce externally publishable research when results advance the broader science of alignment.
- Have strong hands-on experience training, evaluating, or debugging large ML models, especially LLMs.
- Have deep curiosity, interest in alignment, and high agency.
- Bring depth in alignment, interpretability, model behavior, empirical ML, or adjacent research.