The opportunity
Models are becoming increasingly capable—moving from tools that assist humans to agents that can plan, execute, and adapt in the real world. Mitigating the frontier risks resulting from these capabilities is paramount to OpenAI’s ability to continue deploying models safely.
What you'll do
Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems.
Mitigation. Keeping misalignment safeguards, alignment tools, and on track to: adequately address extreme threats that might arise in the future.
Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness: framework , and partnering with other staff to achieve these targets.
Scalable oversight: Establishing practices for model misbehavior monitoring and oversight which remain effective in superhuman model capability regimes, with a focus on bridging from today’s monitoring approaches to future-proof ones.
Automated auditing: As model capabilities increase, we’ll increasingly rely on automated approaches for finding the most severe forms of model misalignments. We’ll both need to sift through large swaths of production traffic to find the most egregious misalignments, and reliably elicit tail risks before deployment.
Rigorous monitorability: Rigorous testing and red-teaming of our measurements of model misbehavior related to loss-of-control (e.g. reward hacking, sandbagging, scheming). This includes better understanding monitorability , and e.g. preparing for potential losses of Chain-of-Thought monitorability.
What they're looking for
- Model behavior science: Design experiments and evaluations to understand the extent to which models are problematically misaligned, or their safety-relevant capabilities lag behind dangerous capabilities. This may include training model organisms of misbehavior for behaviors not currently present in production, or training interventions to increase safety-relevant capabilities.
- Coordination and verification: Prototype technical mechanisms for verifying compliance with potential future AI safety agreements.
- AI R&D risk measurement: Track progress toward automation of technical staff to inform OpenAI’s near-term investments in alignment and security.
- Maintaining and strengthening RSI safety cases: We’re especially interested in identifying and addressing blindspots of mitigation areas which we may have missed.