The opportunity
About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats that could scale to an extreme level of severity. Our work involves: Tracking and prediction.
What they're looking for
- Model behavior science: Design experiments and evaluations to understand the extent to which models are problematically misaligned, or their safety-relevant capabilities lag behind dangerous capabilities. This may include training model organisms of misbehavior for behaviors not currently present in production, or training interventions to increase safety-relevant capabilities.
- Coordination and verification: Prototype technical mechanisms for verifying compliance with potential future AI safety agreements.
- AI R&D risk measurement: Track progress toward automation of technical staff to inform OpenAI’s near-term investments in alignment and security.
- Maintaining and strengthening RSI safety cases: We’re especially interested in identifying and addressing blindspots of mitigation areas which we may have missed.