Researcher, Recursive Self-Improvement SafetyActive$295K–$445K

San FranciscoOperations

The opportunity

About the team Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats that could scale to an extreme level of severity. Our work involves: Tracking and prediction.

What they're looking for

  • Model behavior science: Design experiments and evaluations to understand the extent to which models are problematically misaligned, or their safety-relevant capabilities lag behind dangerous capabilities. This may include training model organisms of misbehavior for behaviors not currently present in production, or training interventions to increase safety-relevant capabilities.
  • Coordination and verification: Prototype technical mechanisms for verifying compliance with potential future AI safety agreements.
  • AI R&D risk measurement: Track progress toward automation of technical staff to inform OpenAI’s near-term investments in alignment and security.
  • Maintaining and strengthening RSI safety cases: We’re especially interested in identifying and addressing blindspots of mitigation areas which we may have missed.