Researcher, Safety Training, National SecurityNew$380K–$500K

The opportunity

The Safety Training research team aims to fundamentally advance our capabilities for precisely implementing safe behavior in AI models, and to leverage these advances to make OpenAI’s deployed models safe and beneficial. This requires a breadth of new ML research to address the…

What you'll do

  • Research and implement methods for safety training, reinforcement learning, and adversarial robustness.

  • Develop evaluations, identify model failure modes, and use findings to improve training.

  • Work with research, engineering, security, and policy partners to support safe, reliable deployment.

  • Bring 4+ years of relevant AI safety research experience, including RLHF, adversarial training, or robustness.

  • Have a degree in computer science, machine learning, or a related field, and: strong deep learning research or engineering skills.

  • Have experience improving model safety for deployment and enjoy collaborative research.

What they're looking for

  • Are motivated by OpenAI’s mission and the responsible use of AI in safety-critical settings.
  • Active TS/SCI clearance or equivalent.