Model Policy Manager, Agentic SafetyNew$207K–$335K

The opportunity

Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency.

What you'll do

  • Identify vulnerabilities that emerge as models interact with tools, data, and: external systems, and translate them into model- and system-level safeguards.

  • Develop threat models and empirical frameworks for understanding harmful outcomes from misaligned behavior.

  • Build frameworks for understanding harmful outcomes arising from model misalignment.

  • Identify the underlying behaviors and system conditions that drive those outcomes.

  • Turn findings into policy frameworks, evaluation criteria, online measurement and safeguards.

  • Develop human data campaigns and gold sets to ground measurement and: evaluation of emerging behaviors and risks.

What they're looking for

  • Partner with research, engineering, security, and product teams to shape: model and system safety, balancing difficult trade-offs between safety, utility, and business risk.
  • Inform deployment decisions, system cards, safeguards reports, and OpenAI’s: broader approach to agentic safety.
  • Build monitoring approaches that detect regressions and emerging risks after deployment.
  • Brings a strong background in AI agent safety, privacy, security,: cybersecurity, or adjacent fields, with the adversarial mindset needed to investigate real-world harmful outcomes.