The opportunity
Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency.
What you'll do
Identify vulnerabilities that emerge as models interact with tools, data, and: external systems, and translate them into model- and system-level safeguards.
Develop threat models and empirical frameworks for understanding harmful outcomes from misaligned behavior.
Build frameworks for understanding harmful outcomes arising from model misalignment.
Identify the underlying behaviors and system conditions that drive those outcomes.
Turn findings into policy frameworks, evaluation criteria, online measurement and safeguards.
Develop human data campaigns and gold sets to ground measurement and: evaluation of emerging behaviors and risks.
What they're looking for
- Partner with research, engineering, security, and product teams to shape: model and system safety, balancing difficult trade-offs between safety, utility, and business risk.
- Inform deployment decisions, system cards, safeguards reports, and OpenAI’s: broader approach to agentic safety.
- Build monitoring approaches that detect regressions and emerging risks after deployment.
- Brings a strong background in AI agent safety, privacy, security,: cybersecurity, or adjacent fields, with the adversarial mindset needed to investigate real-world harmful outcomes.