The opportunity
The Integrity team is responsible for the scaled systems that help identify and respond to bad actors and harm on OpenAI’s platforms. We protect users and the wider world by developing responses to complex harms and building the harm protections needed for new product and model launches.
What you'll do
Define and drive protection strategy for a harm area or launch. Translate: ambiguous risks into goals, priorities and roadmaps, and make trade-offs based on evidence, risk and available capacity.
Directly build and maintain harm detection. Develop production signals,: classifiers and pipelines using SQL, Python, machine learning and language models, integrating with shared engineering platforms.
Lead cross-functional execution. Translate between specialist roles,: establish shared understanding and ownership, resolve technical and organisational dependencies, and communicate recommendations and specific asks clearly.
Investigate emerging harms and abuse patterns. Work with world class domain: experts to distinguish harmful activity from legitimate use, develop detection criteria, and inform policies and review protocols with Policy and Operations partners.
Evaluate and improve the complete system of defence. Measure remaining harm,: detection quality and intervention effectiveness and collaborate with Data Science on more sophisticated metrics. Use expert feedback and operational outcomes to guide improvements.
Protect product and model launches. Define data, detection and review: requirements; establish monitoring; investigate gaps and critical escalations; and drive improvements informed by post-launch findings.
What they're looking for
- Scale effective approaches with language models, agents and reusable tools.: Choose where automation and human judgement are needed, and build documentation and processes that improve reliability and help others extend the work.
- Think in systems. You have enough breadth across technical, analytical,: investigative and operational disciplines to understand how they fit together, recognise gaps and translate between specialist perspectives.
- Bring deep expertise in at least one relevant harm domain, developed through: integrity, trust and safety, security, investigations or related work. You approach unfamiliar problems with curiosity and respect for expert knowledge and you are driven by reducing harm.
- Are proficient in SQL and Python and have experience writing reliable and: efficient production code for detection and data processing. Additional ML or SWE experience is a plus.