The opportunity
As a Safeguards Enforcement Analyst focused on Violence & Extremism, you will be responsible for building and executing operational workflows to assess model behavior, drive enforcement decisions, and develop evals across a technically demanding range of policy areas. Your work…
What you'll do
Design and architect automated enforcement systems and review workflows that: scale effectively while maintaining high accuracy
Develop and maintain evals that measure model performance on these policy: areas, surface regressions, and inform policy and model improvements
Partner with Engineering and Data Science to optimize detection and automated: enforcement systems for potential policy violations
Review flagged content to drive enforcement decisions and surface policy: gaps, with particular attention to novel or technically sophisticated misuse attempts + emerging extremist movements, ideologies, and mobilization tactics
Support the Safeguards policy design team by providing structured feedback on: policy gaps and enforcement ambiguities based on real enforcement scenarios
Develop and maintain enforcement guidelines and reviewer documentation that: enable accurate, consistent enforcement across a wide range of content
What they're looking for
- Subject matter expertise in one or more high-stakes harm areas, such as: weapons and dangerous technology, violent extremism, terrorism, autonomous systems, or critical infrastructure protection
- Familiarity with relevant legal and regulatory frameworks governing dangerous: technology, critical infrastructure, or domestic/international terrorism
- Experience developing evals or red-teaming AI systems, particularly for: harmful content or policy enforcement use cases
- Experience with threat actor profiling and threat intelligence frameworks (e.g., MITRE ATT&CK)
- Experience tracking threat actors, extremist networks, or misuse patterns: across surface, deep, and dark web environments
- Experience with large language models and an understanding of how AI: technology could provide meaningful uplift toward serious harm
- Proficiency in Python for data analysis and workflow automation
- Background in law enforcement, national security, defense, counterterrorism,: or a relevant regulatory environment