The opportunity
Anthropic's Safeguards organization builds the policies, evaluations, and detection and enforcement systems that define and hold the limits on how Claude can be used. In this role, you'll lead our policy design team, managing the teams responsible for radicalization, child…
What you'll do
Lead, develop, and grow the managers and teams responsible for the consumer: harms portfolio, including child safety, user well-being, harmful manipulation, and election integrity
Coordinate policy decisions across the portfolio, and build the mechanisms: that keep them tracked, consistent, and legible — so stakeholders know what was decided, why, and who owns what
Set the strategy for how mitigations built on top of the model: policies, detection and enforcement systems, and product interventions — complement what is trained into the model itself, partnering closely with the alignment training team that owns Claude's character
Prioritize across harm areas competing for the same resources, and make those: tradeoffs and their rationale clear to leadership
Serve as the escalation point for high-severity and ambiguous consumer harms: decisions, including rapid response to emerging risks
Partner with engineering, data science, product, legal, and research across: the model development cycle so consumer harms considerations are represented from training through launch, on every surface where Claude is deployed
What they're looking for
- Subject-matter depth in one or more of the portfolio's harm areas, from: academia, clinical practice, civil society, government, or trust & safety work
- Experience working directly with model training or research teams on model: behavior, or shaping the character of a deployed AI system
- Experience with generative AI safety systems, including LLM-based: classification, evaluation, or enforcement pipelines
- Experience engaging external stakeholders in these domains: child safety organizations, election authorities, mental health experts, or regulators
- Experience using agentic AI tools to scale a team's analysis and operations