Product Manager, Safeguards Rare HarmsActive$305K

The opportunity

Anthropic is dedicated to developing AI assistants that are helpful, harmless, and honest. As usage of our AI services grows, we need to ensure they are not misused.

What you'll do

  • Determine how to build in safety by design upstream and leverage downstream: defenses for Anthropic’s frontier models, AI products, customers on different surfaces - Claude.ai, 1P API, external Cloud providers.

  • Ability to write safety evals and communicate externally about safety.

  • Drive impact via ruthless prioritization by clearly defining problems,: solution options forward, clarity on both business & technical tradeoffs and accordingly clear requirements toward MVP vs. ideal state.

  • Align & collaborate with policy, enforcement, research, engineering and cross functional stakeholders.

  • Understand the AI landscape and ecosystem to plan for mitigation of: deployment risks of increasingly powerful models and determined adversaries.

  • Lead the development of metrics to understand the area, performance,: blindspots to help inform future project planning.

What they're looking for

  • + years in product management with a focus on fast problem understanding,: building roadmaps with tractable progress, ability to get into the details on data, detection & interventions, infrastructure & tools, and/or evals.