Engineering Manager, Safeguards InterventionsActive$405K

Hybrid · San Francisco, CATechnology

The opportunity

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole.

What they're looking for

  • Care about measurement: you've built (or insisted on) the evals that prove a system does what it claims, and you've killed things that didn't.
  • Can drive ambiguous, multi-stakeholder tradeoffs (safety vs UX vs latency vs: cost) to a decision and own the outcome.
  • Care deeply about AI safety and want your team's work to be the reason advanced models can be deployed at all.
  • Have worked in trust & safety, integrity, or abuse-prevention engineering at scale.