The opportunity
Anthropic's Safeguards team is responsible for enforcing our policies, protecting users, and ensuring our platform is not misused. As a Safeguards Enforcement Analyst focused on Safety Evaluations, you'll play a central role in ensuring our models meet safety and policy standards before and after launch.
What you'll do
Support model launch readiness by running evaluations, monitoring and: interpreting results, and surfacing regressions or unexpected behavior changes to relevant stakeholders
Partner closely with policy and domain experts throughout the evaluation: lifecycle — from identifying risks and scoping the right evaluation approach, to coordinating creation of new evals and ensuring existing ones remain current with evolving policies, threat vectors, and model capabilities
Work with cross-functional stakeholders to help manage evaluation outcomes,: including interpreting results and driving mitigations where needed
Think strategically about eval quality to build processes and eval paradigms: that keep evaluations unsaturated, high-signal, and insightful as models improve
Build out processes and frameworks for creating product-specific evaluations: as Anthropic's product surface area expands
Help design and scope tooling improvements that accommodate evolving eval: needs and expand self-serve eval creation and iteration for non-technical users
What they're looking for
- Write and maintain rigorous documentation for evaluation creation, execution,: and interpretation as the team builds out eval tooling and processes