Staff+ Software Engineer, Safeguards EvalsActive$320K

Hybrid · San Francisco, CA | New York City, NYTechnology

The opportunity

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole.

What they're looking for

  • Care deeply about AI safety and want your work to have real impact
  • Expertise in building or contributing to LLM or agent evaluation frameworks,: benchmarks, or automated grading systems
  • Extensive experience in trust and safety, content moderation, or abuse detection systems
  • Experience in red teaming, adversarial testing, or jailbreak research on AI systems
Staff+ Software Engineer, Safeguards Evals at Anthropic | Role Match