Research Engineer / Scientist, AlignmentActive$350K

Hybrid · San Francisco, CATechnology

The opportunity

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole.

What they're looking for

  • Model Welfare: Investigating and addressing potential model welfare, moral status, and related questions. See our program announcement and welfare assessment in the Claude 4 system card for more.
  • Testing the robustness of our safety techniques by training language models: to subvert our safety techniques, and seeing how effective they are at subverting our interventions.
  • Run multi-agent reinforcement learning experiments to test out techniques like AI Debate .
  • Build tooling to efficiently evaluate the effectiveness of novel LLM-generated jailbreaks.