Research Engineer / Scientist, AlignmentActive$350K

The opportunity

Note: For this role, we conduct all interviews in Python and prefer candidates to be based in the Bay Area.

What you'll do

  • Scalable Oversight: Developing techniques to keep highly capable models helpful and honest, even as they surpass human-level intelligence in various domains.

  • AI Control: Creating methods to ensure advanced AI systems remain safe and: harmless in unfamiliar or adversarial scenarios.

  • Alignment Stress-testing: Creating model organisms of misalignment to improve our empirical understanding of how alignment failures might arise.

  • Automated Alignment Research: Building and aligning a system that can speed up & improve alignment research.

  • Alignment Assessments: Understanding and documenting the highest-stakes and most concerning emerging properties of models through pre-deployment alignment and welfare assessments (see our Claude 4 System Card ), misalignment-risk safety cases, and coordination with third-party evaluators.

  • Safeguards Research: Developing robust defenses against adversarial attacks, comprehensive evaluation frameworks for model safety, and automated systems to detect and mitigate potential risks before deployment.

What they're looking for

  • Model Welfare: Investigating and addressing potential model welfare, moral status, and related questions. See our program announcement and welfare assessment in the Claude 4 system card for more.
  • Testing the robustness of our safety techniques by training language models: to subvert our safety techniques, and seeing how effective they are at subverting our interventions.
  • Run multi-agent reinforcement learning experiments to test out techniques like AI Debate .
  • Build tooling to efficiently evaluate the effectiveness of novel LLM-generated jailbreaks.