The opportunity
Note: For this role, we conduct all interviews in Python and prefer candidates to be based in the Bay Area.
What you'll do
Scalable Oversight: Developing techniques to keep highly capable models helpful and honest, even as they surpass human-level intelligence in various domains.
AI Control: Creating methods to ensure advanced AI systems remain safe and: harmless in unfamiliar or adversarial scenarios.
Alignment Stress-testing: Creating model organisms of misalignment to improve our empirical understanding of how alignment failures might arise.
Automated Alignment Research: Building and aligning a system that can speed up & improve alignment research.
Alignment Assessments: Understanding and documenting the highest-stakes and most concerning emerging properties of models through pre-deployment alignment and welfare assessments (see our Claude 4 System Card ), misalignment-risk safety cases, and coordination with third-party evaluators.
Safeguards Research: Developing robust defenses against adversarial attacks, comprehensive evaluation frameworks for model safety, and automated systems to detect and mitigate potential risks before deployment.
What they're looking for
- Model Welfare: Investigating and addressing potential model welfare, moral status, and related questions. See our program announcement and welfare assessment in the Claude 4 system card for more.
- Testing the robustness of our safety techniques by training language models: to subvert our safety techniques, and seeing how effective they are at subverting our interventions.
- Run multi-agent reinforcement learning experiments to test out techniques like AI Debate .
- Build tooling to efficiently evaluate the effectiveness of novel LLM-generated jailbreaks.