The opportunity
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole.
What they're looking for
- Partner with other RL research teams and product teams to translate: capability goals into training envs and evals
- Have experience with fine-tuning large language models for specific domains or real-world use cases
- Have experience with reinforcement learning, reward design, or training data curation for LLMs
- Are comfortable managing technical vendor relationships and iterating quickly on feedback