The opportunity
Anthropic's RL Data Platform team builds the systems that produce, move, and serve the human data Claude learns from: the interfaces humans use to give feedback, the pipelines that turn raw feedback into training signal, and the tooling researchers use to launch, monitor, and…
What you'll do
Design, build, and operate the feedback and data collection interfaces used: by human annotators, domain experts, and internal researchers.
Build and maintain the backend services, APIs, and pipelines that route model: samples to humans and return structured feedback to training.
Own the reliability, latency, and usability of systems that run continuously against live model endpoints.
Partner with RL researchers to translate loosely specified data needs into: well-scoped collection campaigns and the tooling to run them.
Build dashboards, monitoring, and inspection tools so researchers can see: data quality and throughput without asking an engineer.
Identify and remove the bottlenecks between "we want this data" and "it's in the training mix".
What they're looking for
- Experience building annotation, labelling, evaluation, or other human-in-the-loop data tooling.
- Experience with RLHF, preference data, or other human-feedback pipelines for ML systems.
- Experience shipping researcher-facing or other expert-facing internal tools: people love: interviewing users, hunting down friction, measurably improving the experience.
- Experience running experiments on data collection interfaces and using the results to improve data quality.
- Experience working with crowdworker or expert vendor platforms at scale.
- Familiarity with how LLMs are trained and evaluated.