Full-Stack Software Engineer, Reinforcement LearningActive$300K

The opportunity

As a Full-Stack Software Engineer in RL, you'll build the platforms, tools, and interfaces that power environment creation, data collection, and training observability. The quality of Claude's next generation depends on the quality of the data we train it on — and the systems…

What you'll do

  • Build and extend web platforms for RL environment creation, management, and: quality review — including environment configuration, versioning, and validation workflows

  • Develop vendor-facing interfaces and tooling that let external partners: create, submit, and iterate on training environments with minimal friction

  • Design and implement platforms for human data collection at scale, including: labeling workflows, quality assurance systems, and feedback mechanisms that surface reward signal integrity issues early

  • Build evaluation dashboards and observability UIs that give researchers: real-time insight into environment quality, training run health, and reward hacking

  • Create backend services and APIs that connect environment authoring tools,: data collection systems, and RL training infrastructure

  • Build and expand scalable code data generation pipelines, producing diverse: programming tasks with robust reward signals across languages and difficulty levels

What they're looking for

  • Develop onboarding automation and documentation tooling so new vendors and: internal users ramp up in hours, not weeks
  • Partner closely with RL researchers, data operations, and vendor management: to translate ambiguous requirements into well-scoped, well-designed products