Research Engineer / Research Scientist, RL FrontiersNew$500K

The opportunity

Reinforcement learning is how Claude learns to reason, write code, and act autonomously over long horizons. The RL Scaling team works on how RL scales: what happens to throughput, stability, and learning efficiency as models get larger, episodes get longer, and compute grows by…

What you'll do

  • Study how RL training and sampling scale with model size, context length, and: compute, and find the algorithmic and systems changes that keep scaling efficient

  • Develop next-generation model architectures and RL algorithms, and make them run efficiently at frontier scale

  • Take promising small-scale results to frontier-scale runs, and diagnose why: they behave differently when they get there, whether the cause is numerical, algorithmic, or systemic

  • Build the experimental infrastructure that sets research velocity: fast, reproducible comparisons of architecture and algorithm variants at meaningful scale

  • Own end-to-end performance of our largest RL runs, from research code down to the hardware

  • Build performance and cost models for proposed architecture and algorithm: changes, and use them to decide which ideas get scaled

What they're looking for

  • Research experience in reinforcement learning, optimization, or large-scale training, published or otherwise
  • Experience developing RL algorithms for language models
  • Experience with scaling laws or other quantitative models of training efficiency
  • Experience designing or modifying transformer architectures beyond standard configurations
  • Experience scaling training to large fleets of accelerators and debugging the: problems that only appear at scale
  • Deep understanding of numerics in large-scale training, including: low-precision formats and sources of instability
  • Familiarity with how GPU or TPU performance characteristics shape architecture and algorithm choices
  • Experience with C++ or Rust