Research Engineer / Performance Engineer, RL Distributed SystemsNew$500K

The opportunity

Reinforcement learning is how Claude learns to reason, write code, and act autonomously over long horizons. At frontier scale, an RL run is an unusually demanding distributed system.

What you'll do

  • Design, build, and operate the distributed systems that run RL at scale,: across training, sampling, and environment execution

  • Find and remove whatever currently limits the system, whether it's: scheduling, data movement, storage, networking, or coordination

  • Build fault tolerance into every layer: failure detection, isolation, and recovery that keep long-running jobs making progress without human intervention

  • Design resource management and autoscaling so that compute follows demand as a run's needs shift

  • Build observability that makes it possible to understand what a run is doing: and why it slowed down, stalled, or produced unexpected results

  • Build automation that detects and remediates common problems, and design: interfaces that let engineers and automated tools operate runs safely

What they're looking for

  • Experience running ML training or inference infrastructure at scale
  • Experience across several layers of the stack, such as scheduling, storage, networking, and orchestration
  • Experience building schedulers, autoscalers, or resource management systems
  • Experience with container orchestration such as Kubernetes, and with: sandboxed or virtualized code execution at scale
  • Experience with high-performance networking, RDMA, or collective communication libraries
  • Experience building observability or automated remediation for large fleets
  • Experience with async Python frameworks such as Trio or asyncio
  • Familiarity with reinforcement learning or large language model training workloads