Staff Software Engineer, Code RLActive$405K

The opportunity

Code RL at Anthropic drives reinforcement learning efforts behind Claude's coding capabilities, creating and scaling agentic coding environments. This is an engineering role with unusual latitude to set technical direction and standards.

What you'll do

  • Design widely-used APIs, frameworks, and abstractions that other engineers: and researchers build on, with careful attention to interface legibility and principled defaults

  • Embed with research teams on a rotational basis: understand their engineering needs, build systems and APIs that support their work, and transfer ownership so teams can maintain those systems after you rotate off

  • Work directly in research codebases, improving reliability and structure: without slowing down the research they support

  • Anticipate silent failure modes and prevent them structurally through type: safety, well-designed invariants, targeted testing, and refactors that shrink the surface area for bugs

  • Contribute to the reliability of production RL systems, including monitoring,: regression detection, and triage tooling

  • Help define engineering standards, review practices, and design patterns for: a new team, and mentor researchers and engineers in adopting them

What they're looking for

  • Experience building infrastructure, tooling, or frameworks for machine learning research or RL workflows
  • Familiarity with reinforcement learning concepts, agentic systems, or LLM training pipelines
  • Experience building or operating large-scale distributed systems
  • Experience building client libraries or SDKs on top of sandboxed, containerized, or remote execution platforms
  • Experience with large-scale data processing or dataset lifecycle management
  • Experience designing plugin systems or extensible class hierarchies used across an organization
  • Experience embedding with or consulting for other teams, including: successfully handing off systems for others to own
  • Experience defining code standards, lint rules, or static verification: approaches adopted across multiple teams