The opportunity
Large teams of agents are starting to solve problems no single agent can, from rewriting major codebases to formalizing landmark mathematics. Our team studies how these teams scale: what happens to performance, cost and coordination as the number of agents, the compute budget…
What you'll do
Design, run and interpret large-scale experiments on agent teams, reasoning: rigorously about what the data does and doesn't show
Investigate how performance and efficiency change as team size, compute and: task horizon grow, and find the bottlenecks that limit them
Build and scale the systems that run very large agent teams reliably, and: debug the failures that only appear at scale
Design evaluations for long-horizon problems, and keep their results trustworthy
Build the tooling and metrics that let researchers see what a large agent team is doing and why
Partner with research teams across Anthropic so they can run their own: experiments on the platform, and communicate findings clearly
What they're looking for
- Have significant software engineering, ML or research engineering experience
- Have owned something substantial end to end, such as a large system, an: evaluation or benchmark, an agent product, or a research project
- Genuinely enjoy both research and engineering work
- Think quantitatively about complex systems, and think twice before trusting a number