Research Engineer, Model EvaluationsActive$500K

The opportunity

We're looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do. Your work will turn ambiguous notions of "intelligence" into clear, defensible metrics that researchers, leadership, and the public can rely on.

What you'll do

  • Design and run new evaluations of Claude's capabilities: reasoning, agentic behavior, knowledge, safety properties — and produce visualizations that make the results legible to researchers and decision-makers

  • Build and harden the distributed eval execution platform so hundreds of evals: run reliably against checkpoints throughout production RL training runs

  • Own the dashboards researchers and leadership use to monitor model health: during training, improving signal-to-noise, reducing latency, and making regressions impossible to miss

  • Debug anomalous eval results mid-training-run, determine whether the cause is: a model change or an infrastructure issue, and communicate the answer clearly under time pressure

  • Improve the tooling, libraries, and workflows researchers use to implement and iterate on evaluations

  • Partner with research teams across the full lifecycle of a new capability —: from defining what to measure to interpreting results as training progresses

What they're looking for

  • Hands-on experience using large language models such as Claude, including prompting, sampling, and scaffolding
  • Background in data visualization and a track record of building dashboards people actually trust and use
  • Experience developing robust evaluation metrics for language models
  • Experience with observability, monitoring, or experiment-tracking systems
  • Background in statistics and experimental design
  • Experience with large-scale dataset sourcing, curation, and processing
  • Experience running or supporting ML training infrastructure
  • A bias toward picking up slack and operating flexibly across team boundaries