Senior Research Scientist, Model EvaluationActive

The opportunity

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

What you'll do

  • Create ambitious new evaluation benchmarks that push the limits of what our models can accomplish.

  • Work on highly cross-functional teams to translate model feedback into trustworthy, repeatable evaluations.

  • Conduct research to advance the state-of-the-art in LLM evaluation methods,: including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency.

  • Build scalable and reusable tools for digging into model performance.

  • You enjoy rapidly building prototypes that demonstrate the boundaries of what: LLMs are capable of, and you have developed resources to measure those capabilities.

  • You have spent dozens of hours reviewing complex data and LLM outputs to ensure high data quality.

What they're looking for

  • You are obsessive about rigorously measuring AI capabilities, and also about: making sure your measurements actually align with the capabilities you care about.
  • You have strong software engineering skills.
  • A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.
  • Full health and dental benefits, including a separate budget for mental health.