The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
Own the eval strategy for North: define what we need to measure across agent workflows, tool use, enterprise knowledge work such as deep research, document creation or editing, and other human-AI interactions.
Build high-quality evals from the realities of the product: user feedback, production failures, privacy-preserving usage logs, internal dogfooding, customer needs, and forward-looking product goals.
Create systems that continuously turn what North is learning from users,: customers, and feature teams into evals, so measurement keeps pace with the product rather than becoming a static benchmark.
Be the voice of North inside modelling: extract clear insight from eval results, customer-derived data and product context, then turn model failures, capability gaps, and North-specific needs into actionable recommendations for the central modelling teams.
You have improved LLM-powered, agent-powered, or AI-product systems through: evals, feedback loops, data curation, prompting, model adaptation, or model selection.
You care deeply about evaluation as a craft: representative tasks, precise rubrics, clean data, failure analysis, regression tracking, and knowing when a metric is giving false confidence.
What they're looking for
- You have strong applied MLE judgment and can reason clearly about model: behavior, eval validity, product outcomes, and production tradeoffs.
- You are comfortable working close to users and product teams, while: translating messy qualitative signals into measurement that other modelling teams can act on.
- You are self-directed, practical, and motivated by open-ended problems where: the right answer requires both technical depth and product understanding.
- You care about making AI systems genuinely useful in real enterprise: settings, not just better on abstract benchmarks.