The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
Evaluate and rank model outputs: Complete preference and comparison tasks, assessing which responses best conform to project guidelines for accuracy, helpfulness, tone, and safety, and writing clear justifications for your judgments.
Stress-test and break models: Probe models adversarially to surface failure modes, unsafe behavior, and capability gaps, and document reproducible cases that engineering and research teams can act on.
Create datasets: Author high-quality prompts, responses, and exemplars to build training and evaluation datasets, following detailed specifications and editing machine-written or human-written outputs to standard.
Build and apply rubrics and taxonomies: Contribute to the design of grading criteria and rubrics, then apply them consistently to produce structured, high-quality annotations across task types.
Annotate and correct multimodal data: Label, audit, and rectify inaccuracies across text, image, and structured data, maintaining a high standard of data integrity and accuracy.
Calibrate and maintain consistency: Participate in calibration exercises and inter-annotator agreement checks to align on standards, and flag ambiguous or uncovered edge cases rather than guessing, since a single misjudgment replicated at scale degrades a model.
What they're looking for
- Adapt to experimental work: Take on new and evolving task types as project needs shift, applying sound judgment in areas where guidelines are still being developed.
- Report on model performance: Surface and communicate quality and performance trends in model and agent behavior, giving cross-functional partners clear, well-evidenced feedback on where models succeed, fail, and degrade.
- + years of experience in AI data annotation, LLM evaluation, content: moderation, research, or a related analytical role, with exposure to quality assurance, and/or preference ranking.
- Experience applying detailed guidelines to complex and often ambiguous: content, with strong contextual and sociocultural judgment, sensitivity to nuance, tone, and register, and the ability to reason well in cases where there is no single correct answer.