The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
Build new RL environments targeting different agentic capabilities and industry areas
Train and evaluate agents in those environments
Make all the pieces work together: tasks, data, tool implementations, and verifiers
Work across modeling and product to identify gaps in agent performance, and: improve both the agents and the environments
Work with external vendors to create high-quality, expert-built RL: environments, and build tools to ensure high task, data, and verifier quality
Automate the discovery of model capability gaps, and systematically measure: agent performance during evals and training
What they're looking for
- You have engineered agents and optimized them for specific industry use cases
- You have spent dozens of hours reviewing agent trajectories to pinpoint exact: failure points and fix them with model training or harness engineering
- You obsess over measuring agentic capabilities and turning that into a repeatable process
- You have had many debates about what a good outcome from an AI agent should: look like, you translated that into verifier implementations and tuned the reward designs