Product Manager, Agent Harness & ModellingActive$160K–$320K

The opportunity

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

What you'll do

  • Define and own the roadmap for North's agent harness, including the agent: loop, context engineering layer, tool orchestration, sandbox execution, and sub-agent delegation

  • Serve as the primary interface between North engineering and Cohere's: Modeling team, ensuring new harness capabilities are validated before being built and that neither team paints itself into a corner

  • Own North's agentic evaluation framework, ensuring evals are compatible with: both the North harness and Modeling's training infrastructure, and that they serve as a reliable bridge between product and research

  • Engage enterprise customers to surface real-world agentic failures and: translate findings into concrete product and model requirements

  • Partner with the agentic expert modeling team to define goals that maximize: the impact of our models and prioritize model capabilities and evals most critical for delivering on those goals. Establish a clear roadmap for our agentic modeling capabilities.

  • Stay current with the open-source and commercial agent ecosystem and drive: adoption decisions that keep North's architecture aligned with emerging standards.

What they're looking for

  • + years of product management experience as a research PM, in agentic AI systems,, or applied ML products
  • Deep understanding of agentic harnesses, including the latest paid and open source offerings.
  • Deep understanding of modern LLM agent architectures, including multi-agent: systems, tool-augmented reasoning, memory and retrieval, programmatic orchestration, RAG, and long-horizon execution
  • Strong grasp of agentic evaluation design, including how to measure task: completion, failure recovery, and long-horizon reliability, and how to diagnose model vs. scaffolding gaps
  • Technically deep enough to contribute to architecture decisions at the: implementation level: comfortable reviewing and shaping design docs, reasoning about async execution patterns, sandboxed environments, filesystem design, and the tradeoffs that come with building harness capabilities into a production platform
  • Ability to flex between ML research conversations and engineering architecture discussions with equal fluency
  • Track record of shipping platform-layer products with demonstrated impact on: reliability, performance, or capability.