Staff Research Engineer, Model EfficiencyActive

The opportunity

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

What you'll do

  • model architecture and MoE routing optimization

  • decoding and inference-time algorithm improvements

  • software/hardware co-design for GPU acceleration

  • performance optimization without compromising model quality

  • Have a PhD in Machine Learning or a related field

  • Understand LLM architecture, and how to optimize LLM inference given resource constraints

What they're looking for

  • Have significant experience with one or more techniques that enhance model efficiency
  • Strong software engineering skills
  • An appetite to work in a fast-paced high-ambiguity start-up environment
  • Publications at top-tier conferences and venues (ICLR, ACL, NeurIPS)