Member of Technical Staff, Model EfficiencyActive

The opportunity

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

What you'll do

  • + years of experience writing high-performance, production-quality code

  • Strong programming skills in C++ or Python (Rust/Go also welcome)

  • Experience working with large language models and familiarity with the LLM: inference ecosystem (e.g., vLLM, SGLang, etc.)

  • Ability to diagnose and resolve performance bottlenecks across the model execution stack

  • A strong bias for action: you ship fast, measure impact, and iterate

  • GPU programming, CUDA, or low-level systems optimization

What they're looking for

  • Language modeling with transformers (MoE, speculative decoding, KV-cache optimizations)
  • Scaling performance-critical distributed systems (e.g., computation, search, storage)
  • A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.
  • Full health and dental benefits, including a separate budget for mental health.