Audio Inference Engineer, Model EfficiencyActive

The opportunity

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

What you'll do

  • Significant experience developing high-performance audio or machine learning inference systems.

  • Proficiency with programming languages such as C++ and Python.

  • Hands-on experience with deep learning models for audio, speech, or language applications.

  • A bias for action and a strong results-oriented mindset.

  • GPU programming, low-level system optimization, model parallelization techniques over multiple GPUs

  • Have experience with duplex real-time streaming architectures.

What they're looking for

  • Internals of machine learning frameworks for audio (such as PyTorch,: TensorFlow, or specialized audio libraries).
  • Have experience with inference framework like vLLM, SGLang, Tensort-LLM, or: custom distributed inference systems
  • Sequence modeling (e.g., transformers for audio/speech) and end-to-end audio pipeline optimization
  • A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.
Audio Inference Engineer, Model Efficiency at Cohere | Role Match