The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
Build and own the training framework responsible for large-scale LLM training.
Design distributed training abstractions (data/tensor/pipeline parallelism,: FSDP/ZeRO strategies, memory management, checkpointing).
Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100).
Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
Collaborate closely with infra teams to ensure our cluster, container: environments, and hardware configurations support high-performance training.
Investigate and resolve performance bottlenecks across the ML systems stack.
What they're looking for
- Build robust systems that ensure reproducible, debuggable, large-scale runs.
- Strong engineering experience in large-scale distributed training or HPC: systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.
- Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).
- Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines.