The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
Building a high-performance data loading and caching pipeline.
Implementing performance profiling across the ML systems stack
Developing internal metrics and monitoring for training runs.
Building reproducibility and regression testing infrastructure.
Developing a performant fault-tolerant distributed checkpointing system.
Build and own the training framework responsible for large-scale LLM training.
What they're looking for
- Design distributed training abstractions (data/tensor/pipeline parallelism,: FSDP/ZeRO strategies, memory management, checkpointing).
- Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100).
- Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
- Collaborate closely with infra teams to ensure our cluster, container: environments, and hardware configurations support high-performance training.