The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
Design, build, and operate the distributed storage system that feeds model training and evaluation.
Run this system multiple on Kubernetes clusters at petabyte scale.
Work with researchers and training-infra teams on how jobs actually read and: write data, and turn that into throughput, latency, and durability requirements
Work through the networking, I/O, and consistency problems of moving large: datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success
Strong storage fundamentals, including replication, consistency, caching, and data lifecycle management.
Strong coding ability. We work in Python and Go; experience in either is: enough, but you should be willing to pick up the other
What they're looking for
- Experience running stateful systems on Kubernetes, including Persistent: Volumes, CSI drivers, and StatefulSets.
- Hands-on experience with cloud object storage such as S3 as well as POSIX-style filesystems.
- Experience with parallel or HPC filesystems such as Weka, VAST, or Lustre.
- Familiarity with the data-loading and checkpointing patterns used in large-scale model training.