Software Engineer, Data InfrastructureNew$285K–$340K

The opportunity

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

What you'll do

  • Design, build, and operate the distributed storage system that feeds model training and evaluation.

  • Run this system multiple on Kubernetes clusters at petabyte scale.

  • Work with researchers and training-infra teams on how jobs actually read and: write data, and turn that into throughput, latency, and durability requirements

  • Work through the networking, I/O, and consistency problems of moving large: datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success

  • Strong storage fundamentals, including replication, consistency, caching, and data lifecycle management.

  • Strong coding ability. We work in Python and Go; experience in either is: enough, but you should be willing to pick up the other

What they're looking for

  • Experience running stateful systems on Kubernetes, including Persistent: Volumes, CSI drivers, and StatefulSets.
  • Hands-on experience with cloud object storage such as S3 as well as POSIX-style filesystems.
  • Experience with parallel or HPC filesystems such as Weka, VAST, or Lustre.
  • Familiarity with the data-loading and checkpointing patterns used in large-scale model training.