The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
+ years of engineering experience running production infrastructure at a large scale
Experience designing large, highly available distributed systems with: Kubernetes, and GPU workloads on those clusters
Experience with Kubernetes dev and production coding and support
Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving
Experience in designing, deploying, supporting, and troubleshooting in: complex Linux-based computing environments
Experience in compute/storage/network resource and cost management
What they're looking for
- Excellent collaboration and troubleshooting skills to build mission-critical: systems, and ensure smooth operations and efficient teamwork
- The grit and adaptability to solve complex technical challenges that evolve day to day
- Familiarity with computational characteristics of accelerators (GPUs, TPUs,: and/or custom accelerators), especially how they influence latency and throughput of inference.
- Strong understanding or working experience with distributed systems.