The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
+ years of engineering experience running production infrastructure at a: large scale, with a track record of technical leadership
Demonstrated experience leading the architecture and design of large, highly: available distributed systems with Kubernetes and GPU workloads on those clusters
Deep expertise with Kubernetes dev and production coding and support,: including setting team-wide standards and best practices
Extensive experience across GCP, Azure, AWS, OCI, and multi-cloud on-prem /: hybrid serving environments, with the ability to guide strategic infrastructure decisions
Proven ability to lead the design, deployment, support, and troubleshooting: of complex Linux-based computing environments at scale
Experience owning compute/storage/network resource and cost management at an: organisational level, including optimisation strategies
What they're looking for
- Exceptional collaboration and communication skills, with experience mentoring: engineers and leading cross-functional initiatives to build mission-critical systems
- The grit and adaptability to both solve and guide others through complex: technical challenges that evolve day to day
- Strong expertise in the computational characteristics of accelerators (GPUs,: TPUs, and/or custom accelerators), and how to leverage them to drive latency and throughput improvements at scale
- Deep knowledge of distributed systems, with experience establishing patterns: and practices across engineering teams