The opportunity
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
What you'll do
Build self-service systems that automate managing, deploying and operating services.
This includes our custom Kubernetes operators that support language model deployments.
Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems.
Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation.
Build strong relationships with internal developers and influence the: Infrastructure team’s roadmap based on their feedback.
Develop our team through knowledge sharing and an active review process.
What they're looking for
- + years of engineering experience running production infrastructure at a large scale
- Experience designing large, highly available distributed systems with: Kubernetes, and GPU workloads on those clusters
- Experience with Kubernetes dev and production coding and support
- Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving