Site Reliability Engineer, Inference InfrastructureActive

The opportunity

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

What you'll do

  • Build self-service systems that automate managing, deploying and operating services.

  • This includes our custom Kubernetes operators that support language model deployments.

  • Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems.

  • Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation.

  • Build strong relationships with internal developers and influence the: Infrastructure team’s roadmap based on their feedback.

  • Develop our team through knowledge sharing and an active review process.

What they're looking for

  • + years of engineering experience running production infrastructure at a large scale
  • Experience designing large, highly available distributed systems with: Kubernetes, and GPU workloads on those clusters
  • Experience with Kubernetes dev and production coding and support
  • Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving