Engineering Manager, GPU InfrastructureActive$330K–$400K

The opportunity

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

What you'll do

  • Hire, mentor, and grow a team of GPU infrastructure engineers , including: performance, career development, and technical guidance on hard infrastructure problems

  • Own the technical roadmap for the fleet: how we deploy, operate, and scale Kubernetes clusters, including workload scheduling, hardware fault detection, and performance

  • Partner with researchers and ML engineers so the training and inference stack: works well on new GPU architectures

  • Work with cross-functional stakeholders such as Capacity, Finance, Legal,: Security, and other infrastructure teams on planning, cost, compliance, and shared dependencies

  • Drive operational excellence: observability for GPU utilization and reliability, automation of cluster provisioning, cost optimization, and vendor relationships

  • Experience managing engineering or SRE teams, with a focus on technical: mentorship, hiring, and growth, including in remote, distributed settings

What they're looking for

  • A background running large Kubernetes compute fleets in production, including: in multi-cloud environments: multi-cluster operations, scheduling, node health at scale, and familiarity with IaC and infrastructure monitoring
  • You’ve gone deep in one of the layers that make a GPU training fleet work,: whether that’s cluster-wide operations, GPU networking, or hardware, and you’re willing to get hands-on and learn the rest
  • Experience with cost optimization and capacity planning for GPU infrastructure
  • A track record of partnering with researchers or ML engineers, and of making: data-informed tradeoffs across reliability, cost, and delivery