The opportunity
About the Team: Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers,…
What they're looking for
- Build and deeply optimize reliable system software for large-scale compute: systems that run some of the world's most demanding AI workloads
- Design and operate infrastructure across accelerators, CPUs, NICs, switches,: networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health
- Profile, benchmark, and optimize training workloads across compute, memory,: storage, networking, NCCL and collective communication, and cluster scheduling bottlenecks
- Create hardware-aware automation that makes provisioning, firmware and driver: upgrades, incident response, and day-to-day operations faster and less error-prone