The opportunity
Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software,…
What you'll do
Compute Foundations: Build the low-level platform primitives that make heterogeneous hardware, providers, and data centers repeatable, automatable, and operable at scale.
Fleet / Orchestration: Turn raw capacity into reliable, efficient clusters and scheduling systems that researchers and product teams can use with minimal friction and great experience.
Core Network Engineering: Build and operate the high-performance networking fabrics, protocols, and observability needed for the largest training and serving workloads.
Hardware Health and Observability: Detect, diagnose, remediate, and prevent hardware and fleet-health issues so usable compute stays high across providers and accelerator generations.
Storage: Build scalable, performant, durable storage abstractions that keep: data movement and storage access from becoming a bottleneck to research or products.
Agent Infrastructure: Build sandboxed execution infrastructure for agentic workloads across research and production, with strong isolation, reliability, and scale.
What they're looking for
- Build and deeply optimize reliable system software for large-scale compute: systems that run some of the world's most demanding AI workloads
- Design and operate infrastructure across accelerators, CPUs, NICs, switches,: networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health
- Profile, benchmark, and optimize training workloads across compute, memory,: storage, networking, NCCL and collective communication, and cluster scheduling bottlenecks
- Create hardware-aware automation that makes provisioning, firmware and driver: upgrades, incident response, and day-to-day operations faster and less error-prone