Software Engineer, Compute InfrastructureActive$230K–$405K

The opportunity

Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software,…

What you'll do

  • Compute Foundations: Build the low-level platform primitives that make heterogeneous hardware, providers, and data centers repeatable, automatable, and operable at scale.

  • Fleet / Orchestration: Turn raw capacity into reliable, efficient clusters and scheduling systems that researchers and product teams can use with minimal friction and great experience.

  • Core Network Engineering: Build and operate the high-performance networking fabrics, protocols, and observability needed for the largest training and serving workloads.

  • Hardware Health and Observability: Detect, diagnose, remediate, and prevent hardware and fleet-health issues so usable compute stays high across providers and accelerator generations.

  • Storage: Build scalable, performant, durable storage abstractions that keep: data movement and storage access from becoming a bottleneck to research or products.

  • Agent Infrastructure: Build sandboxed execution infrastructure for agentic workloads across research and production, with strong isolation, reliability, and scale.

What they're looking for

  • Build and deeply optimize reliable system software for large-scale compute: systems that run some of the world's most demanding AI workloads
  • Design and operate infrastructure across accelerators, CPUs, NICs, switches,: networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health
  • Profile, benchmark, and optimize training workloads across compute, memory,: storage, networking, NCCL and collective communication, and cluster scheduling bottlenecks
  • Create hardware-aware automation that makes provisioning, firmware and driver: upgrades, incident response, and day-to-day operations faster and less error-prone