Director of Infrastructure EngineeringActive$225K–$325K

The opportunity

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform.

What you'll do

  • Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage. Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation.

  • Architect HPC & Global Networking: Oversee the design, scaling, and operation of Runpod’s global network backbone, as well as ultra-low-latency HPC cluster networks. Drive the implementation and optimization of InfiniBand and RDMA over Converged Ethernet (RoCE) to support massive, multi-node GPU training workloads.

  • Drive Storage Engine Innovation: Direct the architecture and performance tuning of highly scalable, distributed storage systems. Ensure our storage engines can deliver the massive IOPS and throughput required to keep high-end GPUs fed with data during deep learning tasks.

  • Build a High-Output Org: Hire, mentor, and grow highly technical engineering managers and senior ICs (network architects, systems engineers, SREs). Create a culture of ownership, operational excellence, and craft in a remote-first environment.

  • Translate Scale into Strategy: Partner with Program Management and Product to forecast capacity requirements, shape technical roadmaps, and convert massive scale challenges into clear technical scopes, milestones, and measurable outcomes.

  • Continuously Improve Systems & Flow: Drive measurable improvements in infrastructure reliability and delivery metrics, such as deployment frequency, MTTR (Mean Time To Recovery), infrastructure as code (IaC) coverage, and system uptime.

What they're looking for

  • Architectural Stewardship: Provide architectural oversight for bare-metal provisioning, virtualization layers, network fabrics, and storage clusters, ensuring seamless scalability without becoming a bottleneck for your teams.
  • Cross-Functional Partnership: Coordinate cleanly with product delivery and platform teams to ensure the infrastructure primitives they rely on are robust, well-documented, and highly available.
  • Engineering Leadership Experience: 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments.
  • Deep Infrastructure Expertise: 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms.