The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Technical Ownership: You will define the roadmap for the VPC control plane and lead the evolution of network virtualization systems (such as OVN/OVS) that deliver VPCs, subnets, security groups, load balancing, NAT, and internet/hybrid connectivity.
Architecture & Scale: You will oversee the design of distributed control-plane services: intent APIs, state reconciliation, network policy compilation, and southbound programming of hosts and DPUs, while driving scalability well beyond current fleet size.
Production Execution: You will lead reliability engineering, convergence and API-latency benchmarking, regression prevention, and incident response, ensuring operational excellence within 3-6 month execution cycles.
People Leadership: You will mentor and grow a team of senior and staff-level distributed systems engineers, setting technical standards and fostering a high-performance culture of accountability.
Collaboration: You will partner closely with data-plane teams (XDP/eBPF, DPDK, DPU offload) and cloud product teams to deliver low-latency, highly available networking for multi-tenant GPU clusters.
Extensive Experience: At least 10+ years in distributed systems or cloud networking engineering, with 5-7+ years specifically managing senior/staff-level talent.
What they're looking for
- Technical Depth: Deep knowledge of SDN and network virtualization: overlay networking (VXLAN/Geneve), VPC constructs, routing (BGP/EVPN), and control-plane architectures such as OVN/OVS or equivalent.
- Distributed Systems Expertise: Hands-on experience building large-scale control planes: state reconciliation, consensus and consistency trade-offs, API design, and fleet-wide configuration propagation (Go, Kubernetes-style controllers, or similar).
- Reliability Focus: A strong understanding of operating multi-tenant cloud services: SLOs, convergence-time and scale benchmarking, graceful degradation, and blast-radius containment. Execution Mindset: The ability to resolve complex technical challenges in a fast-moving, execution-heavy environment.
- Experience with DPU/SmartNIC-programmed dataplanes, cloud load balancing or: NAT at scale, network security policy engines, or open-source contributions to OVN/OVS or SDN projects