The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Technical Ownership: You will help define the roadmap for the VPC control plane and drive the evolution of network virtualization systems (such as OVN/OVS) that deliver VPCs, subnets, security groups, load balancing, NAT, and internet/hybrid connectivity.
Architecture & Scale: You will contribute to the design of distributed control-plane services: intent APIs, state reconciliation, network policy compilation, and southbound programming of hosts and DPUs, while driving scalability well beyond current fleet size.
Production Execution: You will support reliability engineering, convergence and API-latency benchmarking, regression prevention, and incident response, ensuring operational excellence within 3-6 month execution cycles.
People Leadership: You will mentor and grow a team of mid-level to senior distributed systems engineers, setting technical standards and fostering a high-performance culture of accountability.
Collaboration: You will partner closely with data-plane teams (XDP/eBPF, DPDK, DPU offload) and cloud product teams to deliver low-latency, highly available networking for multi-tenant GPU clusters.
Solid Experience: At least 6-8+ years in distributed systems or cloud networking engineering, with 2-4+ years specifically managing engineering talent.
What they're looking for
- Technical Depth: Strong knowledge of SDN and network virtualization: overlay networking (VXLAN/Geneve), VPC constructs, routing (BGP/EVPN), and control-plane architectures such as OVN/OVS or equivalent.
- Distributed Systems Expertise: Hands-on experience building large-scale control planes: state reconciliation, consensus and consistency trade-offs, API design, and fleet-wide configuration propagation (Go, Kubernetes-style controllers, or similar).
- Reliability Focus: A strong understanding of operating multi-tenant cloud services: SLOs, convergence-time and scale benchmarking, graceful degradation, and blast-radius containment.
- Execution Mindset: The ability to resolve complex technical challenges in a fast-moving, execution-heavy environment.