The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Serve as highest-level escalation point for complex P1/P0 incidents.
Lead cross-functional root cause investigations involving compute, networking: (IB/RDMA/RoCE), storage, and orchestration layers.
Partner with SRE, Software teams (Storage, Networking, Compute, K8) to design: systemic fixes rather than recurring workarounds.
Design and improve node validation, burn-in processes, performance baselining, and release readiness.
Influence Kubernetes architecture, workload orchestration (Slurm, Terraform), and AI/ML cluster stability.
Reduce MTTR and incident recurrence through structural improvements.
What they're looking for
- Troubleshoot NCCL, IB, GPU driver/firmware issues, distributed training failures.
- Support complex AI workloads (training + inference) with performance tuning and observability improvements.
- Act as technical advisor during high-risk customer incidents.
- Deliver executive-ready RCAs with clarity and confidence.