Staff Cloud Support EngineerActive$156K–$190K

The opportunity

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.

What you'll do

  • Serve as highest-level escalation point for complex P1/P0 incidents.

  • Lead cross-functional root cause investigations involving compute, networking: (IB/RDMA/RoCE), storage, and orchestration layers.

  • Partner with SRE, Software teams (Storage, Networking, Compute, K8) to design: systemic fixes rather than recurring workarounds.

  • Design and improve node validation, burn-in processes, performance baselining, and release readiness.

  • Influence Kubernetes architecture, workload orchestration (Slurm, Terraform), and AI/ML cluster stability.

  • Reduce MTTR and incident recurrence through structural improvements.

What they're looking for

  • Troubleshoot NCCL, IB, GPU driver/firmware issues, distributed training failures.
  • Support complex AI workloads (training + inference) with performance tuning and observability improvements.
  • Act as technical advisor during high-risk customer incidents.
  • Deliver executive-ready RCAs with clarity and confidence.