The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Collaborate with cross-functional teams to define and evolve availability: metrics for Crusoe’s cloud platform, including establishing, measuring, and improving SLIs and SLOs
Participate in production incident response, diagnosing and resolving service: disruptions while contributing to post-incident reviews and root cause analysis
Build, operate, and improve observability across Crusoe’s infrastructure: using tools such as Prometheus, Grafana, Alertmanager, and OpenTelemetry
Identify reliability risks, performance bottlenecks, and early indicators of: potential production issues across distributed systems
Develop automation and tooling that reduces operational toil, improves: recovery times, and enables self-healing infrastructure
Partner with compute, networking, storage, and platform teams to strengthen: service resilience and disaster recovery capabilities
What they're looking for
- Contribute to improving operational processes, knowledge sharing, and: reliability best practices across the engineering organization
- Continue growing technical depth through mentorship, training, and hands-on: work operating large-scale AI infrastructure
- + years of experience in Production Engineering, SRE, or large-scale infrastructure operations
- Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems