The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Lead cross-functional efforts to define and evolve availability metrics for: Crusoe's cloud platform, including establishing, measuring, and improving SLIs and SLOs
Drive production incident response, diagnosing and resolving service: disruptions while leading post-incident reviews and root cause analysis
Architect, operate, and improve observability across Crusoe's infrastructure: using tools such as Prometheus, Grafana, Alertmanager, and OpenTelemetry
Identify reliability risks, performance bottlenecks, and early indicators of: potential production issues across distributed systems
Design and develop automation and tooling that reduces operational toil,: improves recovery times, and enables self-healing infrastructure
Partner with compute, networking, storage, and platform teams to strengthen: service resilience and disaster recovery capabilities
What they're looking for
- Define and champion operational processes, knowledge sharing, and reliability: best practices across the engineering organization
- Mentor and grow junior and mid-level engineers, helping build technical depth across the team
- Bachelor's degree in Computer Science, Engineering, or a related technical: field (or equivalent practical experience)
- + years of experience in Production Engineering, SRE, or large-scale infrastructure operations