Staff Production Engineer, Core PEActive$209K–$253K

The opportunity

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.

What you'll do

  • Lead cross-functional efforts to define and evolve availability metrics for: Crusoe's cloud platform, including establishing, measuring, and improving SLIs and SLOs

  • Drive production incident response, diagnosing and resolving service: disruptions while leading post-incident reviews and root cause analysis

  • Architect, operate, and improve observability across Crusoe's infrastructure: using tools such as Prometheus, Grafana, Alertmanager, and OpenTelemetry

  • Identify reliability risks, performance bottlenecks, and early indicators of: potential production issues across distributed systems

  • Design and develop automation and tooling that reduces operational toil,: improves recovery times, and enables self-healing infrastructure

  • Partner with compute, networking, storage, and platform teams to strengthen: service resilience and disaster recovery capabilities

What they're looking for

  • Define and champion operational processes, knowledge sharing, and reliability: best practices across the engineering organization
  • Mentor and grow junior and mid-level engineers, helping build technical depth across the team
  • Bachelor's degree in Computer Science, Engineering, or a related technical: field (or equivalent practical experience)
  • + years of experience in Production Engineering, SRE, or large-scale infrastructure operations