Senior Production Engineer, Core PEActive$172K–$209K

The opportunity

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.

What you'll do

  • Collaborate with cross-functional teams to define and evolve availability: metrics for Crusoe’s cloud platform, including establishing, measuring, and improving SLIs and SLOs

  • Participate in production incident response, diagnosing and resolving service: disruptions while contributing to post-incident reviews and root cause analysis

  • Build, operate, and improve observability across Crusoe’s infrastructure: using tools such as Prometheus, Grafana, Alertmanager, and OpenTelemetry

  • Identify reliability risks, performance bottlenecks, and early indicators of: potential production issues across distributed systems

  • Develop automation and tooling that reduces operational toil, improves: recovery times, and enables self-healing infrastructure

  • Partner with compute, networking, storage, and platform teams to strengthen: service resilience and disaster recovery capabilities

What they're looking for

  • Contribute to improving operational processes, knowledge sharing, and: reliability best practices across the engineering organization
  • Continue growing technical depth through mentorship, training, and hands-on: work operating large-scale AI infrastructure
  • + years of experience in Production Engineering, SRE, or large-scale infrastructure operations
  • Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems