Senior Production Engineer, Managed CloudActive$170K–$205K

The opportunity

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.

What you'll do

  • Design and operate reliable managed AI services with a focus on serving and scaling LLM workloads

  • Define, measure, and improve SLIs/SLOs across to ensure performance and reliability targets are met

  • Collaborate with AI, platform, and infrastructure teams to optimize: large-scale training and inference clusters

  • Automate observability by building telemetry and performance tuning strategies for latency-sensitive services

  • Investigate and resolve reliability issues in distributed AI systems using telemetry, logs, and profiling

  • Contribute to the architecture of next-generation distributed systems purpose-built for AI-first environments

What they're looking for

  • Strong software engineering background: experience building production-grade systems beyond scripting or Bash
  • Demonstrated experience in distributed systems design and implementation
  • SRE mindset and experience (whether or not under the SRE title) including: Defining and measuring SLIs/SLOs
  • Building monitoring and observability systems