Principal Production EngineerActive$261K–$326K

The opportunity

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.

What you'll do

  • Own the reliability and scalability of Crusoe's cloud infrastructure: compute, storage, and networking — defining SLOs, leading incident response, and driving systemic improvements that reduce toil and raise the bar across the platform

  • Build and mature the observability and tooling layer: from network fabric telemetry and storage health to control plane instrumentation and on-call tooling — so the team can detect, diagnose, and resolve issues faster than customers notice them

  • Drive platform reliability improvements across the full cloud stack,: partnering closely with software, hardware, and network engineering teams to influence architecture decisions early, before they become operational debt

  • Act as a trusted advisor to senior leadership, bringing perspective on: observability trends and advocating for the right long-term technology investments.

  • Set the technical standards for how Crusoe's production engineering: organization builds, operates, and scales — defining on-call culture, incident frameworks, and reliability practices that grow with the company

  • Mentor senior and staff engineers, elevate the team's collective technical: depth, and be the person others seek out when the problem is genuinely hard

What they're looking for

  • + years of experience in infrastructure, networking, or production: engineering — with meaningful time at companies operating at internet scale (cloud providers, CDNs, large-scale social/media platforms, or similar)
  • Strong systems fundamentals: Linux, distributed systems, storage, compute scheduling — you understand the full stack from hardware up
  • Hands-on data center experience: you've done physical infra, understand power and thermal constraints, and can reason about reliability at the facility level, not just the server level
  • The ability to write code: not necessarily full-time, but enough to automate what shouldn't be manual, instrument what isn't observable, and build tooling your team will actually use