Senior Incident ManagerPosted today

The opportunity

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.

What you'll do

  • Incident Response Leadership: Lead end-to-end management of high-visibility technical incidents and enterprise customer escalations, ensuring rapid restoration of services and effective communication throughout the lifecycle.

  • Incident Coordination: Coordinate the response effort across teams during an incident — you won't be troubleshooting the technical issue yourself, but you'll work closely with responders to determine which teams are needed and ensure everyone stays focused in the right direction.

  • Strategic Collaboration: Act as a critical bridge between Customer Success and internal engineering/product teams, translating frontline feedback into actionable incident response steps.

  • Customer Communication & Judgment: Own status page communications — determining the right level of detail and the best way to frame updates so customers get clear, useful information. This requires more judgment at a startup, where communication practices are still being defined, versus larger, more established companies.

  • RCA Ownership: Own end-to-end production of customer-facing root cause analyses — translating incident data into clear, structured writing for enterprise customers and executives.

  • Corrective Action Documentation: Ensure proper documentation of internal corrective actions coming out of incidents, driving improvements to the reliability of our products and platform.

What they're looking for

  • Incident Metrics & Reporting: Maintain and report on core incident metrics (response time, resolution time, time to RCA delivery), ensuring incident data is accurate and complete.
  • Knowledge Empowerment: Develop and deliver training materials, internal documentation, and knowledge base articles for teammates and customers.
  • Process Innovation: Design and implement incident response strategies and self-serve support processes to scale the function as the team grows.
  • On-Call Participation: Participate in on-call rotation, 10am–10pm coverage, 5 days per week, providing a reliable safety net for critical service interruptions.