The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Incident Response Leadership: Lead end-to-end management of high-visibility technical incidents and enterprise customer escalations, ensuring rapid restoration of services and effective communication throughout the lifecycle.
Incident Coordination: Coordinate the response effort across teams during an incident — you won't be troubleshooting the technical issue yourself, but you'll work closely with responders to determine which teams are needed and ensure everyone stays focused in the right direction.
Strategic Collaboration: Act as a critical bridge between Customer Success and internal engineering/product teams, translating frontline feedback into actionable incident response steps.
Customer Communication & Judgment: Own status page communications — determining the right level of detail and the best way to frame updates so customers get clear, useful information. This requires more judgment at a startup, where communication practices are still being defined, versus larger, more established companies.
RCA Ownership: Own end-to-end production of customer-facing root cause analyses — translating incident data into clear, structured writing for enterprise customers and executives.
Corrective Action Documentation: Ensure proper documentation of internal corrective actions coming out of incidents, driving improvements to the reliability of our products and platform.
What they're looking for
- Incident Metrics & Reporting: Maintain and report on core incident metrics (response time, resolution time, time to RCA delivery), ensuring incident data is accurate and complete.
- Knowledge Empowerment: Develop and deliver training materials, internal documentation, and knowledge base articles for teammates and customers.
- Process Innovation: Design and implement incident response strategies and self-serve support processes to scale the function as the team grows.
- On-Call Participation: Participate in on-call rotation, 10am–10pm coverage, 5 days per week, providing a reliable safety net for critical service interruptions.