The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Weekend Shift Schedule: Wednesday – Sunday, or Saturday – Wednesday (5-day work week, includes weekend coverage). On-site required for weekday shifts. Compensation reflects a 10% increase for this schedule.
Customer Support: Provide exceptional technical support to customers via Zendesk, meeting SLAs and maintaining high CSAT (95%+).
On-Call Rotation: Participate in a 24/7 on-call rotation to ensure timely resolution of critical issues.
Incident Management: Primary point of contact for incident management, focusing on initial triage, communication, and procedural rigor throughout the incident lifecycle. You'll lead response efforts, ensuring clear communication with both technical and non-technical stakeholders and acting as a customer advocate to minimize disruption.
Troubleshooting: Diagnose and resolve issues related to VMs, hardware failures, and scaling tests using CLI and internal tools.
Alert Triage and Maintenance: Manage alert triage, prepare for maintenance windows, and conduct node delivery testing.
What they're looking for
- Collaboration: Work closely with SRE, Networking, and Storage teams from initial triage to root cause analysis (RCA) delivery.
- Global Teamwork: Adhere to global team collaboration and handoff processes for ticketing and on-call procedures.
- Knowledge Sharing: Develop onboarding/training materials, knowledge base documentation, and standard operating procedures (SOPs).
- Education/Experience: Bachelor's degree in IT, Computer Science, Engineering, or a related field, or 4+ years of equivalent technical experience.