The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Troubleshooting & Repair: Diagnose and resolve hardware failures in complex GPU-based servers (both air and liquid-cooled), ensuring minimal downtime.
Hardware Testing & Qualification: Collaborate with the Infrastructure Systems team to support burn-in/stress testing of new hardware and resolve any issues that arise. Support the qualification of new hardware.
Vendor Management: Open and manage support tickets with hardware vendors, serve as the datacenter liaison for vendor support personnel, and maintain a hardware issue tracker.
Inventory Management: Maintain an accurate spares inventory and replenish stock as needed to ensure quick repairs.
Deployment Support: Assist the Cloud Deployments team with racking and cabling servers, contributing to the efficient expansion of our infrastructure.
Documentation & Communication: Maintain detailed records of hardware issues and resolutions, and communicate effectively with internal teams and vendors.
What they're looking for
- Physical Demands: Work in a physically challenging environment (sound/vibration/thermal) and be able to lift 50 lbs.
- On-Call Support: Provide occasional after-hours support to address critical issues.
- Server Hardware Expertise: Possess significant experience diagnosing and repairing complex GPU-based servers (both air and liquid-cooled).
- Technical Proficiency: Demonstrate a deep understanding of server hardware, BMC-based manageability, BIOS settings, and firmware deployment.