The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Drive the full hardware development and sustaining lifecycle, including: feasibility, bring-up, validation, deployment, and ongoing production support.
Develop and maintain scripting and automation frameworks for hardware: testing, diagnostics, and continuous reliability improvements.
Lead deep troubleshooting and debugging across: PCIe (link training, topology, performance issues)
InfiniBand (fabric debugging, throughput, connectivity issues)
NVMe/storage (performance bottlenecks, firmware interactions, failure analysis)
Conduct rigorous system validation and characterization for GPU, CPU, and high-performance compute platforms.
What they're looking for
- Support E2E integration and solution testing to ensure Crusoe Cloud products: meet performance, reliability, and scalability expectations.
- Collaborate with mechanical, thermal, firmware, software, and manufacturing: teams to resolve system-level issues and enable stable production operation.
- Drive prototyping, qualification, and readiness for high-volume manufacturing: with both internal teams and external vendors.
- Identify opportunities for new hardware technologies, testing methods, and: sustainability improvements aligned with Crusoe’s long-term objectives.