The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Developing and implementing deep-level diagnostics and troubleshooting of: hardware faults within GPU racks and high-density compute systems.
Developing troubleshooting and automation tooling for GPU platforms including: NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
Developing automation and AI agents for executing component-level diagnosis: and remediation for failed or degraded hardware.
In conjunction with data center operations develop innovative tooling and AI: agents for managing the critical environment.
Developing tooling for post-repair validation and testing tools such as: burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.
Own the deployment, monitoring, and operational support of developed tooling,: ensuring solutions maximize GPU fleet availability and performance to drive customer success.
What they're looking for
- Developing automation and operational tooling for facilities management power: as well as direct liquid cooling hardware systems
- Software engineering experience.
- The ability to identify a problem, rapidly develop a scalable solution and ship it.
- Ability to lean in and assist team members working on critical or complex technical initiatives.