Staff Data Center Operations EngineerActive$150K–$170K

The opportunity

Crusoe Cloud operates GPU infrastructure across six production sites globally, with a fleet that spans SuperMicro, HPE, and next-generation ODM platforms as we scale. We're looking for a Staff Data Center Operations Engineer to serve as the senior technical operations resource…

What you'll do

  • Own Tier 2/3 hardware escalations across all Crusoe sites for issues that: exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution

  • Travel to sites as needed for complex platform issues, new hardware bring-ups, and deployment support

  • Identify recurring failure patterns across sites and translate them into: platform feedback, sparing strategy inputs, or OEM improvement requests

  • Root-cause complex hardware issues: PCIe, BMC, thermal, fabric — and produce resolution documentation reusable across the SiteOps org

  • Hand off platform-level findings to the appropriate internal engineering: teams with clear, well-documented escalation packages

  • Develop and maintain deep technical relationships with Crusoe's primary: hardware partners — currently SuperMicro and HPE, with upcoming ODM’s as growing platforms — at the engineering and field escalation level

What they're looking for

  • + years in data center operations, field engineering, or OEM/ODM technical: support with hands-on GPU infrastructure experience
  • Direct hands-on experience deploying and supporting GPU platforms at scale: across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms required
  • Deep familiarity with server platform architecture and OEM escalation and RMA processes
  • Experience leading or contributing to large-scale GPU cluster bring-ups: including rack staging and production handoff
  • Demonstrated ability to build technical relationships with OEM and ODM: engineering teams and drive platform-level issue resolution
  • Experience developing SOPs, runbooks, or field troubleshooting procedures and: delivering technical training to data center technician teams
  • Strong written communication: comfortable producing escalation documentation, platform runbooks, and leadership reporting
  • Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%)