The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
The platform architecture itself. Own the end-to-end design: twin sources of truth, the abstraction and orchestration layer (API gateway, workflow engine, policy engine, state reconciler), and the domain services built on top. You define how the pieces fit together and evolve.
The fleet as one logical computer. A topology graph that models every site,: rack, host, GPU, switch, and cable as a traversable system of record — answering questions like "if I upgrade this spine switch, which customer workloads are affected?" in one query.
Reconciliation as the core loop. Never let the inventory system guess about: runtime state, or telemetry guess about intended design. Detect drift between what should exist and what's actually running, and drive automated remediation — a GitOps-style loop for physical infrastructure.
Policy-gated lifecycle automation. No host reaches production without passing: every gate: BOM validation, attestation, burn-in, network readiness, performance baselines. You design the policy engine that makes fleet growth a software-gated pipeline instead of a manual process.
The full hardware lifecycle as code. New Product Introduction, bare-metal: provisioning, OS and firmware imaging, dependency-aware rolling firmware upgrades with canary stages and rollback, re-imaging, repair & RMA, and re-admission — orchestrated end to end, replayable via event sourcing.
Unified observability plane. One pane that correlates GPU, networking: (InfiniBand/RoCE), storage, orchestration, and workload signals at extreme cardinality — so any engineer can diagnose and recover fast, and the platform can act on what it sees.
What they're looking for
- Site autonomy at global scale. Edge agents that keep every site operating: through network partitions and reconcile with the global graph when connectivity returns — a distributed-systems problem across dozens of sites.
- From detection to resolution, automatically. Anomaly detection, blast-radius: computation from topology, workload migration, quarantine, RMA, burn-in re-validation, and fleet re-admission — closing the loop with no human in the path for known failure patterns.
- Structured multi-team orchestration. Replace email and Slack handoffs between: hardware, network, compute, and validation teams with policy-checked work orders, SLA tracking, and validation gates built into the platform.
- Set the architecture and operating standards for the platform: you define them, you don't inherit them.