The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Drive the end-to-end lifecycle of next-generation compute platforms,: including evaluation, bring-up, validation, deployment, and production readiness.
Define and execute performance characterization and validation strategies for: CPU, GPU, and accelerated computing platforms.
Conduct in-depth workload characterization studies across training and: inference - dense, MoE, long-context, and multimodal models to understand compute, memory, communication, and I/O behavior on target platforms.
Translate workload and platform insights into cluster-level tuning and: configuration recommendations: topology, parallelism strategy, scheduling, power, and software stack settings to maximize delivered performance and efficiency.
Build and maintain workload performance profiles and reference configurations: that guide how clusters are deployed, tuned, and scaled for specific model families and workload classes.
Analyze system and workload performance, identify bottlenecks, and work: across hardware and software layers to drive improvements.
What they're looking for
- Lead complex system-level debugging across compute, memory, storage,: networking, accelerators, and platform firmware.
- Partner with vendors and internal engineering teams on prototyping,: qualification, NPI, and production readiness of new technologies.
- Collaborate across hardware, firmware, networking, software, infrastructure,: reliability, and operations teams to resolve complex platform issues.
- Use data and system-level insights to influence platform architecture,: technology selection, hardware roadmaps, and long-term infrastructure strategy.