The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
End-to-end program delivery: Own multi-quarter release planning, dependency governance, and executive communication across the Managed Inference platform.
Complex, high-risk program management: Drive model version rollouts, inference optimization campaigns, SLA readiness for new GPU hardware, and multi-tenant capacity planning from kickoff through delivery.
Cross-functional alignment: Coordinate across Model Engineering, IaaS, Cloud Foundations, Data Center Operations, and external model providers to keep programs on track and unblocked.
Proactive risk identification: Surface risks across model serving, reliability, capacity constraints, and vendor timelines before they become program-level problems.
Execution frameworks and dashboards: Build lightweight, scalable TPM frameworks suited to Crusoe's pace; maintain real-time execution dashboards and deliver crisp, data-driven executive updates.
Phase 0 planning for model onboarding: Own pre-launch planning for model onboarding on new GPU generations, including firmware and driver readiness, CUDA and ROCm stack validation, and commissioning criteria for inference workloads.
What they're looking for
- Stakeholder leadership: Drive alignment and push back effectively across engineering, product, and operations leadership -- including highly technical stakeholders who have not previously worked with a TPM.
- + years of experience as a Technical Program Manager in fast-paced technical: environments, with a track record of owning complex programs end-to-end across engineering and product organizations.
- LLM inference and model serving knowledge: Working familiarity with batching strategies, quantization approaches, and the tradeoffs that govern latency, throughput, and cost at production scale.
- Multi-tenant systems experience: Familiarity with isolation, quota management, and SLA enforcement across concurrent workloads.