The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Design and operate reliable managed AI services with a focus on serving and scaling LLM workloads
Build automation and reliability tooling to support distributed AI pipelines and inference services
Define, measure, and improve SLIs/SLOs across AI workloads to ensure: performance and reliability targets are met
Collaborate with AI, platform, and infrastructure teams to optimize: large-scale training and inference clusters
Automate observability by building telemetry and performance tuning: strategies for latency-sensitive AI services
Investigate and resolve reliability issues in distributed AI systems using telemetry, logs, and profiling
What they're looking for
- Contribute to the architecture of next-generation distributed systems purpose-built for AI-first environments
- Strong software engineering background: experience building production-grade systems beyond scripting or Bash
- Demonstrated experience in distributed systems design and implementation
- Hands-on work with large language models (LLMs) or AI/ML infrastructure