The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Build at scale: Contribute to and own components behind Crusoe Cloud's managed container registry, serving the images that training and inference workloads run on.
Ship end-to-end: Own features across the full lifecycle, design, implementation, testing, rollout, and operation.
Tackle hard problems: Take on the hard challenges of building a cloud provider, bringing clarity to ambiguous problems and landing simple, pragmatic designs.
Collaborating across teams: Partner with product and platform teams to align on solutions that serve customers and operations alike.
Build for reliability: Write and ship sustainable software with observability baked in, so the platform stays maintainable as it grows, issues surface early, and root causes are quick to find.
Raise the bar operationally: Spot and act on opportunities to make the team's systems easier, safer, and faster to run.
What they're looking for
- Distributed systems fundamentals: You have a solid grasp of how fault-tolerant distributed systems behave under high load, with experience contributing to such services in production.
- A performance mindset: You consider performance in everything you do, bringing hands-on experience profiling and benchmarking real systems, as well as reasoning about the tradeoffs.
- A customer-centric mindset: You keep the end-user top of mind, championing decisions that make the platform more reliable for the people who use it.
- Technical & infrastructure proficiency: You have strong fundamentals in microservices and cloud-native infrastructure technologies (Docker, Kubernetes, Terraform) alongside languages like Go, Rust, Java, or C++.