The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Data Center Scale-Out Automation: Design and drive automation for new on-prem server and site bring-up as Crusoe undergoes rapid, large-scale data center expansion — turning bring-up from a manual effort into a repeatable, optimized pipeline.
Performance & Reliability Deep Dives: Deep dive every component within and beneath our managed services to find systemic opportunities for performance and reliability gains — sharding gaps, concurrency issues, thundering herds — and own them through resolution. Establish platform-wide benchmarks and ensure every domain is architected for 10x scale with HA, graceful degradation, and fault isolation built in from the start.
Rapid Prototyping & Product Viability: Build POCs to evaluate viability of new platform capabilities — AI platforms, managed vector databases, managed relational databases — exploring the full solution space, not just the recommended path, so leadership can make build/extend/buy decisions faster.
Customer & Product-Focused Development: Deep dive AI-native customer profiles (training and inference) to identify what features and intelligent insights — both rule-based and LLM-based — our platform can offer. Partner with customer success and solution engineering to identify missing customer-facing functionality and tooling that simultaneously improves the admin and support experience.
Platform Architecture & Migrations: Investigate and scope high-leverage architectural shifts — e.g., migrating managed services between orchestration platforms (managed vs. internal Kubernetes, public cloud vs. on-prem), quantifying effort, trade-offs, and benefits before the org commits to one-way doors.
Cross-Functional Initiatives: Drive initiatives spanning multiple orgs — e.g., full service isolation for managed services, or partnering with managed Kubernetes, storage, and networking teams to define a permanent air-gapped solution for customers and what it looks like as a managed service offering.
What they're looking for
- Force Multiplication via AI & Tooling: Use AI to augment your own development and prototyping, and build tools and frameworks that multiply output across the entire org — including internal and customer-facing MCP servers — never limiting impact to a single team.
- Mentorship & Culture: Coach senior engineers and introduce systems that uplevel the team without requiring your direct involvement. Create an environment where it's safe to fail and set the team's reputation for being collaborative, responsive, and excellent.
- Distributed Systems Depth: Hands-on expertise designing and operating distributed systems at scale — sharding, replication, consensus, load balancing, concurrency. You've resolved the hard problems, not just read about them.
- Range: Comfort operating anywhere in the stack: from on-prem infrastructure bring-up to customer-facing product features — and switching contexts as priorities shift.