The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Performance Engineering: Platform-Wide: Own performance at the platform level. Establish consistent benchmarks across all domains, identify systemic bottlenecks before they become incidents, and drive solutions that scale. You’re the person who sees the thundering herd problem, the sharding gap, the concurrency issue — and owns it through resolution.
Resiliency & Long-Term Architecture: Ensure every platform domain is architected for long-term scale — HA, disaster recovery, graceful degradation, and fault isolation built in from the start, not bolted on after the fact. Identify one-way door decisions early and make sure the team doesn’t commit without full context.
Operational Excellence: Set the standard for how the team operates at scale. Drive down on-call noise, automate anything done more than once, and build shared patterns and frameworks that let the entire team ship faster with less friction. You’ve been on a customer-facing on-call rotation and lead by example.
Cross-Domain Technical Leadership: Float across platform domains as the team’s senior distributed systems authority. Provide oversight on existing, inflight, and planned systems. Contribute to technical decisions across Control Plane, Storage & State, Edge & Agents, Data Pipeline, and Async & Metering without needing to own any single one.
Cross-Team Influence: Serve as a senior technical voice in conversations with adjacent infrastructure teams — Compute, Managed Orchestration, Networking — ensuring platform-level decisions are made with full context. Build credibility across org boundaries to influence outcomes that affect the broader platform.
Roadmap & Product Collaboration: Engage directly with product and engineering leadership in the earliest, most ambiguous stages of scoping. Help define quarterly roadmaps, surface incremental milestones, and ensure technical decisions are tied to business outcomes.
What they're looking for
- Mentorship & Culture: Actively coach senior engineers. Introduce systems and frameworks that uplevel the entire team without requiring your direct involvement. Identify early signs of burnout, create an environment where it’s safe to fail, and set the team’s reputation for being collaborative, responsive, and excellent.
- Distributed Systems Depth: Deep, hands-on expertise designing and operating distributed systems at scale — sharding, replication, consensus, load balancing, and concurrency. You’ve resolved the hard problems, not just read about them.
- Performance Engineering: Proven track record of impact at the staff level and above. Proficiency in Go or a modern compiled language (Go strongly preferred). You benchmark, you profile, you fix — and you build the culture around doing it right.
- Operational Leadership: On-call experience on a customer-facing team is required. You don’t just respond to incidents — you eliminate the conditions that cause them. You’ve built runbooks, tuned alerts, and reduced noise systematically.