The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Distributed Systems Ownership: Own the architecture and evolution of large-scale telemetry pipelines, including high-volume ingestion, stream processing, time-series and log storage, and low-latency query paths. Design systems that stay correct, available, and cost-efficient as data volume grows 10x.
Scalable, Multi-Tenant Design: Design services that are highly scalable, durable, and fair across tenants. Solve the hard problems in this space: hot shards, high-cardinality data, noisy neighbors, backpressure, retention and compaction at scale, and graceful degradation under load.
Reliability and Operational Excellence: Build for operability from day one. Improve pipeline reliability and data freshness, reduce on-call burden through better system design rather than more process, and participate in a customer-facing on-call rotation, leading by example.
Technical Leadership: Set the technical direction for the team's distributed systems work. Drive design reviews, identify one-way door decisions early, and raise the bar on how the team scopes, builds, and operates systems.
Cross-Team Collaboration: Work with product, compute, networking, and platform teams to make sure observability decisions are made with full context. Represent the team's technical position in cross-org conversations.
Mentorship: Coach senior and mid-level engineers through design work, code: review, and incident response. Build patterns and frameworks that make the team better without requiring your direct involvement.
What they're looking for
- Distributed Systems Depth: Deep, hands-on experience designing and operating distributed systems at scale. You have solved real problems in sharding, replication, consistency, load balancing, and concurrency, not just studied them.
- Observability Data Experience: Experience building or operating large-scale data infrastructure such as time-series databases, log aggregation, streaming pipelines, or distributed tracing backends. Familiarity with technologies like Prometheus, VictoriaMetrics, Loki, OpenTelemetry, Kafka, Vector, or similar.
- Technical Proficiency: Strong programming fundamentals in Go or another modern compiled language (Go strongly preferred). Comfort with Kubernetes, microservices, and CI/CD as the operating environment for everything you build.
- Design Judgment at Staff Level: You proactively scope ambiguous problems, surface non-functional requirements without being prompted, and reason about tradeoffs in terms of customer and business outcomes, not just technical elegance.