Staff Software Engineer, Cloud Monitoring Service (Distributed Systems)Active$215K–$260K

The opportunity

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.

What you'll do

  • Distributed Systems Ownership: Own the architecture and evolution of large-scale telemetry pipelines, including high-volume ingestion, stream processing, time-series and log storage, and low-latency query paths. Design systems that stay correct, available, and cost-efficient as data volume grows 10x.

  • Scalable, Multi-Tenant Design: Design services that are highly scalable, durable, and fair across tenants. Solve the hard problems in this space: hot shards, high-cardinality data, noisy neighbors, backpressure, retention and compaction at scale, and graceful degradation under load.

  • Reliability and Operational Excellence: Build for operability from day one. Improve pipeline reliability and data freshness, reduce on-call burden through better system design rather than more process, and participate in a customer-facing on-call rotation, leading by example.

  • Technical Leadership: Set the technical direction for the team's distributed systems work. Drive design reviews, identify one-way door decisions early, and raise the bar on how the team scopes, builds, and operates systems.

  • Cross-Team Collaboration: Work with product, compute, networking, and platform teams to make sure observability decisions are made with full context. Represent the team's technical position in cross-org conversations.

  • Mentorship: Coach senior and mid-level engineers through design work, code: review, and incident response. Build patterns and frameworks that make the team better without requiring your direct involvement.

What they're looking for

  • Distributed Systems Depth: Deep, hands-on experience designing and operating distributed systems at scale. You have solved real problems in sharding, replication, consistency, load balancing, and concurrency, not just studied them.
  • Observability Data Experience: Experience building or operating large-scale data infrastructure such as time-series databases, log aggregation, streaming pipelines, or distributed tracing backends. Familiarity with technologies like Prometheus, VictoriaMetrics, Loki, OpenTelemetry, Kafka, Vector, or similar.
  • Technical Proficiency: Strong programming fundamentals in Go or another modern compiled language (Go strongly preferred). Comfort with Kubernetes, microservices, and CI/CD as the operating environment for everything you build.
  • Design Judgment at Staff Level: You proactively scope ambiguous problems, surface non-functional requirements without being prompted, and reason about tradeoffs in terms of customer and business outcomes, not just technical elegance.