Staff Network Production Engineer, OperationsNew$195K–$235K

The opportunity

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.

What you'll do

  • Production Reliability: Help own uptime across Crusoe's global edge, backbone, data center, and GPU cluster network, directly supporting AI workloads at scale.

  • Incident Response: Lead and contribute to end-to-end response for high-severity network events, including mitigation, stakeholder communication, and postmortem documentation.

  • Root Cause Analysis: Drive RCAs for production incidents, identify systemic issues, and author remediation plans tracked through to closure.

  • Observability Improvements: Contribute to and improve Crusoe's network monitoring stack using streaming telemetry, SNMP, NetFlow, and tools such as Kentik, Grafana, Prometheus, and ThousandEyes.

  • Operational Standards: Author and maintain runbooks, escalation playbooks, and SOPs used across the operations team.

  • Operational Automation: Write Python-based tooling to reduce toil, automate common remediation workflows, and accelerate mean time to resolution.

What they're looking for

  • SLI/SLO Contribution: Partner with Architecture and SRE teams to define and track network reliability metrics and service level objectives backed by real-time dashboards.
  • Mentorship: Provide technical guidance to Senior engineers and contribute to: a culture of operational excellence and continuous learning.
  • + years of production network engineering experience with a focus on: operations, incident response, and reliability in large-scale or internet-scale environments.
  • Hands-on experience with observability and monitoring tools including: streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, and ThousandEyes.