The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Production Reliability: Support uptime across Crusoe's global edge, backbone, data center, and GPU cluster network, directly supporting AI workloads at scale.
Operational Automation: Write and maintain Python-based tooling to reduce toil, automate common remediation workflows, and speed up mean time to resolution across the operations org.
Incident Response: Participate in high-severity network events — detection, triage, mitigation — and contribute to stakeholder communication and postmortem documentation.
Root Cause Analysis: Perform RCAs using existing and newly built tooling, and help drive remediation items to closure.
Observability: Build automation on top of Crusoe's monitoring stack (streaming telemetry, SNMP, NetFlow, Kentik, Grafana, Prometheus, ThousandEyes) to reduce manual triage and surface actionable signal faster.
Operational Standards: Maintain and improve runbooks, escalation playbooks, and SOPs, flagging opportunities for scripting or automation.
What they're looking for
- SLI/SLO Execution: Help build and maintain the automated dashboards and alerting that back reliability metrics, in partnership with Architecture and SRE teams.
- Collaboration: Work closely with Architecture and SRE teams to execute on operational excellence initiatives and improve day-to-day practices.
- + years of production network engineering experience with a focus on: operations, incident response, and reliability in large-scale or internet-scale environments.
- Strong Python and scripting proficiency: comfortable writing diagnostic tooling and automation scripts, and extending existing automation pipelines.