The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Production Reliability: Help own uptime across Crusoe's global edge, backbone, data center, and GPU cluster network, directly supporting AI workloads at scale.
Incident Response: Lead and contribute to end-to-end response for high-severity network events, including mitigation, stakeholder communication, and postmortem documentation.
Root Cause Analysis: Drive RCAs for production incidents, identify systemic issues, and author remediation plans tracked through to closure.
Observability Improvements: Contribute to and improve Crusoe's network monitoring stack using streaming telemetry, SNMP, NetFlow, and tools such as Kentik, Grafana, Prometheus, and ThousandEyes.
Operational Standards: Author and maintain runbooks, escalation playbooks, and SOPs used across the operations team.
Operational Automation: Write Python-based tooling to reduce toil, automate common remediation workflows, and accelerate mean time to resolution.
What they're looking for
- SLI/SLO Contribution: Partner with Architecture and SRE teams to define and track network reliability metrics and service level objectives backed by real-time dashboards.
- Mentorship: Provide technical guidance to Senior engineers and contribute to: a culture of operational excellence and continuous learning.
- + years of production network engineering experience with a focus on: operations, incident response, and reliability in large-scale or internet-scale environments.
- Hands-on experience with observability and monitoring tools including: streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, and ThousandEyes.