Manager, Production EngineeringActive

The opportunity

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.

What you'll do

  • Team Leadership & Founding Culture: Recruit, mentor, and establish a high-performing Production Engineering footprint in Tel Aviv, setting an uncompromising cultural standard for operational discipline and systems-first engineering.

  • Incident & On-Call Ownership: Partner with US and Dublin teams to run a follow-the-sun global on-call rotation, while championing a strict blameless post-mortem culture that targets systemic failures over human error.

  • Software-Defined Operations: Drive alert-reduction initiatives to improve fleet signal-to-noise ratios, automate routine manual workflows using modern runbook automation (e.g., Temporal), and build predictive monitoring to catch SEV1/SEV2 events before customers do.

  • Collaborative Governance: Act as the ultimate Production Gatekeeper across cross-functional compute, storage, networking, and platform teams, holding a strict line on Production Readiness Reviews and change control.

  • Strategic Reliability Engineering: Protect team bandwidth to ensure engineers spend at least 30% of their time on strategic automation, tooling, and firmware optimization primitives rather than drowning in incident response.

  • Physical-to-Digital Automation: Instill a software-first approach to physical problems, ensuring that any physical intervention occurring twice is successfully converted into a software-defined auto-remediation.

What they're looking for

  • Years of Infrastructure Experience: Minimum of 8+ years of experience working within infrastructure, SRE, or production engineering environments.
  • Engineering Leadership Track Record: Minimum of 2+ years of experience directly leading first-line engineering teams within a high-growth neocloud, hyperscaler, or large-scale distributed environment.
  • Non-Negotiable Coding Proficiency: Strong, hands-on software engineering fundamentals in Go, Python, C++, or a comparable systems language to build automation rather than scale through headcount.
  • Distributed Systems Depth: Expert-level command of Linux internals, container orchestration at scale, and root-cause analysis across complex physical-to-virtual boundaries.