The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Lead the response to critical (SEV-1 / SEV-2) incidents impacting AI: infrastructure, GPU clusters, networking, storage, and data center operations.
Serve as the Incident Commander during major outages, coordinating: engineering, networking, facilities, and vendor teams.
Act as the liaison between leadership and external teams during incidents /: post-incidents to provide updates and status summaries.
Establish clear incident timelines, triage actions, and resolution plans.
Own the incident response lifecycle including: Assisting Technical Triage
Resolution Post-incident review
What they're looking for
- Ensure timely and accurate communication with internal stakeholders and leadership.
- Maintain incident response documentation and operational playbooks.
- Conduct analysis on incidents and identify patterns / trends for improvement: in response and systems reliability.
- Work in an On-Call Rotation to respond to, lead, and coordinate incidents