The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Serve as a senior technical escalation point, troubleshooting the hardest: infrastructure and platform issues down to the hardware, driver, or kernel level when needed
Quickly and accurately distinguish between hardware failures, driver issues,: kernel-level problems, and customer workload misconfiguration, so issues get resolved correctly the first time
Proactively identify process, tooling, and documentation gaps, and go fix: them, not just wait for them to be assigned
Use AI tools effectively to build scripts, automations, or small internal: tools that close real operational gaps (no professional development background required)
Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure
Craft clear documentation of solutions and contribute to evolving support procedures
What they're looking for
- Collaborate closely with engineering teams to turn recurring customer pain points into permanent fixes
- Take escalations from peers while training and mentoring them in the process
- Participate in a rotating on-call schedule, owning major incidents and major customer issues
- Be ready to roll up your sleeves and pitch in wherever needed, especially during fast, high-volume deployments