The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Track, log, and manage all quality issues arising in the data center during: deployment and production environment
Perform root cause analysis (RCA) for every failure (hardware, software, process)
Analyze production system metrics and quality data to detect trends, anomalies, or weak points
Improve turnaround time (TAT) for Return Merchandise Authorization (RMA) processes
Design, monitor, and drive corrective and preventive actions (CAPA)
Implement and verify containment actions to keep systems operational until permanent fixes are applied.
What they're looking for
- Collaborate with operations, hardware, engineering, supply chain, and vendors to resolve quality issues
- Capture and upload failure analysis (FA) reports and related data into Quality Management Systems (QMS)
- Verify quality of spares (incoming and outgoing) to avoid repeat failures.
- Define and track quality KPIs / SLAs and report on quality performance to leadership