The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Ensure new server, storage and network infrastructure is properly racked, labeled, cabled, and configured.
Troubleshoot hardware and software issues in some of the world’s most advanced GPU and Networking systems.
Document and update data center layout and network topology in DCIM software
Work with supply chain & manufacturing teams to ensure timely deployment of: systems and project plans for large-scale deployments
Manage a parts depot inventory and track equipment through the: delivery-store-stage-deploy-handoff process in each of our data centers
Partner with HW Support teams to ensure data center hardware incidents with: higher level troubleshooting challenges are resolved, reported on and solutions are disseminated to the large operations organization.
What they're looking for
- Work with RMA team to ensure faulty parts are returned and replacements are ordered
- Follow installation standards and documentation for placement, labeling, and: cabling to drive consistency and discoverability across all data centers
- Have strong past experiences with critical infrastructure systems supporting: data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management
- Are Fluent in English and Spanish