The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
HPC Deployments: Turns bare metal into production-ready capacity: ensure firmware leveling, system burn-in to shake out early failures, and validation of server performance and correctness, through to the InfiniBand fabric and GPU clusters.
Fleet Reliability: Day-2 operations across the fleet. Keeps systems healthy and keeps as much of the fleet in service for as much of its useful life as possible.
Fleet Orchestration / Data: Owns our production source of truth system. Synchronizes data from upstream systems and holds the line on correctness and quality, because everything automated downstream depends on it.
Fleet Orchestration / Automation: Owns the workflow orchestration system which people use to safely work on fleet systems for workflows that include: locking hosts, running firmware leveling jobs, OS installs, burn-in and validation, and reporting on work in flight and its results.
Fleet Foundation: Builds the host enablement tooling systems: OS and ZTP switch provisioning, firmware management, out-of-band access, and power management.
Lead and grow a distributed team of top-talent engineers responsible for the: deployment and operation of production systems infrastructure.
What they're looking for
- Work cross-functionally to deliver projects and deployments on time, ensuring alignment across stakeholders.
- Identify opportunities for efficiency gains in the tools, processes, and: automation that teams across the organization rely on day to day.
- Give stakeholders clear visibility into project progress, risks, and outcomes.
- Participate in qualification efforts for new technologies entering our production deployments.