The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Own the architecture of Lambda's global network: GPU fabrics, data center fabrics, backbone, and internet edge — along with its security, performance, and availability
Build multi-year roadmaps for network architecture, data center fabrics, and: cloud connectivity, and keep capacity ahead of demand
Set the strategy for our internet presence: peering, transit, IX footprint, and backbone traffic engineering
Partner with security, product, and executive leaders to align the network roadmap with business goals
Own network readiness for every AI data center build including management,: site turn-up, and validation — so capacity lands on schedule
Own the network operating and capital budgets (OpEx and CapEx) and driving: cost per [GPU/port/Gbps] down as we scale
What they're looking for
- Own 24/7 global network operations: uptime, latency, and jitter targets, on-call health, incident response, and blameless postmortems that change how we build
- Establish the telemetry and monitoring architecture that lets us find and: mitigate risk proactively rather than reactively
- Drive an automation-first engineering culture built on Infrastructure-as-Code: pipelines, declarative configuration management, and version-controlled automated testing
- Own company-level goals for network security posture, network health, and traffic engineering capability