The opportunity
AI needs a new infrastructure layer. We're building it at Modal.
What you'll do
+ years of experience writing high-quality production code
Experience operating fleets of physical hardware (bare metal provisioning,: BMC/IPMI, PXE or network boot, firmware) or building the control planes that manage them (the more challenges you've worked through, the better)
Strong cloud skills
Strong knowledge of low-level operating system foundations (Linux kernel,: drivers, networking, file systems, containers, etc.)
Effective at debugging across layers, from BGP flapping, Linux RPS, and vBIOS: bugs to a Python control-plane service
Willingness to step into the thick of it with our on-call rotation and respond to production incidents
What they're looking for
- Experience with GPUs and the NVIDIA software stack in production (drivers,: health monitoring, XIDs, RDMA/NVLink)
- Prior experience with Go
- Automatic remediation of unhealthy machines (power cycling, reimaging, GPU: recovery) to maximize uptime and minimize operator toil.
- Automatic integration of new CPU, GPU, and storage servers into the fleet: while managing hardware and network heterogeneity.