The opportunity
About the Team The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models.
What they're looking for
- Debug issues across layers: PXE/boot-loader, UEFI/BIOS, BMC, OS bring-up, NIC/network reachability, kubelet/control-plane connectivity, storage constraints, and early rack/lab realities.
- BS in CS/EE (or equivalent practical experience).
- + years of experience in systems SW development and building/operating: Linux-based infrastructure in production or pre-production environments.
- Strong, hands-on experience with: Kubernetes cluster operations (node lifecycle, bootstrap/join, debugging control-plane connectivity)