The opportunity
We are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s) team. You will play a crucial role in shaping the architecture, reliability, and automation of our Kubernetes-based infrastructure, which powers mission-critical workloads across our global platform.
What you'll do
Design, build, and maintain scalable control plane services, operators, and: custom Kubernetes controllers; develop automation in Go/Python for end-to-end cluster lifecycle management — provisioning, upgrades, patching, and deletion
Build GPU-aware orchestration systems, working within the platform: architecture to support GPU scheduling and resource allocation
Partner with the Network team on networking solutions for AI workloads: CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and GPUDirect
Write resilient systems that handle failure gracefully: timeouts, retries, backoff, and degraded-mode operation — across large-scale distributed environments
Develop platform services for inference: model serving infrastructure, autoscaling based on inference load, and multi-model deployment patterns
Build internal tools and CLIs that let ML/AI teams deploy and monitor their own inference services
What they're looking for
- Experience building and operating managed Kubernetes services (GKE, EKS, AKS,: or similar) or working on Kubernetes control plane components
- Hands-on experience with NVIDIA's GPU/networking ecosystem: GPU Operator, device plugins, DCGM, MIG, Network Operator, NCCL tuning, or similar
- Familiarity with HPC and traditional job schedulers (Slurm) and: Kubernetes-native batch scheduling (KAI, Volcano, Kueue)
- Familiarity with GPU, InfiniBand, RDMA, or high-performance computing on Kubernetes
- Exposure to storage architecture for AI/ML workloads
- Past contributions to CNCF projects or Kubernetes SIGs a plus