Senior AI Infrastructure Engineer, Physical InfrastructureActive$220K
The opportunity
CorpTech Infrastructure Engineering builds and operates the foundational infrastructure that powers Anduril at large. We give engineers, researchers, and product teams across the company a place to deploy fast, scalable infrastructure without having to become infrastructure experts themselves.
What you'll do
Rack, stack, cable, and bring up GPU compute (H200/B200/B300, NVL72): including physical topology, power, cooling, firmware/BIOS, and burn in validation.
Build and tune the interconnect fabric (NVLink, InfiniBand, RoCE, Spectrum-X): connecting hundreds of GPUs into low latency training and inference clusters.
Integrate high performance parallel storage (VAST, DDN, Weka) to sustain the: throughput demanded by distributed training and terabyte scale multi modal datasets across Anduril's programs.
Automate cluster deployment and configuration end to end, including: infrastructure as code for bring up, firmware/driver management, and fabric config, so new capacity comes online with minimal manual work.
Operate and extend our Kubernetes/Run:AI environment for GPU scheduling,: quota management, and multi tenant workload isolation across research and engineering teams company wide.
Own fleet health: monitoring, alerting, and rapid triage of hardware and network faults (bad transceivers, GPU Xid errors, NCCL/collective failures, RoCE congestion).
What they're looking for
- Experience with NVIDIA NVL72 rack scale systems.
- Experience supporting LLM token serving/inference infrastructure alongside training clusters.
- Network fabric tuning experience (congestion control, adaptive routing, QoS) for RoCE/InfiniBand at scale.
- Familiarity with GPU/network observability tooling (DCGM, fabric telemetry) and automated fault detection.
- Experience supporting infrastructure as a shared platform serving multiple: internal customer teams with differing requirements.