Senior AI Infrastructure Engineer, Physical InfrastructureActive$220K

The opportunity

CorpTech Infrastructure Engineering builds and operates the foundational infrastructure that powers Anduril at large. We give engineers, researchers, and product teams across the company a place to deploy fast, scalable infrastructure without having to become infrastructure experts themselves.

What you'll do

  • Rack, stack, cable, and bring up GPU compute (H200/B200/B300, NVL72): including physical topology, power, cooling, firmware/BIOS, and burn in validation.

  • Build and tune the interconnect fabric (NVLink, InfiniBand, RoCE, Spectrum-X): connecting hundreds of GPUs into low latency training and inference clusters.

  • Integrate high performance parallel storage (VAST, DDN, Weka) to sustain the: throughput demanded by distributed training and terabyte scale multi modal datasets across Anduril's programs.

  • Automate cluster deployment and configuration end to end, including: infrastructure as code for bring up, firmware/driver management, and fabric config, so new capacity comes online with minimal manual work.

  • Operate and extend our Kubernetes/Run:AI environment for GPU scheduling,: quota management, and multi tenant workload isolation across research and engineering teams company wide.

  • Own fleet health: monitoring, alerting, and rapid triage of hardware and network faults (bad transceivers, GPU Xid errors, NCCL/collective failures, RoCE congestion).

What they're looking for

  • Experience with NVIDIA NVL72 rack scale systems.
  • Experience supporting LLM token serving/inference infrastructure alongside training clusters.
  • Network fabric tuning experience (congestion control, adaptive routing, QoS) for RoCE/InfiniBand at scale.
  • Familiarity with GPU/network observability tooling (DCGM, fabric telemetry) and automated fault detection.
  • Experience supporting infrastructure as a shared platform serving multiple: internal customer teams with differing requirements.