Senior Production EngineerActive$220K

The opportunity

The SRE team owns reliability and infrastructure for Anduril's cloud deployments. We operate Kubernetes clusters, Terraform infrastructure, and observability platforms across 10+ production environments supporting active defense contracts.

What you'll do

  • Diagnose and fix stability vulnerabilities in core platform services that: cause cascading failures under multi-replica, multi-tenant operation

  • Implement resilience patterns (leader election, circuit breakers, failure: domain isolation) directly in service code

  • Design multi-replica support for services that currently assume single-instance operation

  • Collaborate with service owners on contract testing and upgrade validation

  • Trace cascading failures across service boundaries and drive them to root-cause fixes

  • Contribute to observability platform improvements to support service stability

What they're looking for

  • Light infrastructure work: Terraform/Kubernetes changes to support service fixes (~20% of time)