Senior Site Reliability Engineer, Fleet InfrastructureActive$220K

Boston, Massachusetts, United States; Washington, District of Columbia, United StatesManufacturing

The opportunity

Anduril products are designed to operate in high stakes environments, in some cases with life-and-death consequences for success or failure. As such, it is critical that Anduril services are reliable and maintainable.

What you'll do

  • Build & operate a robust, high-availability core observability plane: We run Anduril’s centralized observability system for 100s of environments. Engineering teams depend on these systems daily to remain operationally responsible systems - these systems must gracefully scale to meet demand and meet strict uptime requirements.

  • Eagerly engage with our customers: We partner closely with engineering teams to understand their needs, gaps of our systems, and proactively engage with them so that our systems continually improve. We actively work to close gaps early on, and maintain a close working relationship with internal customers.

  • Enable metrics-driven engineering and seamless investigation of production: incidents — the ultimate mandate of the Observability team is to enable our organization to make data-driven decisions in real-time. This also extends to providing the substrate for agentic workflows for root-cause analysis, and enabling users to synthesize information across services and environments when triaging and actioning issues.

  • + years of experience as a Site Reliability Engineer / related role supporting production systems

  • Bachelors degree in Computer Science, Computer Engineering, Electrical: Engineering, or related field (or equivalent experience)

  • Familiarity with common container orchestration systems and cloud: infrastructure (Docker, Kubernetes, experience developing and deploying systems on AWS, GCP, or Azure). Experience with infrastructure components of the observability stack is a plus (ClickHouse/ClickStack, Victoria Metrics, Prometheus, Grafana, ELK stack, etc).

What they're looking for

  • + years of experience as a Site Reliability Engineer / related role supporting production systems
  • Bachelors degree in Computer Science, Computer Engineering, Electrical: Engineering, or related field (or equivalent experience)
  • Familiarity with common container orchestration systems and cloud: infrastructure (Docker, Kubernetes, experience developing and deploying systems on AWS, GCP, or Azure). Experience with infrastructure components of the observability stack is a plus (ClickHouse/ClickStack, Victoria Metrics, Prometheus, Grafana, ELK stack, etc).
  • Experience serving as part of an on-call rotation to support high-availability systems
  • User-focused engineering mindset: ability to translate user needs into technical solutions while balancing user experience with engineering constraints
  • U.S. Person status is required as this position needs to access export controlled data