Site Reliability Engineer, SpaceActive$220K
The opportunity
Anduril's Space team is dedicated to expanding our AI-powered capabilities into the final frontier, enhancing Space Domain Awareness, Space Control, and Command and Control for U. S.
What you'll do
Serve as a senior technical subject matter expert on a growing Product: Operations team supporting a live, operational DoD system, helping set the standard for how the team operates, communicates, and solves problems
Stand up, harden, and maintain classified Linux-based infrastructure across: globally distributed sites in air-gapped environments
Apply and enforce STIG compliance and system hardening standards across all: deployed environments this is not checkbox work, you will own the security posture
Build, maintain, and evolve infrastructure as code using Ansible/Puppet: automating provisioning, configuration management, and deployment pipelines
Write and maintain operational tooling, automation scripts, and workflow optimizations in Bash and Python
Architect and maintain comprehensive observability across the stack using: Splunk/ELK/Open Search logging, alerting, and tracing for distributed systems
What they're looking for
- Manage and operate containerized workloads in Kubernetes at scale in production, classified environments
- Provide real-time incident response, rapid root-cause analysis, and: resolution of complex issues spanning application code and backing infrastructure
- Collaborate cross-functionally with Software Engineering, Mission Ops, and: TPMs to validate operational readiness of new deployments, features, and integrations
- Provide 24x7 on-call support as part of a team supporting the Space Domain: Awareness mission on behalf of the United States Space Force