Senior Site Reliability EngineerNew$220K

Costa Mesa, California, United StatesManufacturing

The opportunity

We are seeking a highly skilled and mission-driven Site Reliability Engineer (SRE) to join our Mission Autonomy team. In this critical role, you will be responsible for ensuring the reliability, scalability, performance, and operational excellence of our cutting-edge autonomous systems.

What you'll do

  • Manage and expand specialized on-site infrastructure: Administer and grow on-premises developer servers, Hardware-in-the-Loop (HITL) systems, and other compute resources.

  • Design, implement, and maintain highly available, fault-tolerant, and resilient autonomous systems

  • Identify and eliminate performance bottlenecks in software and: infrastructure, ensuring low-latency, high-throughput, and real-time responsiveness for mission-critical operations.

  • Develop and implement comprehensive monitoring, logging, tracing, and: alerting solutions to provide deep insights into system health and behavior at scale.

  • Automate away manual operational tasks, from provisioning and deployment to testing and recovery.

  • Develop and implement strategies for scaling our services and infrastructure: to meet evolving mission demands, including distributed systems and edge deployments.

What they're looking for

  • Work closely with security teams to integrate best practices into our: operational processes and infrastructure, ensuring the integrity and confidentiality of our autonomous systems.
  • Create clear, concise, and comprehensive documentation, runbooks, and playbooks for operational procedures.
  • Integrate open-source, commercial, and Anduril-internal tooling to create: effective solutions for software delivery.
  • Collaborate with Anduril's Developer Platform, Networking, and Security teams: to support integration with broader Anduril systems.