Staff Site Reliability EngineerPosted today$253K

Costa Mesa, California, United StatesManufacturing

The opportunity

Anduril Industries is a defense technology company with a mission to transform U. S.

What you'll do

  • Set the reliability architecture for CorpTech Platform’s production: environment, including observability infrastructure, deployment systems, incident management, capacity planning, and failure-domain isolation.

  • Design and operate the observability platform, including metrics, distributed: tracing, structured logging, alerting, and dashboarding infrastructure that gives engineering teams real production visibility at scale.

  • Own deployment infrastructure and release-safety mechanisms, including: progressive rollout systems, canary analysis, automated rollback, and deployment gates that let the organization ship fast without gambling on production stability.

  • Define and govern SLO frameworks that make reliability measurable and: actionable, creating shared language between SRE, product engineering, and leadership for making trade-off decisions.

  • Identify systemic reliability risks across the platform and drive: infrastructure or automation investments that eliminate entire failure classes rather than patching individual symptoms.

  • Establish production-readiness standards and review processes that: engineering teams adopt during design and pre-launch, embedding reliability into the development lifecycle rather than bolting it on after release.

What they're looking for

  • Lead incident response for complex, multi-system failures and drive: post-incident processes that produce durable systemic improvements rather than shallow fixes.
  • Develop reliability patterns for AI-enabled systems, including monitoring for: model behavior drift, non-deterministic output degradation, and graceful fallback under novel failure modes.
  • Provide technical direction to SRE and infrastructure engineers, facilitate: cross-team architecture decisions, and represent reliability concerns in platform-level planning and prioritization.
  • Drive capacity planning, cost optimization, and performance engineering for: the platform’s critical paths and shared infrastructure.