Staff Site Reliability EngineerPosted today$253K
The opportunity
Anduril Industries is a defense technology company with a mission to transform U. S.
What you'll do
Set the reliability architecture for CorpTech Platform’s production: environment, including observability infrastructure, deployment systems, incident management, capacity planning, and failure-domain isolation.
Design and operate the observability platform, including metrics, distributed: tracing, structured logging, alerting, and dashboarding infrastructure that gives engineering teams real production visibility at scale.
Own deployment infrastructure and release-safety mechanisms, including: progressive rollout systems, canary analysis, automated rollback, and deployment gates that let the organization ship fast without gambling on production stability.
Define and govern SLO frameworks that make reliability measurable and: actionable, creating shared language between SRE, product engineering, and leadership for making trade-off decisions.
Identify systemic reliability risks across the platform and drive: infrastructure or automation investments that eliminate entire failure classes rather than patching individual symptoms.
Establish production-readiness standards and review processes that: engineering teams adopt during design and pre-launch, embedding reliability into the development lifecycle rather than bolting it on after release.
What they're looking for
- Lead incident response for complex, multi-system failures and drive: post-incident processes that produce durable systemic improvements rather than shallow fixes.
- Develop reliability patterns for AI-enabled systems, including monitoring for: model behavior drift, non-deterministic output degradation, and graceful fallback under novel failure modes.
- Provide technical direction to SRE and infrastructure engineers, facilitate: cross-team architecture decisions, and represent reliability concerns in platform-level planning and prioritization.
- Drive capacity planning, cost optimization, and performance engineering for: the platform’s critical paths and shared infrastructure.