Staff Site Reliability Engineer, Federal (TS/SCI)Active

The opportunity

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era.

What you'll do

  • Design, build, and operate large-scale cloud infrastructure and production services.

  • Participate in an on-call rotation supporting highly available customer-facing systems.

  • Lead incident response efforts and drive post-incident reviews focused on systemic improvements.

  • Define, measure, and improve Service Level Indicators (SLIs), Service Level: Objectives (SLOs), and error budgets.

  • Partner with engineering teams to improve service availability, scalability, performance, and resilience.

  • Continuously improve observability through metrics, logging, tracing, dashboards, and alerting.

What they're looking for

  • Experience operating SaaS platforms serving large-scale customer workloads.
  • Experience working within Kubernetes-based microservices environments.
  • Experience supporting globally distributed production environments.
  • Experience with GitOps and ArgoCD.
  • Experience implementing AI-assisted operational tooling or automation workflows.
  • U.S. soil status: the employee must be on U.S. soil, which means the 50 states, the District of Columbia, or outlying areas of the United States, as defined in Federal Acquisition Regulation (FAR) 2.101
  • U.S. Security Clearance status: the employee must be able to obtain and maintain a U.S. security clearance (Secret or Top Secret) to the extent required by U.S. Government contracts.
  • Supporting Your Well-Being