Staff Reliability EngineerNew

The opportunity

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era.

What you'll do

  • Design and Own the resilience, health and availability of our entire global: corporate network domain, managing operational responsibilities such as responding to alerts, monitoring health indicators, and executing reliability projects to ensure an "Always Secure. Always On." environment.

  • Drive Strategic Reduction of systemic toil and technical debt across multiple: teams by introducing process efficiencies, automating network operations, and building scalable self-service operational tooling.

  • Collaborate and Influence closely with cross-functional: stakeholders—including Business Technology, Workplace, Security, and executive leaders—challenging assumptions with grace and cascading relevant information to project teams.

  • Make Critical Decisions and lead the resolution of complex network operations: issues and alerts from a systems perspective, anticipating potential business challenges, monitoring leading indicators, and preventing future outages.

  • Foster Learning and Talent by defining success for the whole team,: cultivating an open and transparent environment, and actively mentoring team members through the P4 level to develop their skills and operational engineering best practices.

  • Typically requires 8+ years of related experience in a professional role with: a Bachelor’s degree; or 6+ years with a Master’s degree; or 3+ years with a PhD; or equivalent experience.

What they're looking for

  • Deep expertise in AWS Networking and Palo Alto Networks solutions as core required technical competencies.
  • Comprehensive operational experience in Distributed Systems & Networking: fundamentals, including quick incident response to alerts, monitoring system health, and managing protocols such as WiFi, DNS, DHCP, VLANs, VPN, ACLs, Routing, and Firewall Policies.
  • Strong proficiency in core technical skills: Cloud Platforms, IaC (e.g., Terraform/Ansible), Observability tools (e.g., Prometheus/Grafana), Programming (Python/Go), and Service Reliability Management (SLOs/SLIs) for large-scale enterprise environments.
  • Proven track record of managing operational availability, delivering: multi-quarter objectives, and executing technical projects within defined budgets and strategic VMTs (Vision, Mission, Targets).