Senior Infrastructure Reliability EngineerActive$220K

The opportunity

Infrastructure Reliability Engineering (IRE) is a small but growing team responsible for the infrastructure and operations behind the core developer tools and on-prem compute platforms used across the entire engineering organization. We own the services every engineer depends on…

What you'll do

  • Serve as a primary owner for critical services, including on-call and knowledge-sharing across the team

  • Own the lifecycle of core self-hosted developer tools (e.g., RunAI, GitHub: Enterprise Server, CircleCI, JFrog Artifactory/Xray)

  • Design and implement automated systems for patching, backups (with validation), and upgrades

  • Scale infrastructure to support a fast-growing engineering org

  • Use Infrastructure-as-Code (Terraform) to manage environments

  • Operate and troubleshoot systems using Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure)

What they're looking for

  • Experience with RKE2 (or other bare-metal Kubernetes distro such as k3s, kubeadm, or OpenShift) and Cilium
  • Prior experience with GitHub Enterprise Server, JFrog Artifactory/Xray, or CircleCI
  • Experience with GitOps workflows and tooling (e.g., ArgoCD/FluxCD)
  • Experience maintaining highly available, scalable internal tools
  • Exposure to security best practices, compliance requirements, or auditing
  • Experience supporting large, rapidly scaling engineering organizations
  • Experience with monitoring and observability platforms (e.g., Datadog, Prometheus, Grafana)
  • Background in SRE or hybrid SWE/DevOps roles