Senior Infrastructure Reliability EngineerActive$220K
The opportunity
Infrastructure Reliability Engineering (IRE) is a small but growing team responsible for the infrastructure and operations behind the core developer tools and on-prem compute platforms used across the entire engineering organization. We own the services every engineer depends on…
What you'll do
Serve as a primary owner for critical services, including on-call and knowledge-sharing across the team
Own the lifecycle of core self-hosted developer tools (e.g., RunAI, GitHub: Enterprise Server, CircleCI, JFrog Artifactory/Xray)
Design and implement automated systems for patching, backups (with validation), and upgrades
Scale infrastructure to support a fast-growing engineering org
Use Infrastructure-as-Code (Terraform) to manage environments
Operate and troubleshoot systems using Docker, Kubernetes, and cloud platforms (AWS, GCP, Azure)
What they're looking for
- Experience with RKE2 (or other bare-metal Kubernetes distro such as k3s, kubeadm, or OpenShift) and Cilium
- Prior experience with GitHub Enterprise Server, JFrog Artifactory/Xray, or CircleCI
- Experience with GitOps workflows and tooling (e.g., ArgoCD/FluxCD)
- Experience maintaining highly available, scalable internal tools
- Exposure to security best practices, compliance requirements, or auditing
- Experience supporting large, rapidly scaling engineering organizations
- Experience with monitoring and observability platforms (e.g., Datadog, Prometheus, Grafana)
- Background in SRE or hybrid SWE/DevOps roles