The opportunity
Elastic, the Search AI Company, enables everyone to find the answers they need in real time, using all their data, at scale — unleashing the potential of businesses and people. The Elastic Search AI Platform, used by more than 50% of the Fortune 500, brings together the…
What you'll do
Engineering software to automate large-scale systems: building internal tools and services, not just running scripts.
Optimizing the reliability and lifecycle of hosts across multiple cloud providers.
Strengthening our observability posture: crafting alerting and monitoring systems that drive incident prevention over incident response.
Scaling global infrastructure and evolving the infrastructure management processes to meet growing demand.
Contributing to code reviews, sharing your work, planning what we need to do: next, and both mentoring and being mentored by teammates.
Being part of a balanced SRE on-call rotation: responding to incidents, improving runbooks, participating in postmortems, and championing reliability improvements.
What they're looking for
- Production experience operating large-scale cloud compute (hundreds of hosts or more) via automated workflows.
- Deep experience with Linux systems: you are at home in the terminal debugging at the OS level.
- Experience building software with Golang. You are also comfortable reviewing: others' code and offering constructive feedback.
- Proficiency working with containerized workloads in production.