Senior Site Reliability Engineer (Capacity) - Platform InfrastructurePosted today$46K–$183K

The opportunity

Elastic, the Search AI Company, enables everyone to find the answers they need in real time, using all their data, at scale — unleashing the potential of businesses and people. The Elastic Search AI Platform, used by more than 50% of the Fortune 500, brings together the…

What you'll do

  • Assess current and future capacity requirements based on workload demands to: ensure seamless scaling of resources. Develop and maintain accurate capacity models that predict resource needs and align with business objectives. Collaborate with teams to implement proactive measures that prevent capacity shortages and bottlenecks.

  • Implement effective strategies for optimizing resource usage across our cloud: environments. Ensure that compute resources are utilized efficiently to enhance performance and support seamless scalability.

  • Analyze capacity metrics and trends to guide effective resource allocation: decisions. Develop insightful reporting tools that provide clear visibility into capacity and performance, helping to optimize our compute resources.

  • Operate an autoscaling framework that accommodates various customer workloads: seamlessly. Optimize infrastructure performance across over 60 regions in Elastic Cloud. Collaborate with development teams to implement scaling best practices effectively.

  • + years with cloud infrastructure and capacity management

  • Knowledge of performance monitoring and optimization techniques

What they're looking for

  • Understanding of cloud scaling challenges and solutions
  • Proficiency with incident investigation and troubleshooting processes
  • experience with compute auto-scaling processes and capacity reservations across the three major CSPs
  • solid software and platform engineering background