The opportunity
As a Senior Site Reliability Engineer in IaaS, you will help shape the next generation of Algolia’s production infrastructure.
What you'll do
Lead the design and evolution of Cloud Baseline capabilities across cloud: providers, including identity and access, networking, account structure, security, auditability, tagging, inventory, and cost visibility.
Design and automate cloud infrastructure foundations that enable a growing: fleet of production Kubernetes clusters.
Lead complex infrastructure initiatives, such as cloud-environment: standardisation, cluster lifecycle automation, upgrade strategies, or infrastructure-drift reduction.
Ensure cloud and Kubernetes foundations, lifecycle operations, and: operational guardrails are reliable and scalable enough to support large-scale workload migration without compromising customer experience.
Treat the platform as a product: define clear interfaces, reusable modules, self-service workflows, documentation, and reliable operational standards for the engineers who consume it.
Build automated guardrails for security, compliance, reliability, and safe: change management, allowing teams to move faster without weakening production protections.
What they're looking for
- Improve platform efficiency through capacity planning, rightsizing,: autoscaling, resource governance, and clear cost visibility.
- Use automation and AI-assisted engineering tools where appropriate to improve: fleet-scale analysis, infrastructure documentation, and safe, repeatable operational changes.
- Mentor engineers, share knowledge, and raise the quality of infrastructure: design and operations across the team.
- Collaborate with Infrastructure, Security, FinOps, and engineering teams to: align technical decisions and deliver high-impact platform capabilities.