The opportunity
Algolia was built to help users deliver intuitive search experiences across websites and mobile applications. Our Search API serves thousands of customers in more than 100 countries, answering billions of queries every month.
What you'll do
Own and evolve production infrastructure supporting AI-related workloads and services at scale
Design and operate highly available Kubernetes-based platforms
Drive reliability through SLOs, observability, capacity planning and production guardrails
Lead complex production investigations and turn findings into durable architectural improvements
Improve shared infrastructure across networking, databases, service communication and compute
Build better CI/CD, progressive delivery, automation and developer experience
What they're looking for
- Drive cloud infrastructure efficiency and FinOps initiatives
- Participate in and improve on-call and incident response
- Mentor engineers and raise the technical bar for reliability and production engineering
- Strong hands-on production experience with at least one major cloud provider: GCP, AWS or Azure