The opportunity
As a Site Reliability Engineer in IaaS, you will help build the next generation of Algolia’s production infrastructure.
What you'll do
Build and improve Cloud Baseline capabilities, including identity and access,: networking, security, resource inventory, tagging, and auditability.
Develop and maintain infrastructure as code and automation for cloud: environments and Kubernetes infrastructure.
Contribute to reliable, repeatable cloud and cluster lifecycle operations.
Help build self-service capabilities, reusable modules, and clear: documentation that make the safe path the easy path for platform consumers.
Reduce manual work and configuration drift through automation, testing, GitOps practices, and standardisation.
Use automation and AI-assisted engineering tools where appropriate to improve: infrastructure analysis, documentation, and safe, repeatable changes.
What they're looking for
- Improve observability, monitoring, alerting, capacity management, and operational documentation.
- Investigate production issues, participate in the on-call rotation, and turn: lessons learned into lasting improvements.
- Work with Infrastructure, Security, FinOps, and engineering teams to deliver: reliable, secure, and cost-aware platform capabilities.
- Hands-on production knowledge of AWS or GCP.