Site Reliability Engineer, AI PlatformActive

The opportunity

Algolia was built to help users deliver intuitive search experiences across websites and mobile applications. Our Search API serves thousands of customers in more than 100 countries, answering billions of queries every month.

What you'll do

  • Build and operate production infrastructure supporting AI-related workloads and services

  • Operate and improve highly available Kubernetes-based platforms

  • Improve reliability through SLOs, observability, alerting and capacity management

  • Investigate production issues and turn findings into durable fixes and improvements

  • Work across networking, databases, compute and service infrastructure

  • Improve CI/CD pipelines, deployment automation and developer experience

What they're looking for

  • Build and maintain infrastructure using Infrastructure as Code
  • Participate in on-call, incident response and operational improvements
  • Collaborate with experienced engineers across AI Platform and progressively: take ownership of broader production areas
  • Solid hands-on Kubernetes knowledge, including workloads, resource management, and production operations