The opportunity
Our Platform team exists to accelerate the rest of engineering. As our product teams build out Immutable Audience, our job is to make them faster and safer, removing toil, clearing blockers, and building the paved roads and self-service tooling that let teams ship without reinventing the foundations each time.
What you'll do
Own and evolve our AWS and Kubernetes (EKS) foundations, from VPC upwards
Develop and release Infrastructure as code
Partner closely with security, product and data teams, understanding what: they're building and how the platform supports them
Define SLOs, SLIs, monitoring, alerting and incident response practices.
Set the bar for observability excellence within the organisation
Measure systems' health, scalability and performance metrics and identify areas of improvement.
What they're looking for
- Ensure our incident management processes, automation, and remediation are world-class
- Work as a deep technical expert on the services we have ownership over (e.g.,: cleanup tech debt, maintenance, new architecture patterns)
- Hands-on ownership of data reliability and observability: freshness, anomaly and schema-drift monitoring and data SLOs
- Deep AWS fundamentals: EKS/Kubernetes as a day-one runtime, Lambda/serverless, networking, multi-AZ