The opportunity
At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements.
What you'll do
Architect and evolve the Kubernetes platform, service mesh, networking,: storage, and CI/CD pipelines; ship golden paths and IaC modules.
Define platform SLOs; use error budgets to guide reliability vs. velocity: trade‑offs; drive incident learning and readiness reviews.
Improve developer velocity (build/deploy times, flaky tests, local dev ergonomics) with measurable results.
Lead capacity planning and commitments; build guardrails for cost, security,: and compliance with Security/FinOps partners.
Write high‑impact code, automation, and tooling; mentor across teams and: raise the bar on operational excellence
Embed AI in the developer experience—from code generation to observability: and incident response—so teams ship faster and safer by default.
What they're looking for
- Experience: 10+ years building and operating cloud platforms (compute,: networking, storage, runtimes like Kubernetes), with a track record of multi‑region HA and SLO rigor.
- Technical expertise: Deep in Kubernetes, service mesh, Terraform/IaC, CI/CD, and production observability; you ship golden paths and guardrails that lift the whole org.
- Experience with databases and storage systems, including SQL and NoSQL: databases, and object, block, or file storage platforms.
- AI tools & automation: You’ve brought AI into platform engineering—from copilot‑assisted workflows and intelligent test generation to AIOps for incident triage, anomaly detection, and runbook automation—with clear security and cost boundaries.