The opportunity
At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements.
What you'll do
Define enterprise scalability fitness functions (latency/throughput/error: rates) and a scorecard; align teams to SLOs and budgets.
Design/implement sharding and partitioning strategies, caching/back‑pressure,: multi‑region readiness, and high‑volume migration paths.
Build lightweight enablement: benchmarks, profiling harnesses, reproducible testbeds; pair with teams to land fixes.
Lead scalability reviews and readiness gates that accelerate—not: block—delivery; drive incident deep dives tied to systemic fixes.
Communicate clearly to execs and engineers, tying technical work to business impact and customer outcomes.
Integrate AI into scale and resiliency work—from proactive anomaly detection: to synthetic load and guided runbooks—so performance improvements stick and incidents don’t repeat.
What they're looking for
- Experience: 12+ years scaling multi‑tenant SaaS with a reputation for: removing major bottlenecks and proving impact with data.
- Technical expertise: Performance engineering, capacity planning, sharding/partitioning, caching/back‑pressure, multi‑region readiness, and high‑volume migrations; you turn hotspots into robust patterns.
- AI tools & automation: You apply AI to scale work—profiling assistance, workload modeling, synthetic traffic generation, anomaly detection, and runbook copilots—always with explicit guardrails and observability.
- Cross‑org influence: You align teams through fitness functions, scorecards, and readiness gates that accelerate—not block—delivery; you communicate tradeoffs crisply to execs and engineers.