The opportunity
At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements.
What you'll do
Build and operate foundational, security-critical services with a strong: emphasis on availability, scalability, latency, and fault tolerance
Apply software engineering principles to automate infrastructure, reduce: operational toil, and improve system reliability at scale
Design, implement, and evolve systems using SRE best practices
Define and refine SLIs, SLOs, and error budgets to guide engineering decisions
Improve observability, alerting, and incident response to reduce mean time to detection and recovery
Participate in on-call rotations with a focus on sustainable operations and automatic remediations
What they're looking for
- Perform quantitative analysis to understand system behavior, capacity constraints, and scaling limits
- Identify systemic risks and reliability bottlenecks and drive long-term, preventative solutions
- Collaborate closely with product, platform, and security engineers to: influence architecture early and ship reliable systems
- Mentor and pair with other engineers, helping raise the bar for reliability,: operational maturity, and engineering excellence