The opportunity
At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements.
What you'll do
Build, operate, and improve production systems with a focus on reliability, scalability, and performance
Apply software engineering principles to automate operational tasks and reduce manual toil
Contribute to the design and implementation of systems using established SRE best practices
Help define and measure SLIs and SLOs for services you support
Improve observability through metrics, dashboards, logging, and tracing
Participate in on-call rotations and respond to production incidents with guidance and support
What they're looking for
- Assist with incident investigation and contribute to post-incident reviews and follow-up actions
- Perform basic analysis around system behavior, capacity usage, and scaling characteristics
- Identify reliability issues or operational pain points and work with teammates to address them
- Collaborate with product, platform, and security engineers to ship reliable systems