The opportunity
At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements.
What you'll do
Contribute to the architecture and evolution of backend services that power: product recommendations across Klaviyo experiences (email, SMS, KAgent, onsite, etc.), meeting standards for reliability, performance, and clear APIs.
Contribute to and maintain robust, large-scale data processing pipelines: (e.g., using Apache Spark or similar frameworks) that transform raw events and catalog data into high-quality features and inputs for recommendation models, ensuring data quality and lineage.
Collaborate closely with ML engineers and product stakeholders to: productionize recommendation models —defining high-level interfaces, feature contracts, and deployment patterns for batch and/or real-time inference systems.
Contribute to the development of the vector database that powers: recommendation, semantic search, and agentic use cases.
Ensure data and service observability (metrics, logging, tracing, dashboards): to facilitate recommendations that are correct, explainable, fast, and highly available for all customers.
Work with Product to break down projects into clear milestones, balancing the: need for rapid experimentation with technical soundness and long-term maintainability.
What they're looking for
- Lead data-driven decision making and A/B testing efforts —ensuring: recommendation systems are instrumented with the right metrics, and independently interpreting results to guide future product and engineering iterations.
- Participate in on-call and incident response for the systems you own, driving: major post-incident follow-ups that substantially improve the resilience and operability of our recommendation stack.
- Integrate AI into your and the team’s development workflow from the ground up: —for example, using AI to accelerate development, automate complex tests, or build smarter monitoring and debugging tools.
- Share knowledge, mentor junior engineers, and define best practices on: working with large-scale data frameworks, distributed systems, and integrating ML into production systems.