The opportunity
At Braze, we have found our people. We’re a genuinely approachable, exceptionally kind, and intensely passionate crew.
What you'll do
Architect and Operate Ingress Fleets: Configure, tune, and operate Braze's high-performance NGINX routing, proxying, and ingress controller layers that manage massive, real-time API ingestion traffic
Own and Expand Scaling Routines: Own, expand, and tune automated scaling routines for our high-throughput, API-based services, leveraging RED (Request, Error, Duration) metrics, Horizontal Pod Autoscalers (HPA), and customized scaling policies to ensure seamless traffic handling under extreme spikes
Translate Product Requirements: Partner directly with product engineering teams owning API features to translate product requirements into resilient, highly available, and scalable technology stacks
Manage SLIs, SLOs, and Error Budgets: Establish meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for API services, helping teams navigate error budgets to balance rapid feature deployment with stability
Systems Design & Capacity Planning: Conduct in-depth systems design, bottleneck profiling, and capacity planning to ensure Braze meets its strict enterprise-grade SLAs
Proactive On-Call Rotation: Participate in a PagerDuty on-call rotation, utilizing shifts not just to resolve structural alerts, but to update runbooks and proactively prevent incidents from recurring
What they're looking for
- Blameless Retrospectives: Lead root-cause analysis (RCA) and blameless retrospectives for availability and performance incidents, translating operational learnings into permanent system improvements
- Professional Experience: 5+ years of experience as a DevOps, or Site Reliability Engineer in a high-scale production environment
- NGINX Mastery: Deep hands-on experience configuring, troubleshooting, and operating high-performance NGINX proxying, routing, and ingress controller layers under heavy traffic loads
- Kubernetes Orchestration: In-depth, hands-on proficiency with Kubernetes administration, cluster networking, container orchestration, cluster scheduling, and container deployment