The opportunity
At Braze, we have found our people. We’re a genuinely approachable, exceptionally kind, and intensely passionate crew.
What you'll do
Design and operate Braze’s MongoDB infrastructure to meet strict: enterprise-grade SLAs, with deep ownership of availability, durability, and query performance
Build proactive monitoring and alerting that fires on symptoms: before customers feel impact – with rich MongoDB-specific observability (oplog lag, replication health, lock contention, index hit rates, etc)
Lead capacity planning and sharding strategy as data volumes and query patterns evolve
Drive root-cause analysis on MongoDB incidents and translate findings into permanent system improvements
Partner with product engineering teams to review schema designs, index: strategies, and aggregation pipelines - catching scalability anti-patterns before they reach production
Build self-service tooling, automation, and runbooks that let engineers: interact with MongoDB safely and efficiently without needing to page the platform team
What they're looking for
- Define and enforce connection pool sizing, write-concern defaults, and: read-preference standards across the fleet
- Manage MongoDB cluster lifecycle (provisioning, upgrades, failovers,: decommissions) on Kubernetes using the MongoDB Enterprise Kubernetes Operator, with infrastructure defined as code via Terraform and Ansible
- Develop and maintain automated backup, restore, and point-in-time recovery: workflows - tested regularly against real workloads
- Contribute to internal platform tooling in Ruby and/or Go that reduces: operational toil across the SRE organization