Senior Site Reliability EngineerActive

The opportunity

At Braze, we have found our people. We’re a genuinely approachable, exceptionally kind, and intensely passionate crew.

What you'll do

  • Design and operate Braze’s MongoDB infrastructure to meet strict: enterprise-grade SLAs, with deep ownership of availability, durability, and query performance

  • Build proactive monitoring and alerting that fires on symptoms: before customers feel impact – with rich MongoDB-specific observability (oplog lag, replication health, lock contention, index hit rates, etc)

  • Lead capacity planning and sharding strategy as data volumes and query patterns evolve

  • Drive root-cause analysis on MongoDB incidents and translate findings into permanent system improvements

  • Partner with product engineering teams to review schema designs, index: strategies, and aggregation pipelines - catching scalability anti-patterns before they reach production

  • Build self-service tooling, automation, and runbooks that let engineers: interact with MongoDB safely and efficiently without needing to page the platform team

What they're looking for

  • Define and enforce connection pool sizing, write-concern defaults, and: read-preference standards across the fleet
  • Manage MongoDB cluster lifecycle (provisioning, upgrades, failovers,: decommissions) on Kubernetes using the MongoDB Enterprise Kubernetes Operator, with infrastructure defined as code via Terraform and Ansible
  • Develop and maintain automated backup, restore, and point-in-time recovery: workflows - tested regularly against real workloads
  • Contribute to internal platform tooling in Ruby and/or Go that reduces: operational toil across the SRE organization