The opportunity
At Braze, we have found our people. We’re a genuinely approachable, exceptionally kind, and intensely passionate crew.
What you'll do
Operate production MongoDB clusters running in Kubernetes across multiple regions and clouds
Monitor cluster health, performance, replication, and storage usage, and proactively address degradation
Debug and resolve availability, latency, or data consistency issues in partnership with SREs
Participate in on-call rotations as a subject matter expert for MongoDB infrastructure
Automate routine operations such as database user management, cluster: scaling, version upgrades, and backup restores
Maintain and improve internal tooling used for MongoDB observability, lifecycle management, and access control
What they're looking for
- Collaborate with Platform Engineering on Kubernetes-based deployment patterns and GitOps workflows
- Implement and maintain point-in-time recovery, daily snapshots, and cross-region backup strategies
- Regularly validate restores and participate in resilience exercises and disaster recovery planning
- Partner with Security and Compliance teams to ensure backups meet RTO/RPO targets and audit requirements