The opportunity
At Braze, we have found our people. We’re a genuinely approachable, exceptionally kind, and intensely passionate crew.
What you'll do
Identify and drive the transformative initiatives that change how the team: runs ML in production, whether that's replatforming our queueing and orchestration, overhauling deployment and cloud identity, or retiring a generation of infrastructure
Build and ship at high velocity. Staff at Braze is a hands-on delivery role;: you carry the most complex infrastructure initiatives yourself from design through production. Current examples include multi-region model serving fleets, the pipelines that keep hundreds of customer-specific models healthy, and the CI and deployment tooling that moves it all safely
Own the platform's technical vision and production quality bar. Set direction: for how models are trained, deployed, served, and observed; lead incident response for ML systems; and drive the reliability and cost work that keeps the platform efficient at scale
Drive initiatives that span teams. Our platform builds on shared: infrastructure, deployment tooling, and data systems owned with partner teams, and you carry the technical relationships with those teams
Raise the team's engineering quality through design review, code review, and: production readiness for ML systems, and mentor other senior engineers and data scientists
Connect technical decisions to customer and business outcomes, and represent: the team's technical perspective to product and engineering leadership
What they're looking for
- + years building and operating distributed systems in production, with depth: in deployment and operations. You have designed services for scale and reliability, owned CI/CD and infrastructure as code, and run what you built under production load
- Hands-on experience with ML workloads in production. Training pipelines,: model serving, feature systems, or ML platform tooling all count; deep modeling experience is a plus rather than a requirement
- A technical leader who has owned direction for a team, led multi-quarter: initiatives across team boundaries, and grown senior engineers, all while keeping a high personal output
- Deep working knowledge of Kubernetes and cloud infrastructure, including: identity and access management, networking, and the cost profile of what you run