The opportunity
Benefits include medical, dental, and vision coverage, flexible vacation, a 401(k) plan, meals on in-office days in the US and more.
What you'll do
Own the deployment and operations of SmithDB across cloud environments: including cluster lifecycle management, blue/green and rolling upgrades, and automated failover
Build and maintain the infrastructure tooling (Terraform, Kubernetes, Helm,: or equivalent) that provisions, configures, and scales SmithDB nodes
Own the Kubernetes infrastructure that runs our distributed database services: (multi-tenant, high throughput, low latency)
Build and improve deployment pipelines, rollout strategies, and infrastructure-as-code for the storage layer
Drive reliability engineering efforts: incident response, postmortems, SLOs, and disaster recovery for a system operating at massive scale
Manage capacity planning and cost efficiency: model growth, rightsize resources, and ensure SmithDB can absorb traffic spikes from our largest customers without manual intervention
What they're looking for
- Build the CI/CD pipeline for database infrastructure changes: safe, tested, and fast promotion from dev through staging to production
- Collaborate closely with SmithDB internals engineers to translate new engine: features into production-ready infrastructure changes and ensure safe, low-risk rollouts
- + years of experience in infrastructure, platform engineering, or SRE with hands-on
- Strong hands-on experience with Kubernetes and cloud infrastructure (AWS/GCP/Azure)