Staff Software Engineer, InfrastructureNew

The opportunity

As a Staff Infrastructure Engineer on the Compute & Storage Platform team in Toast, you will design, build, and operate the foundational distributed systems powering the Toast platform as the company expands into further international markets and broadens its product offerings to encompass retail shopping.

What you'll do

  • Online Databases and Storage: Architect, deploy, and manage low-latency, highly available storage solutions (RDS Postgres and Aurora Postgres; DynamoDB; OpenSearch; S3) to power applications across the Toast product ecosystem.

  • High Availability & Realtime Performance: Ensure sub-second end-user latency, near-zero-downtime reliability, and high throughput across storage engines to support Toast's products across POS, phone, computer, drive-thru, and kiosk interfaces.

  • Infrastructure for AI-Driven PDLC: Automate infrastructure provisioning and compute pipelines.

  • Cost Optimization & Performance Engineering: Right-size compute/storage resources, implement auto-scaling policies, and manage AWS infrastructure costs without compromising platform reliability or latency SLAs.

  • Cloud & Infrastructure as Code: Hands-on experience building and managing scalable cloud infrastructure on AWS using Infrastructure as Code (Terraform, CloudFormation, etc.).

  • Software Development: Fluency in Python , Go , or a similar language for infrastructure automation, tooling, or backend development.

What they're looking for

  • Databases: Deep experience operating and tuning enterprise databases at: scale. This should include at least one variety of relational database (such as Postgres, Oracle, or MySQL) and at least one type of nonrelational database (such as MongoDB, DynamoDB, Opensearch).
  • Containerization: Experience working with containerized environments and orchestration tools (e.g., Docker, ECS, Kubernetes).
  • Solid understanding of Site Reliability Engineering (SRE) principles,: including monitoring, alerting, health checks, and performance troubleshooting in high-traffic environments.
  • Observability & Telemetry with Datadog, Prometheus