The opportunity
AI needs a new infrastructure layer. We're building it at Modal.
What you'll do
Identifying architectural changes to improve reliability and performance.
Fostering a culture of reliability across Modal’s engineering organization.
Defining and implementing operational processes such as deployments, upgrades, etc.
Operating systems like Kubernetes, Postgres, Redis, etc.
Participating in on-call rotations, and responding to production incidents.
+ years of experience writing high-quality production code.
What they're looking for
- + years of on-call experience for critical production services.
- Strong cloud skills, and deep familiarity with at least one hyperscaler cloud (AWS preferred).
- Familiarity with auto scaling, fleet management, and capacity planning at scale.
- Experience operating databases, monitoring, CI/CD, and other infrastructure, at scale