The opportunity
Our mission at Duolingo is to develop the best education in the world and make it universally available. It’s a big mission, and that’s where you come in!
What you'll do
Collaborate with internal teams to identify sources of instability in: distributed systems and drive operational excellence
Support core infrastructure (i.e understand, diagnose, and debug these systems in production)
Provide system design consulting, develop software platforms/frameworks, and: conduct launch reviews and root cause analysis
Maintain and document sustainable postmortem/incident response practices
Advocate for and implement changes that improve reliability, scalability, and velocity
Reduce the burden of toil with iterative development of tooling and automation
What they're looking for
- Collaborate with engineering teams to release new features and become an authority on our services
- + years of experience within site reliability engineering/DevOps of a product with millions of users
- Experience identifying and solving issues in large-scale distributed systems
- Experience with Java, Kotlin, Python or Go