The opportunity
MongoDB Atlas is the premier multi-cloud database-as-a-service built and operated by the makers of MongoDB. The Cloud Operations Engineering team at MongoDB is a worldwide team responsible for the consistent operational success of every MongoDB Atlas customer.
What you'll do
Successfully coordinate with a global team of Cloud Operations Engineers who: are tasked with ensuring our uptime guarantees to our Atlas customer base
Help scale the worldwide Cloud Operations Engineering team with the strategic: implementation of new processes and tools
Assist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidents
Monitor and detect emerging customer-facing incidents on the Atlas platform;: assist in their proactive resolution
Automate routine monitoring and troubleshooting tasks
Diagnose live incidents, differentiate between platform issues versus usage: issues, and take the next steps toward resolution
What they're looking for
- Experience with being an on call DevOps, SRE, or Cloud Operations engineer (at least 2 years)
- Expertise with Linux system administration, configuration, troubleshooting
- Experience in monitoring, system performance data collection and analysis, and reporting
- Expertise with networking technologies like DNS, TCP/IP, etc
- Knowledge of database operations and concepts
- Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g. GCP, Azure)
- Capability to write small programs/scripts to solve both short-term systems problems
- A CS/CE degree or equivalent experience