The opportunity
MongoDB Atlas is the premier multi-cloud database-as-a-service built and operated by the makers of MongoDB. The Cloud Operations Engineering team at MongoDB is a worldwide team responsible for the consistent operational success of every MongoDB Atlas customer.
What you'll do
Successfully coordinate and collaborate with a global team of Cloud: Operations Engineers who are tasked with ensuring our uptime guarantees to our Atlas customer base
Help scale the worldwide Cloud Operations Engineering team with the strategic: implementation and refinement of new processes and tools
Assist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidents
Monitor and detect emerging customer-facing incidents on the Atlas platform;: assist in their proactive resolution
Automate routine monitoring and troubleshooting tasks
Diagnose live incidents, differentiate between platform issues versus usage: issues, and take the next steps toward resolution
What they're looking for
- Experience with being an on call DevOps, SRE, or Cloud Operations engineer (at least 2 years)
- Expertise with Linux system administration, configuration, troubleshooting
- Experience in monitoring, system performance data collection and analysis, and reporting
- Knowledge of database operations and concepts
- Expertise with networking technologies like DNS, TCP/IP, etc
- Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g. GCP, Azure)
- Knowledgeable about a wide range of web and internet technologies
- Capability to write small programs/scripts to solve both short-term systems problems