The opportunity
Coordination Systems provides foundational distributed systems building blocks for internal Datadog platforms. Our services cover sharding, consensus, resource protection, configuration distribution, and much more.
What you'll do
Lead a core team of 5 engineers (distributed, with majority in NYC)
Lead ceremonies, prioritize and delegate project
Stay hands-on with the code, e.g. isolated features, small remediations, investigation follow ups
Stay actively involved in operations, incidents, root cause analysis, etc.
Constantly promote a culture of operational excellence, organizing gamedays,: conducting operational reviews, staying proactive with reliability
Strong distributed systems skills, able to understand and account for a: variety of failure modes, well-versed in end-to-end o11y, validation testing, simulation setup, etc.
What they're looking for
- Worked on platform teams before, providing critical infrastructure to internal stakeholders
- Experienced in handling significant incidents, both as a responder and follow-up owner
- Strong cross-team collaboration and stakeholder management skills
- Strong foundation in coding, experienced with performance optimizations and profiling