The opportunity
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation.
What you'll do
Lead, hire, onboard, and develop a distributed engineering team working asynchronously.
Set priorities with Site Reliability Engineering, Product Engineering, and: GitLab Dedicated teams, and help the team deliver observability services iteratively.
Own the reliability, scalability, and cost of the team's metrics, logging,: alerting, and capacity planning platforms.
Reduce noisy or missing alerts and telemetry gaps, and use SLOs, error: budgets, and self-service instrumentation to help engineers maintain the health of their services.
Guide technical decisions about time-series storage, high-cardinality: metrics, log pipelines, and distributed tracing.
Participate in the Incident Manager On Call (IMOC) rotation, coordinating the: response to high-severity incidents affecting GitLab.com .
What they're looking for
- Keep the team's on-call rotation sustainable through coverage across time: zones, useful runbooks, better alerts, and follow-through on post-incident actions.
- Use AI tools and agents to support engineering workflows and incident triage,: reviewing their output while engineers retain responsibility for decisions.
- Experience leading an observability, platform engineering, or site: reliability engineering team operating at scale, including supporting people in a distributed, asynchronous environment.
- Technical knowledge of metrics systems such as Prometheus and long-term: storage, logging platforms such as Elasticsearch or cloud-native services, and alerting design.