Systems EngineerActive

The opportunity

Cloudflare engineering delivers code to production at a tremendous pace, and depends on automated testing to do so without incidents. The SLO team builds and runs the internal platform and tooling that empowers other engineering teams to set up Service Level Indicators (SLIs)…

What you'll do

  • Build the Platform: Create and maintain production reliability testing infrastructure and availability reporting.

  • Define Reliability Metrics: Measure uptime metrics like correctness, availability, and latency SLIs/SLOs. Develop, document, and execute SLI/SLO plans to verify systems continue to operate as expected.

  • Collaborate Cross-Functionally: Collaborate with engineering teams to understand how their systems function and interact with other Cloudflare systems in production at a huge scale.

  • Communicate & Improve: Provide clear and concise feedback to engineering and product teams as an excellent communicator. Help drive continued improvements in the software development and reliability measurement processes.

  • Experience : Proven track record as a software engineer or similar role with: a deep understanding of developing and maintaining distributed systems.

  • System Design: Experience designing, implementing, and maintaining secure and highly-available distributed systems.

What they're looking for

  • Programming Languages: Programming experience with one of the following languages: Go, Rust, or Python.
  • Reliability Metrics: Deep understanding and hands-on experience measuring uptime metrics like correctness, availability, and latency using SLOs/SLIs.
  • Experience working with Clickhouse, Prometheus, GraphQL, and Postgres.
  • Experience working with data pipelines with a focus on reliability and scale.