Site Reliability Engineer (US - Pacific time)Active

The opportunity

Transparency: Everyone can read about our roadmap, how we pay (or even let go of) people, our strategy, and how we work, in our public company handbook . Internally, we share revenue, notes and slides from board meetings, and fundraising plans, so everyone has the context they need to make good decisions.

What you'll do

  • PostHog Code , the only AI devtool that understands your product, not just your codebase.

  • A built-in data warehouse , so users can query product and customer data together using custom SQL insights.

  • PostHog AI , an AI-powered analyst that answers product questions, helps: users find useful session recordings, and writes custom SQL queries.

  • Product-led . More than 450,000 organizations have installed PostHog, mostly: driven by word-of-mouth. We have intensely strong product-market fit.

  • Default alive . Revenue is growing incredibly quickly, and we're very: efficient. We raise money to push ambition and grow faster, not to keep the lights on.

  • Well-funded. We've raised more than $180m from some of the world's top: investors. We're set up for a long, ambitious journey.

What they're looking for

  • Deep hands-on experience with Kubernetes in production (EKS preferred).: You've debugged node pressure, networking issues, and deployment failures at scale (thousands of nodes)
  • Strong experience operating production infrastructure on AWS. Not just one: account, but understanding organizational boundaries, IAM, and networking between many
  • Experience automating infrastructure using Terraform or Terragrunt at scale,: including module design and state management
  • Solid understanding of Linux systems (disk, memory, networking, failure modes)
  • Experience supporting stateful systems (databases, queues, storage systems, etc.)
  • Ability to debug and reason about performance and reliability issues in production
  • You're comfortable owning systems end-to-end, including on-call responsibilities