Principal Software EngineerActive$502K

The opportunity

The Observability team owns the internal platform that gives every HubSpot engineer real visibility into how their systems behave in production. We build and operate the distributed tracing, metrics, alerting, and logging infrastructure that spans hundreds of microservices,…

What you'll do

  • Observability Platform Architecture: Define the patterns and evolution of HubSpot’s core telemetry platform — distributed tracing, metrics, and structured logging — at a scale that spans hundreds of services and billions of daily events. Set the standards for how instrumentation is done across a large, polyglot engineering organization.

  • AI & Agentic Observability: Lead the technical strategy for tracing and understanding AI agents and ML-powered systems in production. Define the primitives, telemetry standards, and debugging workflows that help product engineers understand what their models and agents are doing — and build trust in those systems over time. This is greenfield and consequential work.

  • High-Cardinality, High-Throughput Systems: Architect telemetry pipelines and storage systems that handle high-cardinality data at high throughput without blowing up cost or query latency. Make principled tradeoffs between sampling, fidelity, retention, and developer ergonomics.

  • Hands-on, High-Leverage Builder: Ship production code. Lead design reviews and take high-impact initiatives end-to-end, from prototype to production system at scale. Stay close to the systems you build and be the person who can debug the hardest problems when they surface.

  • Developer Experience & Adoption: Design the instrumentation APIs and libraries that product engineers reach for, making correct observability the path of least resistance. Drive OpenTelemetry adoption across a large, polyglot codebase. Build the tooling that turns raw telemetry into actionable signal for teams operating at speed.

  • Production Intelligence & Reliability Patterns: Define patterns for SLO/SLI design, alerting philosophy, and how teams graduate from reactive to proactive incident response. Push for simplicity in a domain that wants to get complicated, and consistency where tooling can drift across a large organization.

What they're looking for

  • Technical Leadership & Influence: Partner with infrastructure, platform, and product engineering teams to understand their signal gaps and close them. Influence technical strategy alongside engineering leadership, translating observability constraints and opportunities into product and operational decisions. Mentor senior engineers and tech leads, driving thoughtful design decisions and capturing learnings from major incidents and large-scale migrations.
  • Platform-Builder Experience: Proven experience building observability or telemetry tooling for internal engineering teams, rather than simply consuming it. You understand how to architect developer platforms that serve thousands of engineers across a large organization, backed by deep operational instincts and hard-earned expertise.
  • Telemetry Systems Depth: Deep expertise navigating trade-offs in telemetry pipeline design across high-cardinality data, dynamic sampling, query latency, retention economics, and data fidelity. Strong technical fluency with OpenTelemetry, distributed tracing engines, metric backends, and large-scale log ingestion infrastructure.
  • Incident Automation & Operational Excellence: Proven track record linking telemetry signals directly to automated operational workflows. You have designed architecture for real-time telemetry triggers that power automated remediation, dynamic runbooks, or AI-assisted root-cause diagnosis across microservices environments.
Principal Software Engineer at HubSpot | Role Match