Member of Technical Staff - Reliability EngineeringActive$240K–$290K
The opportunity
Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed,…
What you'll do
You own the bar. You define what "reliable" means at Fireworks: SLOs, error budgets, production readiness, on-call expectations. Then you drive adoption across engineering.
You own the process and the tooling. Incident management, postmortems,: observability standards, failure testing, guardrails, and automation are yours end to end.
Every team owns the reliability of what they build. You make that ownership: practical. Structured logging and aggregation, metrics and tracing that work the same way everywhere, alerting that routes to the right owner, dashboards that answer "why is this slow."
You choose where the leverage is. You have a wide view of the platform and: the latitude to spend your time where it changes outcomes most.
Define reliability standards: SLOs, error budgets, production readiness criteria. Not written in a vacuum: you instrument the systems and read the real telemetry the numbers come from.
Own the reliability toolchain: Logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling.
What they're looking for
- Keep customer experience from falling through the cracks: Per-service reliability is necessary but not sufficient. A customer can hit a bad experience while every system sits inside its SLO. You make sure those failures get an owner and a fix.
- Own the seams: The hardest failures live between systems: retries that amplify load, timeouts that do not compose, dependencies nobody mapped. You find them before customers do and drive fixes through the teams that own them.
- Run incident management: Coordinate live production issues, run blameless postmortems, and track follow-ups to completion.
- Reduce toil: Automate repetitive operational work so growth does not turn into an unsustainable on-call load.