The opportunity
Replit enables people to build software with AI. The systems underneath that experience must support safe production changes, measurable reliability, and predictable performance as usage grows.
What you'll do
Observability. Build and operate metrics, logs, traces, and alerting: capabilities. Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements.
Incident Management. Own incident tooling and practices, coordinate: cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and repeat failures.
Load Testing. Build and maintain load/failure testing capabilities. Validate: critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners.
Performance Engineering. Lead deep engagements with internal teams on SLOs: and end-to-end performance. Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners—not just recommendations.
Stay technically engaged. Review designs and production changes, debug: difficult failure modes, and use AI coding tools—including Replit—to prototype and automate. Apply rigorous review and verification to AI-generated changes.
Build and grow a high-ownership engineering team. Coach engineers, develop: technical leaders, manage performance, and hire against agreed needs. Make distributed collaboration, mentoring, and backup coverage deliberate rather than relying on a few permanent escalation points.
What they're looking for
- Measure outcomes and close the loop. Track rollout safety, recovery time,: repeat incidents, critical-path latency/throughput, test coverage, and improvements arising from cost/capacity analysis. Agree success measures and continuing ownership with partner teams.