The opportunity
WorkOS builds modern developer tools and APIs that make it easy for companies to become Enterprise Ready. Our platform powers authentication, identity, authorization, and other critical infrastructure that developers need to securely scale their products to large organizations.
What you'll do
Design and evolve the systems, tooling, and processes that improve the reliability and performance of WorkOS
Collaborate with product and infrastructure teams to ensure services are: production-ready, observable, and resilient to failure
Define and measure SLIs/SLOs to guide reliability improvements
Write and optimize backend systems (in TypeScript) with a focus on: performance, maintainability, and graceful degradation
Improve our incident response process, lead postmortems, and drive follow-through on reliability risks
Develop internal tools and automations that make it easier to operate and scale our systems
What they're looking for
- Participate in our on-call rotation—responding to, resolving, and learning from production incidents
- Contribute to design and architecture discussions with a focus on operability and long-term sustainability
- Document systems, share learnings, and help grow a reliability-minded engineering culture
- Experience operating and scaling production systems in cloud environments (we use AWS)