The opportunity
ChatGPT is a rapidly evolving system: new capabilities ship continuously, product surfaces change quickly, and usage patterns shift week-to-week. Supporting that pace requires infrastructure that can handle real production constraints—high concurrency, unpredictable traffic…
What you'll do
You might work on one or more of the following areas (without being: restricted to any single area): Platform foundations & frameworks: Core libraries, service frameworks, and shared components that standardize system building, integration, and evolution.
Scalability & performance primitives: Patterns and infrastructure that reduce tail latency, improve throughput, and keep costs predictable as demand increases.
Reliability guardrails: Mechanisms that prevent outages by design—rate limiting, load shedding, dependency isolation, backpressure, safe fallbacks, and robust regression controls.
Developer productivity via golden paths: Paved roads for common workflows (data access patterns, service integration, request lifecycle) that are fast, safe, and easy to use.
Observability & debugging systems: Instrumentation, metrics models, and investigative tooling that turn vague symptoms (“it’s slow”) into precise, actionable diagnoses.
Safe change management: Deployment and rollout systems that support rapid iteration with confidence—progressive delivery, automated verification, and fast rollback strategies.
What they're looking for
- Interface and contract design across boundaries: Clean APIs and stable contracts that reduce coupling and enable independent evolution across a complex ecosystem.
- Build and evolve infrastructure platforms used by many engineers and services.
- Translate real-world constraints into clean abstractions: simple APIs, enforceable contracts, safe defaults.
- Drive improvements in reliability and performance through principled design,: measurement, and iterative hardening.