The opportunity
At OpenAI, we’re building safe and beneficial artificial general intelligence. We deploy our models through ChatGPT, our APIs, and other cutting-edge products.
What you'll do
Design, build, and operate OpenAI’s multi-tenant caching platform used across: inference, identity, quota, and product experiences.
Define the long-term vision and roadmap for caching as a core infra: capability, balancing performance, durability, and cost.
Collaborate with other infra teams (e.g., networking, observability,: databases) and product teams to ensure our caching platform meets their needs.
Have 5+ years of experience building and scaling distributed systems, with a: strong focus on caching, load balancing, or storage systems.
Have deep expertise with Redis, Memcached, or similar solutions, including: clustering, durability configurations, client-side connection patterns, and performance tuning.
Have production experience with Kubernetes, service meshes (e.g., Envoy), and autoscaling systems.
What they're looking for
- Think rigorously about latency, reliability, throughput, and cost in designing platform capabilities.
- Thrive in a fast-paced environment and enjoy balancing pragmatic engineering: with long-term technical excellence.