The opportunity
OpenAI's data and storage infrastructure spans data platforms, online databases, and file/object storage. These systems underpin data ingestion and processing, durable persistence, indexing and retrieval, and product file experiences.
What you'll do
Translate model, product and data-platform needs into precise access: patterns, consistency, durability, freshness, availability and scalability requirements. Connect memory, history, retrieval and resumable work to capability and end-to-end latency.
Partner with engineering to transform data and storage architecture into: repeatable scale units: standardized provisioning, placement, routing, data movement and readiness checks that bring storage, compute and networking online together. Tie each expansion to the workloads it can serve.
Lead cross-stack programs connecting ingestion and processing, databases and: indexes, and file/object storage. Make data ownership, schema compatibility, change-data-capture, replay and consumer-readiness contracts explicit so the full data path remains correct and usable.
Make cost and efficiency architectural inputs. Evaluate physical versus: logical footprint, index and replication amplification, redundant copies, tiering, caching and network movement against the cost of serving useful workloads.
Drive resilience and recovery programs with explicit failure scenarios and: validation. Distinguish database backup, failover and point-in-time recovery from execution/workspace save-and-restore; verify correctness, recovery time, safe resumption and isolation from live traffic.
Coordinate lifecycle correctness across files, objects, databases and data: platforms, including metadata, retention, deletion and snapshots. Incorporate privacy, access control, auditability and residency requirements into the design and consumer contracts.
What they're looking for
- Lead adoption and major migrations through compatibility checks,: representative workload testing, staged cutovers, rollback and operational handoff. Improve APIs, guardrails and self-service so new capacity and capabilities can be consumed predictably.
- Measure architecture changes through product and platform outcomes: task completion and continuity, data freshness, query/retrieval and snapshot latency, throughput, reliability and cost/efficiency. Use those results to drive durable performance and operational improvements.
- Have independently owned complex production programs in data platforms,: databases or storage infrastructure and can explain the architectural decisions, your contribution and the resulting impact.
- Have deep working knowledge of hyperscaler/cloud storage technologies, such: as Amazon S3 or Azure Blob Storage, including their performance, placement, resiliency and cost constraints.