The opportunity
Training Runtime builds the distributed systems that power OpenAI's largest model training runs - most recently GPT-5. 5!
What you'll do
Design and build a unified dataset read platform for multiple current and future training frameworks.
Define dataset APIs, storage-format expectations, registration/versioning,: and migration paths that make data access reproducible and maintainable.
Build reliability into the read path, including stateful iteration, caching,: fast restart, recovery, and clear operational contracts.
Build terminal and web-based visualizers that let teams inspect text,: multimodal, and reinforcement learning data late in the pipeline, where bugs are most visible.
Write and review production code in core data loading, service, caching, and reliability paths.
Partner with teams working on training frameworks, reinforcement learning,: multimodal models, storage, runtime, and cluster infrastructure.
What they're looking for
- Have built or owned dataset, data loading, storage, or distributed training: infrastructure at large scale (e.g. torch.utils.data )
- Care equally about API design, debugging ergonomics, performance, and bit-level correctness.
- Understand the failure modes of large distributed training jobs and know how: data systems can create or prevent them.
- Have experience with stateful iterators, checkpoint/restart semantics,: caching, remote services, or high-throughput storage reads.