The opportunity
OpenAI's research training infrastructure powers how our frontier models are trained and evaluated. The Simulation team sits at the intersection between the agentic harness that powers OpenAI's products and the research infrastructure where GPT-next is trained, ensuring that our…
What you'll do
Design, build, and evolve the integration between the Codex harness that: powers OpenAI's products and research training infrastructure used for training GPT-next
Build a platform for our LLMs to train and be evaluated in simulated: environments that mimic their deployment setting as closely as possible, on every axis: agentic harness, compute substrate, timing, tools, data sources, humans in the loop, and more
Own major integration surfaces end-to-end, from architecture and API design: through rollout, operations, and long-term maintenance
Build reliable execution systems that can support demanding training workloads at scale
Partner closely with research, agent, infrastructure, and platform teams to: support new training use cases and harness capabilities
Design clean, stable interfaces and workflows for highly technical internal: users who move quickly and expect strong ergonomics
What they're looking for
- Prevent one-off workarounds from becoming long-term technical debt by: establishing durable abstractions and clear ownership
- Raise the bar for correctness, reliability, operational rigor, and: engineering judgment across a critical research-facing system
- Have significant experience building and scaling backend or infrastructure systems in fast-moving environments
- Bring deep strength in API design, systems design, and engineering fundamentals