The opportunity
About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users.
What you'll do
Design and iterate on agent behaviors across real-world coding tasks and long-horizon workflows.
Work closely with research to develop and run evals to measure agent: performance, regressions, failure modes, and edge cases.
Improve performance through prompting, tool-use strategies, context: construction, and model-facing experimentation.
Analyze failures in production and systematically improve robustness and reliability.
Build feedback loops and data systems that get better real-task data into evaluation and research.
Work with product teams to shape user-facing agent experiences and the interfaces the agent depends on.
What they're looking for
- Help define what “good” looks like for agents completing complex tasks end-to-end.
- Have experience building or shipping machine learning or LLM-powered products.
- Are strong in Python and comfortable with modern ML tooling.
- Have worked on model evaluation, fine-tuning, or prompt design.