The opportunity
OpenAI’s API Multicloud team is responsible for extending OpenAI’s API platform into strategic cloud environments, starting with AWS . The team’s mission is to distribute OpenAI’s API broadly and safely by enabling key API technologies in AWS-native environments, in close…
What you'll do
Partner with strategic customers and internal teams to define target model: behaviors, diagnose failure modes, and translate real-world needs into training, evaluation, and system requirements.
Build and scale production ML systems for model customization, post-training,: and fine-tuning-as-a-service workflows.
Investigate whether training and customization workflows are producing the: intended outcomes, and identify changes to data, evaluation, training, or infrastructure that improve performance.
Partner with backend and infrastructure engineers to integrate ML: capabilities into AWS-native API environments.
Feed learnings from partner deployments back into the platform by proposing: and implementing improvements to post-training systems, tooling, APIs, and developer workflows.
Work closely with Research and Applied teams to bring model improvements,: training workflows, and evaluation best practices into production.
What they're looking for
- Help design systems that allow strategic partners and enterprise customers to: safely customize OpenAI models for high-value use cases.
- Debug and improve complex systems spanning model behavior, training data,: APIs, distributed infrastructure, and customer-facing product surfaces.
- Operate with high ownership in a 0→1 environment where requirements are: ambiguous, systems are evolving quickly, and reliability matters.
- Master’s or PhD in Computer Science, Machine Learning, or a related field, or equivalent practical experience.