The opportunity
The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional…
What you'll do
Build a system for mining production conversations and product signals to: identify representative multimodal workflows, user needs, and failure modes.
Establish and maintain evaluations for the highest-priority multimodal: behaviors and use cases, with clear coverage, quality standards, and ownership.
Package production signals into decision-ready data and evaluations that: research teams can use to improve model behavior.
Measure whether model, prompt, configuration, and product changes produce: meaningful improvements in multimodal evaluations and user outcomes.
Close gaps between research and production environments, including system: prompts, sampling behavior, multimodal configurations, inference differences, and other sources of parity drift.
Create a repeatable process for reproducing product failures with research: partners and validating fixes in the shipped experience.
What they're looking for
- Lead multimodal capacity planning by forecasting demand, translating it into: GPU and serving needs, and managing headroom and reallocation tradeoffs for voice and image-generation workloads.
- Improve the tooling and operating processes used to plan, launch, and operate: multimodal capabilities as demand and model behavior evolve.
- Coordinate targeted multilingual data collection across research, Human Data, and external vendors.
- Drive cross-functional programs that multimodal launches depend on, including: multimodal actor recruitment and selection and voice-related product partnerships across vehicles, smart speakers, and headphone ecosystems.