The opportunity
OpenAI is at the center of some of the highest-impact multimodal work in AI. ChatGPT serves a massive global audience, and enables diverse interactions via text, speech, and visuals.
What you'll do
Define and advance multimodal safety research for text, vision, and audio,: connecting perception and semantic understanding to safe model behavior.
Build training and evaluation methods for VLMs, including post-training,: safety evals, and interventions that help models respond safely and appropriately in varied contexts.
Collaborate closely with Personal AGI, Consumer Devices, and product/model: teams to translate research into safer ambient, embedded, and personalized multimodal experiences.
Have a track record of building or advancing multimodal models, with depth in: vision-language models, video understanding, image generation, audio, or multimodal reasoning—and fluency across both perception and language.
Understand how multimodal systems work end to end, from encoders, projection: layers, and modality fusion to cross-modal reasoning, scaling, and inference tradeoffs.
Have improved frontier model behavior through post-training, using approaches: such as SFT, RL, data curation, synthetic data, evaluation, and rigorous error analysis.
What they're looking for
- Bring strong research and engineering judgment to open-ended safety problems:: you can form testable hypotheses, design decisive experiments and evaluations, diagnose model failures, and translate findings into robust improvements.