The opportunity
The Future of Computing Research team is an applied research team within the Consumer Devices group focused on developing new methods, models, and evaluation frameworks that support our vision for the future of computing. We work at the frontier of multimodal AI, helping turn…
What you'll do
Develop RLHF and post-training methods for multimodal models.
Build reward models and preference-learning pipelines for adaptive, personalized model behavior.
Design datasets, rubrics, and evaluation frameworks that capture user: preferences, contextual appropriateness, and long-term value in realistic tasks.
Run experiments on policy improvement using explicit feedback, implicit signals, and model-based grading.
Work on long-horizon evaluation problems, where model quality depends not: just on a single response but on whether behavior improves outcomes over time.
Collaborate closely with safety researchers to ensure that adaptation and: personalization remain aligned, interpretable, and bounded by clear constraints.
What they're looking for
- Prototype and iterate quickly on training recipes, reward formulations, data: pipelines, and evaluation suites for product-relevant behaviors.
- Help define how OpenAI measures success for personalized AI systems including: trust, appropriateness, and long-term user benefit.
- Have a strong background in machine learning research, with experience in: RLHF, reward modeling, preference optimization, or post-training for large models.
- Have worked on one or more of: reinforcement learning, ranking, recommender systems, personalization, memory, or human-in-the-loop evaluation.