The opportunity
OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we partner closely…
What you'll do
Design and implement inference infrastructure for large-scale multimodal models.
Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs.
Enable experimental research workflows to transition into reliable production services.
Collaborate closely with researchers, infra teams, and product engineers to: deploy state-of-the-art capabilities.
Contribute to system-level improvements including GPU utilization, tensor: parallelism, and hardware abstraction layers.
Have experience building and scaling inference systems for LLMs or multimodal models.
What they're looking for
- Have worked with GPU-based ML workloads and understand the performance: dynamics of large models, especially with complex data like images or audio.
- Enjoy experimental, fast-evolving work and collaborating closely with research.
- Are comfortable dealing with systems that span networking, distributed: compute, and high-throughput data handling.
- Have familiarity with inference tooling like vLLM, TensorRT-LLM, or custom model parallel systems.