The opportunity
OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models.
What you'll do
Prototype and enable OpenAI's AI software stack on new, exploratory accelerator platforms.
Optimize large-scale model performance (LLMs, recommender systems,: distributed AI workloads) for diverse hardware environments.
Develop kernels, sharding mechanisms, and system scaling strategies tailored to emerging accelerators.
Collaborate on optimizations at the model code level (e.g. PyTorch) and below: to enhance performance on non-traditional hardware. Perform system-level performance modeling, debug bottlenecks, and drive end-to-end optimization.
Work with hardware teams and vendors to evaluate alternatives to existing: platforms and adapt the software stack to their architectures.
Contribute to runtime improvements, compute/communication overlapping, and: scaling efforts for frontier AI workloads.
What they're looking for
- + years of experience working on AI infrastructure, including kernels, systems, or hardware-software co-design
- Hands-on experience with accelerator platforms for AI at data center scale: (e.g., TPUs, custom silicon, exploratory architectures).
- Strong understanding of kernels, sharding, runtime systems, or distributed scaling techniques.
- Familiarity with optimizing LLMs, CNNs, or recommender models for hardware efficiency.