The opportunity
OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models.
What you'll do
Design and implement the LLM inference runtime for frontier models running on custom silicon.
Build scheduling, continuous batching, memory management, KV-cache: management, and execution orchestration for high-performance inference.
Develop distributed execution strategies across chips, hosts, and racks,: including model partitioning, communication, and synchronization.
Optimize end-to-end latency, throughput, memory efficiency, and hardware: utilization across diverse model architectures and serving workloads.
Partner with kernel, compiler, architecture, and silicon teams to co-design: interfaces and remove performance bottlenecks across the stack.
Enable new model features, execution patterns, numerical formats, and: hardware capabilities in a reliable production runtime.
What they're looking for
- Create profiling, observability, benchmarking, and performance-modeling tools: that make runtime behavior measurable and actionable.
- Debug complex correctness, performance, and reliability issues spanning model: code, runtime software, communication layers, and hardware.
- Turn workload insights into clear requirements for future generations of silicon and system architecture.
- Have strong systems programming experience in C++, Rust, Python, or: comparable performance-oriented environments.