The opportunity
OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models.
What you'll do
Design and implement the low-level device runtime for OpenAI custom silicon.
Build kernel-launch scheduling, command submission, queueing, dependency tracking, and completion handling.
Manage device memory spaces, allocation, virtual-to-physical mappings, data: movement, and lifetime across concurrent workloads.
Implement synchronization primitives, events, barriers, streams, and ordering: guarantees that are correct and efficient.
Define clean interfaces between the runtime, drivers, firmware,: compiler-generated code, kernels, and higher-level execution systems.
Use event-based, cycle-accurate simulators to develop, validate, debug, and: performance-tune runtime behavior before and after silicon availability.
What they're looking for
- Diagnose concurrency, memory-ordering, deadlock, race, correctness, and: performance issues across software and hardware boundaries.
- Build tests, tracing, profiling, observability, and reproducible workloads: for runtime correctness and performance.
- Partner with architecture and silicon teams to turn workload and simulator: insights into hardware-software interface improvements.
- Have strong low-level systems programming experience in C, C++, Rust, or comparable environments.