The opportunity
Our mission at OpenAI is to discover and enact the path to safe, beneficial AGI. To do this, we believe that many technical breakthroughs are needed in generative modeling, reinforcement learning, large-scale optimization, active learning, and other areas.
What you'll do
Build and optimize OpenAI's inference stack for AWS Trainium.
Develop high-performance kernels for critical model operations and workloads.
Extend and improve compiler support to efficiently target Trainium hardware.
Build the systems necessary to execute and optimize the model forward pass on Trainium.
Profile workloads and identify bottlenecks across kernels, compiler-generated: code, runtime, and model execution.
Partner with inference and ML systems teams to bring new models and architectures onto Trainium.
What they're looking for
- Work across the hardware/software boundary to unlock performance and: capabilities from specialized AI accelerators.
- Own complex performance and systems problems end-to-end, from investigation through production deployment.
- + years of relevant engineering experience, ideally in ML systems, compilers,: kernels, runtimes, or performance engineering.
- Strong systems programming fundamentals and experience writing performance-critical software.