The opportunity
OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models.
What you'll do
Define the model CI strategy for frontier models running on OpenAI custom: silicon and integrated accelerator systems.
Build production CI/CD pipelines for functional regression, performance: benchmarking, software stability, and release qualification.
Design representative test matrices across models, configurations, hardware: generations, software components, and deployment environments.
Create reliable performance baselines, regression detection, bisect and: triage workflows, and clear ownership for failures.
Establish GitOps practices for reproducible configuration, promotion,: rollback, auditability, and environment consistency.
Shape monorepo architecture, dependency management, build and test: boundaries, change validation, and developer workflows at scale.
What they're looking for
- Develop scalable orchestration, artifact management, caching, scheduling,: observability, and capacity controls for accelerator-backed CI.
- Partner with model, compiler, kernel, runtime, firmware, validation, and: hardware teams to translate release risks into automated gates.
- Improve CI reliability, speed, debuggability, and cost efficiency while: maintaining high confidence in production software.
- Have built or operated large-scale CI/CD, developer infrastructure, test: automation, or production engineering systems.