The opportunity
Anthropic's Inference organization serves Claude to millions of users and enterprise customers with the speed, reliability, and efficiency that frontier AI demands. We build across GPUs, TPUs, and Trainium, and the complexity of our development environment grows with every platform we add.
What you'll do
Set technical direction for Inference Developer Productivity, owning the: architecture and roadmap for toolchains, dev environments, and CI/CD across GPU (CUDA), TPU, and Trainium platforms
Be the technical owner of accelerator toolchain management: compilers, drivers, libraries, frameworks, kept current, compatible, and well-tested so Inference engineers focus on model serving instead of environment archaeology
Design and build infrastructure for efficient accelerator usage during: development, including devbox environments, pre- and post-land validation automation, and shared tooling that reduces the cost of working across heterogeneous hardware
Define and instrument productivity metrics for the Inference org, building: the dashboards and alerting that surface regressions early (smoke tests red for extended periods, build times creeping up, toolchain breakages) and drive them to resolution
Proactively hunt down bottlenecks, toil, and friction across Inference: engineering workflows, then design and build the systems that eliminate them
Act as the technical counterpart to Anthropic's central Infrastructure org,: aligning on shared developer productivity initiatives, contributing Inference-specific requirements, and making the call on build vs. adopt
What they're looking for
- Experience with ML compiler toolchains (XLA, Triton, NeuronX) or accelerator: driver/firmware management at scale
- Background building or running shared development environments (devboxes,: remote development, ephemeral environments) for hardware-dependent workflows
- Experience with CI/CD systems at scale, particularly for workloads involving accelerator hardware
- Familiarity with Kubernetes-based development and job scheduling environments
- Prior tech lead experience on a developer productivity or platform: engineering team at a fast-growing AI/ML company