Staff+ Software Engineer, Inference VelocityActive$405K

The opportunity

Anthropic's Inference organization serves Claude to millions of users and enterprise customers with the speed, reliability, and efficiency that frontier AI demands. We build across GPUs, TPUs, and Trainium, and the complexity of our development environment grows with every platform we add.

What you'll do

  • Set technical direction for Inference Developer Productivity, owning the: architecture and roadmap for toolchains, dev environments, and CI/CD across GPU (CUDA), TPU, and Trainium platforms

  • Be the technical owner of accelerator toolchain management: compilers, drivers, libraries, frameworks, kept current, compatible, and well-tested so Inference engineers focus on model serving instead of environment archaeology

  • Design and build infrastructure for efficient accelerator usage during: development, including devbox environments, pre- and post-land validation automation, and shared tooling that reduces the cost of working across heterogeneous hardware

  • Define and instrument productivity metrics for the Inference org, building: the dashboards and alerting that surface regressions early (smoke tests red for extended periods, build times creeping up, toolchain breakages) and drive them to resolution

  • Proactively hunt down bottlenecks, toil, and friction across Inference: engineering workflows, then design and build the systems that eliminate them

  • Act as the technical counterpart to Anthropic's central Infrastructure org,: aligning on shared developer productivity initiatives, contributing Inference-specific requirements, and making the call on build vs. adopt

What they're looking for

  • Experience with ML compiler toolchains (XLA, Triton, NeuronX) or accelerator: driver/firmware management at scale
  • Background building or running shared development environments (devboxes,: remote development, ephemeral environments) for hardware-dependent workflows
  • Experience with CI/CD systems at scale, particularly for workloads involving accelerator hardware
  • Familiarity with Kubernetes-based development and job scheduling environments
  • Prior tech lead experience on a developer productivity or platform: engineering team at a fast-growing AI/ML company