The opportunity
Anthropic is investing $50 billion in American computing infrastructure, including datacenters custom-built for our workloads, and this role sits at the heart of that effort. As a Staff Engineer on the Datacenter Server Lifecycle team, you will own the end-to-end operational…
What you'll do
Build automation to support datacenter fleets at scale.
Define and own the end-to-end system lifecycle strategy: from provisioning and deployment through operation, maintenance, refresh, and decommissioning – and maintain automation and operational procedures for common lifecycle events (e.g., hardware failures, firmware upgrades, fleet rotations).
Partner closely with Infrastructure Security to design and enforce trusted: compute standards across the server lifecycle.
Work closely with our Networking team to ensure end-to-end connectivity across all sites.
Build and maintain tooling to track machine health, configuration, and: operational status across the full datacenter fleet.
Hands-on experience with server hardware, including rack deployment, cabling,: troubleshooting, and understanding failure modes at scale.
What they're looking for
- + years of experience in datacenter infrastructure management, or a closely related discipline.
- Hands-on experience with GPU or AI accelerator hardware (e.g., NVIDIA: A100/H100, Google TPUs, or AWS Trainium) and an understanding of their operational demands.
- Familiarity with modern provisioning and OS tooling such as LinuxBoot and NixOS
- Experience building or contributing to datacenter automation or fleet management platforms.
- Experience building and deploying server operating system distributions across large server fleets.
- Background in large-scale capacity planning and hardware refresh strategy,: ideally at a hyperscaler or large cloud provider.
- Experience with trusted compute and hardware security concepts such as secure: boot, TPM, hardware attestation, and firmware verification — or a strong desire to develop deep expertise in this area.