Compute Server Platform ArchitectActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Own the architecture for all server roles in Cerebras clusters, including: definitions of server types, configurations, and lifecycle strategy.

  • Define and maintain server formulas (counts and ratios per CS-3 count,: cluster size, and workload type) including capacity planning and headroom policy.

  • Specify platform configurations: CPU SKU and core strategy, our vendor roadmap (e.g., AMD, Intel, ARM), memory topology (channels, DIMM type, capacity), PCIe topology and lane budgeting, NIC selection/placement, and local NVMe policy where applicable.

  • Translate software and runtime flows into measurable hardware requirements: (CPU utilization, memory bandwidth/latency, bursty IO patterns, queueing and concurrency limits) and communicate clear guardrails back to software teams.

  • Develop performance and scaling models; validate with microbenchmarks and: workload-level experiments; identify bottlenecks and drive cross-stack fixes.

  • Define the OS, BIOS, firmware, and driver baseline for each server type;: there are other teams that follow these recommendations and apply them on our fleet.

What they're looking for

  • Stay current on emerging server technologies (CPU generations, new memory: technologies, CXL, NVMe evolutions, SmartNIC/DPU capabilities where relevant) and run proof-of-concept evaluations to determine when to adopt.
  • Lead technical vendor engagements (OEM/ODM and component vendors): influence roadmap, request platform knobs, and drive joint debugging on performance or reliability issues.
  • Define qualification and acceptance criteria (performance, stability,: operability) and partner with the Infrastructure Hardware TPM to execute qualification plans and land changes cleanly into production.
  • Support bring-up and rare deployment debugging in lab and staging: environments; drive root-cause analysis for regressions spanning firmware, drivers, OS, and runtime behavior.