Senior Staff AI Accelerator Performance ArchitectActive$175K–$275K

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Own and evolve performance models and modeling methodologies for: next-generation accelerator and system architectures.

  • Build and extend analytical, simulation-based or trace-driven models across: workloads, architectural features and product generations.

  • Analyze important AI workloads, from individual kernels through end-to-end: inference and training execution, to determine where time, bandwidth, compute and capacity are spent.

  • Identify hardware and software bottlenecks and quantify opportunities to: improve latency, throughput, utilization and energy efficiency.

  • Evaluate proposed architectural features and determine their expected: performance return across representative workloads.

  • Study how models and kernels map onto the underlying compute, memory and communication architecture.

What they're looking for

  • + years of experience in performance analysis, performance modeling or: architecture exploration for CPUs, GPUs, AI accelerators or other high-performance computing systems.
  • Strong understanding of hardware architecture developed through hardware,: compiler, kernel, runtime or system-performance work.
  • Experience developing analytical, simulation-based or trace-driven: performance models using Python, C++ or similar environments.
  • Solid understanding of processor architecture, memory systems, interconnects,: parallel execution and hardware resource constraints.
  • Ability to move between kernel-level behavior and end-to-end application or system performance.
  • Experience profiling workloads, forming performance hypotheses and validating them with quantitative evidence.
  • Understanding of how software mapping and programmability affect realized hardware performance.
  • Ability to communicate modeling assumptions, uncertainty, bottlenecks and recommendations clearly.