Hardware Analytics EngineerActive$225K

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Design and optimize scalable data pipeline architectures for multi-terabyte: hardware telemetry, reliability analytics, and performance optimization.

  • Architect, develop, and optimize hyperscale data pipeline frameworks and ETL: processes to aggregate, process, and analyze multi-terabyte hardware performance and telemetry streams, including utilization, power, thermal, acoustic, and reliability metrics across heterogeneous compute, storage, and AI server platforms, ensuring hardware performance compliance and operational reliability.

  • Design and implement hardware performance analysis and anomaly detection: systems using Python, SQL, Tableau, Hive, and Spark to forecast hardware failure curves, identify performance bottlenecks, and generate prescriptive recommendations for hardware and system optimization.

  • Lead hardware characterization experiments and thermal/cooling A/B studies to: evaluate operational envelopes, delivering validated strategies that reduce carbon footprint, improve water usage efficiency, and maintain or enhance system reliability.

  • Engineer telemetry ingestion, monitoring, and visualization systems to: provide real-time, high-fidelity hardware health data to hardware, firmware, and datacenter operations teams, enabling data-driven decision-making at scale.

  • Define, operationalize, and maintain custom efficiency and reliability: metrics; perform root cause analysis of systemic failures using large-scale statistical and machine learning methods; and deploy solutions that improve platform scalability, energy efficiency, and sustainability.

What they're looking for

  • Collaborate with cross-functional engineering teams to troubleshoot complex: failures, isolate defective components, and implement systemic fixes across CPU, GPU, DRAM, PCIe, networking, and storage subsystems.
  • Support the evolution and optimization of next-generation AI platforms and: silicon products, including hardware subsystems (CPU, GPU, DRAM, PCIe, networking, and storage), to meet the performance, scalability, and efficiency demands of large language model training and inference workloads.
  • Large-scale data pipeline architecture and ETL, distributed data processing: (Hive, Spark), and dashboard development;
  • Python, SQL, Tableau, Linux, and automation scripting;