AI Infrastructure Operations EngineerActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Assist with deployment and bring-up of CS-X systems, cluster servers, and: networking hardware; • Execute power-on sequencing, readiness checks, and validation tests. • Monitor hardware telemetry, alerts, and dashboards. • Perform first-line troubleshooting and structured escalation. • Collect logs, telemetry, and observations during incidents.

  • Participate in incident response under senior engineer guidance; • Use: existing monitoring, telemetry, and incident tracking tools. • Provide feedback on tooling and process gaps.

  • Build working knowledge of Cerebras system architecture; • Learn cluster: hardware and networking fundamentals. • Shadow senior engineers during complex debugging. • Progress toward independent ownership of defined workflows.

  • No people management; • No final escalation authority. • No ownership of: cluster architecture, hardware design, or tooling architecture.

  • Build a breakthrough AI platform beyond the constraints of the GPU.

  • Publish and open source their cutting-edge AI research.

What they're looking for

  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.