The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Characterize failure signatures and perform hands-on troubleshooting at the: system, board, and component level.
Drive failure analysis to root cause—down to silicon, PCB, or external: components—using structured, first-principles problem solving.
Remotely diagnose failures in deployed systems using Linux command-line: tools, system logs, telemetry, diagnostic utilities, and remote-access workflows.
Analyze and document results clearly, and provide actionable recommendations: to design and manufacturing teams.
Interface across internal groups (ASIC, systems, and manufacturing) and with: external vendors to resolve issues and close corrective actions.
Participate in design reviews for next-generation systems, contributing: recommendations spanning electrical design-for-test (DFT), mechanical, and thermal improvements.
What they're looking for
- Bachelor’s degree in Electrical Engineering or a related field.
- + years of experience in hardware bring-up, system integration, board/system debug, or failure analysis.
- Strong analytical, diagnostic, and problem-solving skills, grounded in first principles.
- Strong drive to self-educate and operate with autonomy in fast-moving, ambiguous environments.