The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Core Inference Observability: Design and implement end-to-end telemetry systems across the software stack, providing deep visibility into inference performance and enabling rapid iteration before and after deployment.
Benchmarking Infrastructure: Architect, build, and scale the automation that generates, analyzes, and visualizes performance data used to inform business decisions across engineering and leadership.
Performance Analysis: Dive deep into system behavior, dissect performance bottlenecks, and deliver actionable insights that directly influence which features ship and how they evolve.
Feature Integration: Partner closely with Core Platform teams to define rigorous testing methodologies that validate inference features for peak performance.
Bachelor’s or Master’s degree in Computer Engineering, Systems Engineering, or a related field.
Proficiency in Python and/or C++ programming.
What they're looking for
- Proven experience in building and scaling automated infrastructure.
- Strong background in throughput and performance optimization techniques,: especially in complex, large-scale systems.
- Excellent problem-solving skills and a strong analytical mindset.
- Demonstrated ability to dive deep into new domains.