The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Design and implement distributed runtime components to efficiently manage large-scale execution workloads.
Develop and optimize high-performance data and communication pipelines that: fully utilize CPU, memory, storage, and network resources.
Enable scalable execution across multiple compute nodes, ensuring high concurrency and minimal bottlenecks.
Collaborate closely with ML and compiler teams to integrate new model: architectures, training regimes, and hardware-specific optimizations.
Diagnose and resolve complex performance issues across the software stack: using profiling and instrumentation tools.
Contribute to overall system design, architecture reviews, and roadmap planning for large-scale AI workloads.
What they're looking for
- + years of experience developing high-performance or distributed system software.
- Strong programming skills in C/C++, with expertise in multi-threading,: memory management, and performance optimization.
- Experience with distributed systems, networking, or inter-process communication.
- Solid understanding of data structures, concurrency, and system-level: resource management (CPU, I/O, and memory).