The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Design and implement runtime components and high-performance kernels required by novel Core ML algorithms.
Translate research prototypes into efficient implementations for the Cerebras: platform, including reference implementations and comparisons on GPUs where useful.
Profile and debug performance across the ML framework, compiler, runtime, communication, and kernel layers.
Optimize computation, memory movement, communication, and concurrency for: large-scale training and low-latency inference.
Develop benchmarks, instrumentation, and automated tests that validate: functionality, performance, and numerical correctness.
Collaborate closely with Core ML researchers and compiler, runtime, kernel,: and inference engineers to evaluate design alternatives and deliver end-to-end capabilities.
What they're looking for
- Contribute to software architecture and roadmap decisions by identifying: recurring limitations and high-leverage platform improvements.
- Bachelor’s, Master’s, PhD, or equivalent practical experience in Computer: Science, Computer Engineering, Electrical Engineering, or a related field.
- Experience developing high-performance systems software, ML systems,: runtimes, compilers, or computational kernels.
- Strong programming skills in C++ and Python.