The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Drive and provide technical guidance to a team of software engineers working: on complex machine learning integration projects.
Design and implement ML features (e.g., structured outputs, biased sampling,: predicted outputs) that improve performance of generative AI models at inference time.
Design and implement high-throughput, low-latency multimodal inference models: that support delivery of image, audio, and video inputs and outputs.
Maintain our scalable serving backend for handling many concurrent requests per minute.
Scale our inference service by implementing detailed observability throughout the entire stack.
Analyze and improve latency, throughput, memory usage, and compute efficiency: on the service and the implementation of various features.
What they're looking for
- Optimize software to accelerate generative LLM inference by achieving high throughput and low latency.
- Stay up-to-date with advancements in machine learning and deep learning, and: apply state-of-the-art techniques to enhance our solutions.
- Evaluate trade-offs between different approaches, clearly articulate design: choices, and develop detailed proposals for implementing new features.
- Uncover, scope, and prioritize significant areas of technical debt across the: software stack to ensure continued high quality of the inference service.