Sr. Member of Technical StaffActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Design and develop software features that support system resiliency and high: availability, including automated recovery mechanisms and fault-tolerant architecture across distributed environments.

  • Develop and maintain cloud-based deployment workflows for AI inference: software using AWS tools and services to support low-latency and scalable system performance.

  • Develop Python-based scripts and APIs to streamline data preprocessing,: inference execution, and post-processing for real-time inference tasks.

  • Use parallel programming techniques (e.g., multi-threading, asynchronous: processing) to maximize resource efficiency on AWS compute instances.

  • Develop software components to support visualization and analysis of system: performance metrics, enhancing the monitoring and usability of inference services.

  • Develop inference software in Docker containers and define Kubernetes: orchestration strategies that ensure software reliability and efficient scaling.

What they're looking for

  • Develop automated scripts to detect and mitigate common failure modes, improving software system reliability.
  • Debug issues related to model deployment, container orchestration, and: networking configurations, documenting steps to reproduce and root-cause defects.
  • Triage and resolve defects in the software service by analyzing logs,: metrics, and distributed traces using tools like AWS CloudWatch, Grafana, or custom Python scripts.
  • Work with Product Management and User Experience teams to define requirements: for inference service interfaces, including configuration, monitoring, and event logging.
Sr. Member of Technical Staff at Cerebras | Role Match