The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Implement infrastructure to support high-performance, low-latency inference service.
Deploy and configure Kubernetes services to ensure scalability and reliability of inference workloads.
Optimize resource allocation and auto-scaling policies to handle variable: inference demand while minimizing operational costs.
Integrate inference services with containerized environments using Docker and Kubernetes for orchestration.
Ensure high availability and fault tolerance by implementing multi-region: deployments and disaster recovery strategies.
Develop Python-based scripts and APIs to streamline data preprocessing,: inference execution, and post-processing for real-time inference tasks.
What they're looking for
- Master's degree (or foreign equivalent) in Computer Science or a related field.
- One (1) year of experience as a Software Developer, Student/Intern (Software: Developer), Member of Technical Staff (Software Engineer), Software Engineer, or a related occupation.
- Employer accepts full-time or equivalent part-time experience gained before,: during, or after graduate studies.
- Docker and Kubernetes;
- ActiveMQ and Kafka;
- Python or Groovy;
- JavaScript or TypeScript;
- SQL, OracleDB, and Redis; and