The opportunity
At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning.
What you'll do
Design, build, and scale services that orchestrate Ray clusters across cloud: and on-prem environments, supporting both VM-based and Kubernetes-based deployments
Optimize control plane components for large-scale, distributed AI/ML workloads
Build intelligent scheduling and resource management systems for heterogeneous compute clusters
Develop features to enhance the reliability, performance, scalability, and: observability of Anyscale-managed Ray workloads
Support and optimize accelerator integration (e.g., GPUs, TPUs).
Handle container image management and dependency resolution for distributed workloads
What they're looking for
- Participate in code reviews, design and architecture discussions
- Provide on-call support, working closely with customer and field teams to troubleshoot infrastructure issues
- Collaborate with leading distributed systems and machine learning experts to: push the boundaries of AI infrastructure
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience