The opportunity
At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning.
What you'll do
Design, build, and improve the core systems that power Ray Data , with a: focus on performance, scalability, and reliability.
Design and optimize distributed execution across different stages of data: pipelines in heterogeneous environments.
Build data loading and processing solutions for production training and inference workloads.
Solve challenging problems in distributed execution, scheduling, resource: management, data partitioning, fault tolerance, and performance optimization.
Make system-level architectural decisions and reason through tradeoffs in: areas such as resource allocation, execution models, batch vs. streaming workloads, and consistency and availability.
Work with customers and new-age AI-native companies to understand and solve: challenges in scaling their AI workloads.
What they're looking for
- + years of experience building production-grade software, infrastructure, or: developer-facing systems, with strong Python engineering experience.
- + years of experience personally owning core architectural decisions within a: distributed data or compute engine, rather than primarily operating or using a platform someone else designed.
- Deep experience with distributed systems internals, such as scheduling, fault: tolerance, data partitioning, distributed execution, performance optimization, or database and query engine internals.
- A track record of reasoning through system-level tradeoffs and defending: architectural decisions, such as batch vs. streaming, static vs. dynamic resource allocation, or consistency vs. availability.