The opportunity
At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning.
What you'll do
Own a major technical area of Ray Core end to end, defining the roadmap,: identifying the most important technical problems, and driving execution through production.
Lead large, technically complex projects spanning multiple engineers, teams, and/or organizations.
Set technical direction and make architectural decisions for distributed: computing infrastructure used by demanding production workloads.
Design, build, and evolve core distributed-systems primitives rather than: simply integrating existing platforms or frameworks.
Work on problems involving areas such as distributed execution, scheduling,: resource management, fault tolerance, concurrency, networking, storage, or system performance.
Stay hands-on with implementation and debugging in a systems-oriented: codebase, while raising the technical bar for the engineers around you.
What they're looking for
- + years of software engineering experience, with a track record of increasing technical ownership.
- Experience leading substantial projects end to end, including defining the: problem, creating a roadmap, making architectural decisions, driving implementation, and owning the outcome in production.
- Experience leading projects that are larger than a single-engineer effort,: typically spanning multiple engineers and lasting multiple quarters.
- Deep experience with distributed systems and computer systems.
- Strong systems programming experience in languages such as C++, Rust, Java, or similar lower-level languages.
- Strong understanding of systems concepts such as multithreading/concurrency,: distributed coordination, resource management, fault tolerance, performance, or networking.
- Experience building foundational systems such as databases, streaming: systems, distributed runtimes, operating systems, schedulers, storage systems, Spark, Kafka, or similar infrastructure is highly relevant.
- A track record of mentoring engineers and raising the technical bar of the teams around you.