The opportunity
At Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning.
What you'll do
Build the next generation of developer tooling and MLOps capabilities on Ray,: designed for both developers and coding agents.
Develop an agent-first CLI and cohesive SDK, API, and MCP surfaces with: self-discovery, structured errors, dry-run support, and consistent behavior across platform resources.
Work across the Anyscale Workspaces stack to improve the path from local code: to distributed execution, including environments, dependencies, images, authentication, workload submission, and debugging.
Build cohesive experience, tools and frameworks for the AI development: lifecycle, including data preparation, fine-tuning and post-training, evaluation, production serving, dataset management, experiment tracking, and lineage.
Build the path from a trained model to a reliable production endpoint,: including model registration, deployment workflows, performance benchmarking, and LLM-specific service metrics.
Surface observability across the CLI, SDK, UI, and agent-facing interfaces so: users can diagnose failures across jobs, tasks, actors, nodes, and GPUs.
What they're looking for
- Design open integrations using durable standards such as OpenAI-compatible: APIs and OpenTelemetry, along with stable Jobs and Services interfaces for external orchestrators.
- Design and operate the highly available backend services and platform: architecture that power these capabilities across serverless and bring-your-own-cloud environments.
- Work closely with users and field teams to scope, ship, and iterate on the: product, and collaborate with distributed systems and machine learning experts to push the boundaries of AI infrastructure.
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience