Software Engineer (Site Reliability Engineer)Posted today

The opportunity

At  Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing  Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning.

What you'll do

  • Develop a unified perspective on how cloud components are utilized across the: company, taking into account diverse needs and requirements.

  • Ensure that deployment methodologies align with the company's reliability goals.

  • Build systems that promote understanding of production environments,: facilitating quick identification of issues through robust observability infrastructure for metrics, logging, and tracing.

  • Create monitoring and alerting systems at different levels, enabling teams to: easily contribute and enhance the overall monitoring capabilities.

  • Establish testing infrastructure to support the team in writing and executing tests effectively.

  • Develop tools for measuring service level objectives (SLOs) and define organization-wide SLOs.

What they're looking for

  • Implement best practices and on-call systems, ensuring efficient incident: management and up-leveling the incident management system at Anyscale.
  • Coordinate the creation and deployment of cloud-based services, including: tracking deployments and establishing effective communication channels for issue resolution.
  • At least 3 years of relevant work experience in a similar role.
  • Stock Options