Senior Site Reliability Engineer, PaaSActive

The opportunity

The Senior Site Reliability Engineer (IC4) position within the Platform As a Service team presents an exciting opportunity for a seasoned professional to enhance scalable infrastructure with a focus on CI/CD, Observability, and application hosting. In this role, you will bridge…

What you'll do

  • CI/CD Development and Maintenance: Contribute to the design, optimization, and maintenance of the CI/CD pipelines to improve the speed, reliability, and efficiency of the development lifecycle. Assist in driving standardization across various services hosted on the platform.

  • Observability Enhancement: Lead efforts to improve the observability of critical systems, working closely with cross-functional teams to ensure actionable monitoring and alerting frameworks are in place. Help troubleshoot complex issues and optimize system reliability.

  • Kubernetes and Cloud Management: Contribute to the development and operation of our Kubernetes-based architecture. Ensure systems are resilient, scalable, and optimized for performance. Actively participate in enhancing cloud-based solutions for API management and microservices.

  • System Optimization and Scaling: Collaborate with team members to ensure system scalability, operability, and performance. Lead initiatives to optimize resource utilization, focusing on cost efficiency while maintaining high system availability.

  • Mentorship and Knowledge Sharing: Mentor mid-level engineers (IC3) by providing guidance on technical challenges and SRE best practices. Support team growth by fostering knowledge-sharing sessions and helping establish processes that drive operational excellence.

  • Cross-Team Collaboration: Work closely with product, software, and other SRE teams to ensure that platform goals align with broader business objectives. Drive initiatives aimed at enhancing platform stability, security, and scalability.

What they're looking for

  • Strong Programming Skills: Proficient in Golang and Python with a solid understanding of software craftsmanship. Knowledge of Ruby is a plus.
  • Experience in CI/CD Pipelines: Hands-on experience in building and maintaining CI/CD pipelines using tools like GitHub Actions, CircleCI, or alternatives. Familiarity with best practices for ensuring build and deployment reliability.
  • Observability: Experience designing and implementing monitoring, alerting, and observability frameworks that provide actionable insights. Strong troubleshooting skills in production environments.
  • Kubernetes and Cloud Infrastructure: Proven experience in managing and optimizing Kubernetes-based architectures and working with public cloud providers such as GCP, AWS, or Microsoft Azure.