The opportunity
We are seeking a Site Reliability Engineer to join the Data Storage Team within the Platform Engineering Domain.
What you'll do
Develop and Implement: Build and maintain platform solutions that improve the reliability, security, and scalability of our Database Platform infrastructure.
Operational Excellence: Apply cloud infrastructure, networking, and CI/CD best practices to support our operational data storage solutions.
Collaborate: Work closely with Product Teams to facilitate seamless database: provisioning and ensure operational efficiency.
Knowledge Sharing: Participate in technical discussions and contribute to documentation, fostering a culture of transparency and continuous learning.
Execution: Contribute to the team’s roadmap by delivering high-quality code and infrastructure improvements.
Incident Response: Participate in on-call rotations and troubleshooting efforts to maintain a stable and predictable environment.
What they're looking for
- Observability: Maintain and enhance observability solutions (metrics, logging, tracing) to ensure the platform meets its SLOs and availability targets.
- Hands-on experience in maintaining cloud-based infrastructure, particularly in AWS.
- Practical experience with operational data storage solutions (e.g., PostgreSQL, AmazonRDS, S3 or similar).
- Hands-on experience with container technologies (Docker) and orchestration (Kubernetes/EKS)