The opportunity
Render's mission is to eliminate the undifferentiated work that goes into building software products by offering an easy-to-use, powerful cloud platform for developer teams of all sizes.
What you'll do
Own Render's core compute infrastructure across multiple cloud providers,: regions, and data centers. You'll shape how we evolve our compute platform as we rapidly scale.
Design and build capabilities that give users greater performance and: flexibility in how their services are built, deployed, perform, and stay available even when underlying resources go down.
Investigate challenging cloud and compute issues across the stack, from the: kernel and data plane to our kubernetes cluster, control plane, and other orchestration mechanisms.
Improve the performance and reliability of our infrastructure through: systematic profiling, experimentation, and tuning.
Partner with engineers across the company to build a platform that is stable, predictable, and secure.
Participate in our on-call rotation. Help continuously improve how we detect,: respond to, and learn from incidents.
What they're looking for
- At least 7 years of experience building and operating large-scale platform or compute infrastructure.
- Deep expertise in operating, scaling, and enhancing Kubernetes clusters or: similar resource/container orchestration system.
- Experience developing in Go, Rust, or similar languages to develop custom: infrastructure components, scheduling, controllers, that apply business logic to resource management.
- Comfort going broad and deep in a complex systems, making tradeoffs to: improve performance and efficiency without sacrificing reliability.
- Strong experience designing, debugging, and operating distributed systems.
- Experience planning and executing rapid, high-risk upgrades, changes with minimal downtime to user services.