Senior Systems Engineer, Workers AIActive
The opportunity
You'll design and build the core infrastructure that powers AI inference across Cloudflare's global network — real-time voice, frontier open LLMs, and customer-deployed models running on a heterogeneous fleet of GPUs and next-generation accelerators in hundreds of cities…
What you'll do
Platform Architecture: Develop and maintain core components of the serverless inference platform to ensure high availability and scalability for Cloudflare users.
Optimization & Performance: Optimize the model scheduling system to significantly increase efficiency and resource utilization across our inference infrastructure, and implement improvements to request routing logic to reduce latency for end-users.
System Reliability: Drive significant, measurable improvements in the platform's reliability and resilience by identifying and mitigating systemic risks.
Observability: Expand and refine the observability stack (metrics, logging, tracing) and fine-tune alerts to proactively identify and resolve production issues.
Technical Leadership: Lead complex, cross-functional technical projects from initial concept and design through final deployment and operationalization.
Mentorship: Act as a mentor to junior engineers and actively contribute to: cultivating a strong, collaborative engineering culture within the team.