The opportunity
We challenge the status quo because we know a better supply chain isn’t just possible—it’s essential. Better for our customers.
What you'll do
Design, build, and operate core infrastructure across our primary GCP: environment and secondary AWS footprint, leveraging managed services to deliver scalable, reliable, and cost-effective solutions
Own and evolve our Kubernetes (GKE) platform—cluster lifecycle management,: node pool strategy, networking, RBAC, and workload reliability—ensuring developer teams can ship safely and quickly
Build and maintain Terraform-based infrastructure-as-code across cloud: environments, enforcing standards for modularity, state management, and repeatable, auditable provisioning
Manage and improve our Kafka and Strimzi streaming infrastructure, supporting: high-throughput event pipelines that are foundational to project44's real-time data capabilities
Lead and execute cost optimization initiatives—rightsizing compute, improving: resource utilization, reducing waste, and creating visibility and accountability for cloud spend across engineering teams
Champion observability and reliability practices: SLOs/SLIs, incident response, on-call processes, and postmortems — to reduce MTTR and prevent recurrence.
What they're looking for
- Partner with application engineering teams to define and implement platform: standards for deployment, scaling, service mesh, secrets management, and infrastructure security
- Participate in a team on-call rotation, debugging and resolving: infrastructure and platform incidents with urgency and rigor
- Contribute to technical design discussions, architecture reviews, and: documentation—helping set a high bar for quality, reliability, and operational excellence across the engineering organization
- Mentor and support other engineers through code reviews, pairing, and technical guidance