Member of Technical Staff (Software Engineer, Cloud Infrastructure)Active$220K–$405K
The opportunity
The Cloud Infrastructure team owns the foundational cloud primitives and deployment models that power Perplexity's products, from multi-tenant public cloud to single-tenant and on-premises solutions for enterprise customers.
What you'll do
Own the roadmap and technical strategy for agent-driven cloud infrastructure management.
Design and operate Perplexity’s cloud networking fabric, including VPC: architectures, private connectivity, and peering with hyperscalers and neocloud providers to support low-latency, high-throughput AI workloads.
Architect and scale compute platforms (Kubernetes/EKS, autoscaling groups,: and mixed CPU/GPU fleets) to efficiently serve online request traffic and background workloads across regions.
Build and maintain secure, isolated deployment topologies for multi-tenant,: single-tenant, and customer‑owned cloud (BYOC) environments, including cross-account networking, identity, and policy guardrails.
Implement and evolve multi-region strategies for availability, failover, and: data locality, including traffic routing, regional capacity planning, and disaster recovery playbooks.
Partner with security to deliver enterprise controls such as BYOK/KMS: integrations, network isolation, and auditability required for regulated customers.
What they're looking for
- Deep experience designing and operating cloud infrastructure on AWS (VPC: design, routing, security groups, load balancing, private connectivity).
- Strong background with Kubernetes/EKS and container orchestration, including: multi‑cluster, multi‑region, or multi‑account setups.
- Hands-on experience with cloud networking and peering (VPC peering, Transit: Gateway, private link/service endpoints, or similar constructs in neoclouds).
- Experience building or operating secure, isolated environments for enterprise: customers (single‑tenant, BYOC, or on‑prem), ideally including BYOK/KMS integrations and compliance constraints.
- Proficiency with infrastructure as code (Terraform) and strong software: engineering skills in at least one of Python, Go, or Rust for automation and tooling.
- Strong debugging and incident management skills across distributed systems: (networking, compute, and platform layers), with a track record of driving root‑cause analysis and long‑term fixes.
- + years of industry experience building and operating production cloud: infrastructure, including leading the design of complex systems or migrations.