The opportunity
We are looking for a Network Architect to join our Cluster Engineering Team and help shape the front-end datacenter and interconnect fabric for the current and next generations of our AI clusters. You will partner closely with hardware vendors, internal networking teams, and…
What you'll do
Design and architect front-end network fabrics for AI/ML and HPC clusters,: optimizing for high resource utilization, low latency, and high-throughput communication.
Build proof-of-concept implementations of new network designs and features,: and drive them from prototype through production rollout.
Identify and resolve performance and efficiency bottlenecks across the host-NIC-fabric
Automate the deployment, configuration, and validation of network: infrastructure using Python, including topology provisioning, fabric bring-up, config generation, and regression Strong programming skills are essential; this role builds tools, not just runbooks.
Stand up and operate SRE-grade telemetry and observability for the cluster: network: streaming telemetry (gNMI, OpenConfig, sFlow/IPFIX), metrics pipelines, alerting, and incident workflows. Define the SLIs/SLOs that govern network reliability and drive blameless post-incident analysis.
Lead network debugging in large distributed-systems environments spanning: multiple platforms and protocols, including deep dives into RoCEv2, PFC/DCQCN, ECMP hashing, congestion behavior, and packet-level forensics.
What they're looking for
- Lead cross-functional, multi-phase technical projects spanning hardware,: firmware, host networking, and cluster software.
- Collaborate with vendors and industry partners to shape network hardware and feature roadmaps.
- Represent the company in industry forums, standards bodies, and technical communities.
- Serve as the central point of contact for network reliability issues across the cluster.