Network EngineerActive

The opportunity

We are looking for a Network Architect to join our Cluster Engineering Team and help shape the front-end datacenter and interconnect fabric for the current and next generations of our AI clusters. You will partner closely with hardware vendors, internal networking teams, and…

What you'll do

  • Design and architect front-end network fabrics for AI/ML and HPC clusters,: optimizing for high resource utilization, low latency, and high-throughput communication.

  • Build proof-of-concept implementations of new network designs and features,: and drive them from prototype through production rollout.

  • Identify and resolve performance and efficiency bottlenecks across the host-NIC-fabric

  • Automate the deployment, configuration, and validation of network: infrastructure using Python, including topology provisioning, fabric bring-up, config generation, and regression Strong programming skills are essential; this role builds tools, not just runbooks.

  • Stand up and operate SRE-grade telemetry and observability for the cluster: network: streaming telemetry (gNMI, OpenConfig, sFlow/IPFIX), metrics pipelines, alerting, and incident workflows. Define the SLIs/SLOs that govern network reliability and drive blameless post-incident analysis.

  • Lead network debugging in large distributed-systems environments spanning: multiple platforms and protocols, including deep dives into RoCEv2, PFC/DCQCN, ECMP hashing, congestion behavior, and packet-level forensics.

What they're looking for

  • Lead cross-functional, multi-phase technical projects spanning hardware,: firmware, host networking, and cluster software.
  • Collaborate with vendors and industry partners to shape network hardware and feature roadmaps.
  • Represent the company in industry forums, standards bodies, and technical communities.
  • Serve as the central point of contact for network reliability issues across the cluster.