Data Center

The Network Your GPUs Are Waiting On

Updated: 07 October 2026

Cisco Nexus fabric for AI workloads
5 Minutes Read

Cisco Nexus 9000 for AI Workloads: What an AI-Ready Fabric Actually Requires 

Every data center brochure printed this year says "AI-ready". Ask what number makes it ready and the conversation goes quiet. So here is a number to start with: in NVIDIA's enterprise reference architecture for GPU clusters, each eight-GPU server is provisioned with 400Gbps of inter-GPU network bandwidth per GPU. Not the whole rack. Per GPU. 

That is what AI does to a network, and it is why "AI-ready" without a figure attached is decoration. GPUs are the most expensive components most enterprises will ever rack, and in distributed training they spend their time either computing or waiting for each other across the network. A fabric that makes them wait converts your GPU budget into idle heat. The network is not a supporting act in an AI build; it decides what your GPUs are worth. 

This piece sets out what an AI fabric actually requires, what Cisco ships for it, and, because the doctrine of this publication is honesty, when you do not need any of it. 

What Does an AI Workload Demand from the Network? 

AI training demands three things from a network that ordinary enterprise workloads do not: enormous east-west bandwidth between GPUs, near-zero packet loss, and consistently low tail latency. Distributed training runs collective operations in which every GPU exchanges results with every other before the next step begins, so the slowest path sets the pace for the whole cluster and a single congested link idles every GPU at once. 

Traditional data center traffic is bursty, north-south heavy and tolerant of the occasional dropped packet. Training traffic is synchronised, all-to-all and intolerant: RDMA transports lose badly when packets drop. This is why an AI backend fabric is designed as its own network with its own rules, rather than as more ports on the campus-adjacent fabric you already run. 

What Makes an Ethernet Fabric Lossless? 

A lossless Ethernet fabric uses RoCEv2 (RDMA over Converged Ethernet) for direct GPU-to-GPU memory transfers, with two congestion mechanisms working together: priority flow control (PFC), which pauses traffic classes before buffers overflow, and explicit congestion notification (ECN), which marks packets early so senders slow down before pause becomes necessary. Load balancing spreads the massive parallel flows across every available path. 

Tuning these mechanisms is where AI fabrics are won and lost. PFC applied carelessly can spread congestion instead of containing it; ECN thresholds set wrong either drop packets or waste bandwidth. This is engineering, not box-buying, and it is the substance behind a claim the brochures make in two words. 

What Does Cisco Ship for AI Fabrics? 

Cisco's AI fabric line centres on the Nexus 9364E-SG2, a Silicon One switch built for GPU backend networks, deployed inside Cisco's AI POD reference designs that follow NVIDIA's enterprise architecture. The numbers, from Cisco's published designs: 

What The Figure
Nexus 9364E-SG2 capacity 64 ports of 800G Ethernet; 51.2 Tbps in 2RU
Port flexibility Each port runs at 800/400/200/100G
Burst handling 256 MB on-die buffer
Reference architecture NVIDIA 2-8-9-400 (8 GPUs, 400Gbps per GPU)
Scale per leaf Up to 32 eight-GPU servers
Reference cluster 256 GPUs: 32 servers under one fabric
Frontend provisioning 12.5Gbps per GPU for storage, 25Gbps per GPU for user traffic
Fabric technologies RoCEv2, PFC, ECN, EVPN-VXLAN, dynamic load balancing

Note the last row of context that table implies: an AI build is two fabrics, not one. The backend fabric carries GPU-to-GPU training traffic and is engineered lossless. The frontend fabric carries storage and user traffic at those per-GPU figures, and it is a job the Nexus 9300 switches you may already run are built for. Nexus Dashboard manages both, which matters when the person operating the AI pod is the same team running everything else. 

Do You Actually Need Any of This? 

Most enterprises running AI today do not need an 800G backend fabric, and a partner who quotes one for a handful of inference GPUs is selling you the brochure. Inference and small fine-tuning workloads on one or two GPU servers run comfortably on a well-built 25/100G leaf-spine, the kind of fabric a current Nexus 9300 deployment already provides. The dedicated lossless backend earns its cost when training or fine-tuning spans multiple GPU servers, because that is when the all-to-all traffic pattern appears and the network starts deciding your training times. 

The test is arithmetic, not fashion. Count the GPUs, name the workload, and ask whether GPU-to-GPU traffic crosses a server boundary. A Chennai GCC running retrieval and inference for internal copilots on two GPU servers needs clean 100G access and good storage paths, not a training fabric. The same GCC fine-tuning models across eight servers has crossed the line, and now the backend fabric is the difference between a two-day training run and a five-day one. Spend where the waiting is. 

How Do You Add an AI Pod Without Rebuilding the Data Center? 

An AI fabric arrives as a pod with a border, exactly like any other new environment in a running estate. The backend fabric is self-contained by design, so it does not disturb what exists; the frontend connects to your current network at layer 3 with ordinary route policy. Your existing switching, whatever vendor built it, keeps doing its job. The design questions that need answering before the purchase order are the unglamorous ones: power and cooling for dense GPU racks, storage throughput at those per-GPU figures, cabling for 400/800G optics, and who operates the pod on day two. 

We covered the general coexistence rules in our multi-vendor guide; an AI pod is their most extreme case, and their easiest, because isolation is the reference design. 

The Number Test 

When a vendor says "AI-ready", ask three questions: how many Gbps per GPU, end to end; what keeps the fabric lossless under an all-to-all collective; and who tunes it when the first training run stalls? The brochures that go quiet have answered you. 

Proactive Data Systems is a Cisco Penta-Preferred Partner with 35 years in Indian enterprise infrastructure, 100+ engineers and a 24x7 NOC, and we design GPU fabrics the same way we design everything: assessment first, numbers attached. Send us your GPU count and workload, and we will come back with the fabric design it actually needs and a budgetary quote, including the honest answer if that design is the network you already own. An engineer, not a salesperson, will respond.

 

Specifications and reference figures are taken from Cisco's published AI POD and Nexus 9000 documentation and NVIDIA's enterprise reference architecture; verify current specifications before design. This is an independent guide from Proactive Data Systems, a Cisco partner; NVIDIA marks belong to their respective owners.

Frequently Asked Questions

Ethernet works, and it is where enterprise reference designs are heading. NVIDIA's enterprise reference architecture endorses Ethernet-based configurations, and Cisco's AI PODs run RoCEv2 lossless Ethernet on Nexus 9000. The largest research-scale training clusters still often run InfiniBand; for enterprise training and fine-tuning, standards-based Ethernet keeps one operating model across the whole estate.
In NVIDIA's 2-8-9-400 enterprise reference architecture, each GPU in an eight-GPU training server is provisioned 400Gbps of backend inter-GPU bandwidth, plus 12.5Gbps for storage and 25Gbps for user traffic on the frontend. Inference-only workloads need far less; the backend figure applies when training traffic crosses between servers.
For inference and single-server fine-tuning, usually yes: a well-designed 25/100G leaf-spine handles those comfortably, with storage throughput the more common constraint. Multi-node training is the boundary; once GPU-to-GPU traffic crosses server boundaries, a dedicated lossless backend fabric earns its cost.
A Cisco AI POD is a validated building block for AI infrastructure: UCS GPU servers, Nexus 9000 switching for backend and frontend fabrics, and management through Nexus Dashboard, following NVIDIA's enterprise reference architecture. It gives an enterprise a tested design to size against rather than a fabric to invent from scratch.

Whitepapers

E-Books

Contact Us

We value the opportunity to interact with you, Please feel free to get in touch with us.

 

 

 

 

Share a few details to get started.

We'll get back to you shortly.