Updated: July 16, 2026
A GPU cluster is a choir, not a room full of soloists.
Each GPU only produces something useful when it sings in time with the others, and the thing keeping them in time is the network between them. Get the fabric wrong, and you have bought the finest voices in the world and a choir that cannot hear itself. The GPUs sit there, waiting, half-idle, while the invoice for them does not wait at all.
This is the part of AI infrastructure that decides how much of your GPU spend you actually get to use. And for years the answer was simple: for serious AI, you ran InfiniBand. That answer is now genuinely contested, which is why it is worth understanding properly.
Because GPUs in a training run work in lockstep, and the network is where the lockstep lives or dies.
Training a large model means the GPUs constantly exchange and combine their results, an operation the field calls all-reduce, over and over, thousands of times. If the network is slow, congested, or drops a packet under load, every GPU waits for the slowest exchange to finish. One stall does not slow one GPU. It stalls the whole cluster.
So the network is not plumbing behind the compute. It is a multiplier on it. A fabric that keeps the GPUs fed and synchronised turns your accelerators into a machine; a fabric that cannot becomes the ceiling on everything you paid for.
InfiniBand is a networking technology built from the start for high-performance computing, and it has been the default for large-scale AI for good reason.
It is lossless by design, delivers very low latency, on the order of a microsecond, and stays consistent as clusters grow. Its decisive trick at large scale is a feature that performs part of the all-reduce inside the switches themselves, combining results in-flight rather than shuffling everything back and forth between nodes. On big clusters, that is a real and measurable edge. InfiniBand's reputation is earned: at the frontier of model training, it has been the safe, proven choice.
The catch is that it is largely a single-vendor world, with its own hardware, its own skills, and a price to match.
RoCE, RDMA over Converged Ethernet, is the technology that lets ordinary Ethernet behave more like InfiniBand: low-latency, direct memory-to-memory transfers over the network everyone already knows.
For a while, RoCE was the pragmatic-but-second-best option. It worked, but it was harder to run cleanly at scale, prone to congestion problems that InfiniBand did not have. Two things changed that. NVIDIA built Spectrum-X, an Ethernet fabric that ports many of InfiniBand's tricks onto Ethernet and, on smaller clusters, comes within a few per cent of it. And the industry formed the Ultra Ethernet Consortium, a who's-who of AMD, Intel, HPE, Broadcom, Cisco, Meta and Microsoft, explicitly to build an open Ethernet fabric that matches InfiniBand for AI without the proprietary lock-in.
The result is a genuinely shifting market. InfiniBand held around 80% of AI cluster networking in 2023. By 2025, Ethernet had taken the lead in new AI back-end deployments. The gap did not close by accident; the whole industry pushed it closed.
The table is the honest state of play. Treat the assessments as informed opinion on a fast-moving field.
| Factor | InfiniBand | Ethernet (RoCE / Spectrum-X) |
|---|---|---|
| Latency | Lowest, ~1–2 µs, very consistent | Higher historically, now close when well-tuned |
| Performance at large scale | Still ahead, aided by in-network aggregation | Excellent to very large; narrows the lead each year |
| Ecosystem | Largely single-vendor | Open, multi-vendor, and consolidating around standards |
| Skills | Specialist InfiniBand expertise | Ethernet skills your team likely already has |
| Cost | Premium | Generally lower; leverages the wider market |
At the frontier, InfiniBand keeps its edge.
If you are building among the largest training clusters, where thousands of GPUs must stay perfectly synchronised, InfiniBand's consistency and its in-switch aggregation still deliver the best and most predictable performance. For an organisation whose entire competitive position rests on training time at extreme scale, that edge can justify the premium and the single-vendor commitment. This is why the biggest AI labs have leaned on it for so long. It is the safe answer when nothing but the maximum will do.
For almost everyone else, Ethernet's case is getting harder to argue against.
It runs on the networking model your team already understands, it draws on a competitive, multi-vendor market rather than a single supplier, and it typically costs less. With Spectrum-X and the Ultra Ethernet standards, its performance for the great majority of enterprise AI workloads is now close enough that the remaining gap rarely justifies the price and the lock-in. And it converges with the rest of your data center network rather than standing apart as a specialised island. For an enterprise building AI infrastructure it will actually have to operate and afford, that combination is compelling, which is exactly why the market has tilted its way.
Start with your scale and your honesty about it.
If you are training at the absolute frontier and every microsecond of synchronisation is a competitive weapon, InfiniBand still earns its keep. If you are like most enterprises, building serious but not record-breaking AI, running a team that already knows Ethernet, and weighing cost and vendor lock-in alongside raw performance, Ethernet with RoCE or Spectrum-X is increasingly the rational choice. For Indian enterprises in particular, the wider skills pool, the multi-vendor pricing, and the avoidance of single-vendor dependence often tip the balance toward Ethernet, unless the workload genuinely demands the frontier. There is no universal answer, only the fabric that fits your scale, your team and your budget, on a field that is moving quickly.
The fabric decision is easy to get wrong by defaulting to what the biggest labs use, or by under-building to save on the network and then throttling the GPUs it feeds. Sizing the cluster, choosing InfiniBand or Ethernet on the evidence for your scale, and building the spine-and-leaf fabric that carries it, is where networking depth pays off.
Proactive Data Systems designs and builds AI-ready data center networking for Indian enterprises, across InfiniBand and high-speed Ethernet fabrics on Cisco, NVIDIA and other platforms, matched to your cluster rather than a single answer. We are a Cisco Preferred Cloud and AI Partner, Dell Platinum Partner and NetApp Preferred Partner, with 35 years in enterprise IT, deep networking expertise, more than 1,500 organisations served, and a 24/7 service desk in India. To choose the fabric that fits your AI, you can ask Proactive for a data center networking assessment.
Sources:
Market shift (InfiniBand ~80% of AI cluster networking in 2023;
Ethernet leading new AI back-end deployments by 2025) and performance comparisons (InfiniBand ~1–2 µs latency and in-network SHARP aggregation;
RoCEv2 and NVIDIA Spectrum-X closing the gap;
Ultra Ethernet Consortium UEC 1.0): AI networking analyses, 2025–2026. Figures move quickly; verify current benchmarks before acting.
![]()
Disclaimer: This is an independent comparison for general guidance on a fast-evolving technology, not a recommendation for any specific environment. Performance figures are drawn from published benchmarks and vary by configuration, scale and tuning. Verify current specifications before deciding. InfiniBand, NVIDIA Spectrum-X, RoCE and related marks are trademarks of their respective owners; this comparison is not endorsed by or affiliated with any of them and should be reviewed by technical and legal teams before publication.
We'll get back to you shortly.