AI & High-Performance Networking · 1 October 2026
What changes when you design networks for GPU clusters
GPU clusters break assumptions that enterprise networks are built on: traffic patterns, oversubscription, failure domains and the cost of a single slow link.
A GPU cluster is not a faster data center network. It is a different network with different physics. The design habits that serve enterprise and general data center environments often work against you here, and the differences show up in four places.
Traffic is collective, not client-server
Enterprise traffic is largely north-south and bursty. GPU training traffic is east-west and collective: all-reduce, all-gather, all-to-all across the whole job. Every node in a training group talks to every other node, constantly, in patterns the network must serve simultaneously.
That changes what “good” means. Average utilisation is not the metric; the behaviour of the slowest flow is. A single congested link or a single lossy hop can hold back an entire training job, because collectives wait for the slowest participant.
Oversubscription is a different decision
In a general-purpose data center, oversubscription is normal and cost-effective: most workloads do not saturate the fabric. In a GPU cluster, the collective traffic can genuinely consume the fabric, so a design that looks efficient on paper becomes a bottleneck the moment the job scales.
This does not mean “build non-blocking everywhere” — that is an expensive reflex. It means the oversubscription ratio is a workload decision you make deliberately, per tier, with the failure cost understood. Training fabrics usually justify much lower oversubscription than storage or general compute tiers.
Failure domains matter more, not less
The usual availability conversation assumes that a link failure degrades capacity gracefully. In a GPU cluster it can also stall the job. Redundancy design therefore has to consider not just “can traffic reroute” but “does rerouting preserve the performance the collective depends on.”
Practical consequences:
- rail-optimised topologies that keep each GPU’s traffic on a predictable rail;
- multipath that is genuinely equivalent, not just reachable;
- congestion control and telemetry good enough to prove the rerouted path still performs.
The fabric is shared with storage
Modern AI clusters do not separate compute and storage as cleanly as a classical data center. Checkpointing, dataset loading and model serving put heavy demand on storage networking at the same time as training. If the storage fabric and the GPU fabric share capacity without an explicit plan, one will quietly degrade the other.
Treat storage networking as a first-class part of the AI fabric design, not an afterthought bolted on after the GPUs are racked.
What this means for network engineers
The skills transfer, but the priorities change:
- latency and jitter behaviour matters more than average throughput;
- lossless design (PFC/ECN or InfiniBand) becomes foundational, not optional;
- telemetry has to be per-flow and per-hop, or you cannot explain a slowdown;
- validation is at the workload level, not just link-up/link-down.
The failure mode is designing for availability when the workload needs deterministic performance. Those are different targets, and the network has to be designed for the one that actually matters. Get that framing right and the rest of the design follows.