infrastructure//data center//interconnect
An interconnect is the set of links that join processors, servers and racks in a data center so they can exchange data, and in an AI cluster it decides how many GPUs can work on one problem before they spend more time waiting for each other than computing. A model trained across thousands of GPUs exchanges gradients or activations at every step (distributed training), so the links carry terabytes per second and their latency enters every iteration.
An interconnect is the set of links that join processors, servers and racks in a data center so they can exchange data, and in an AI cluster it decides how many GPUs can work on one problem before they spend more time waiting for each other than computing. A model trained across thousands of GPUs exchanges gradients or activations at every step (distributed training), so the links carry terabytes per second and their latency enters every iteration.
Two physical media carry the traffic, and the choice between them is a matter of distance:
Copper carries electrical signals and is cheap, needs no conversion and uses little power over short runs, but at hundreds of gigabits per second the signal degrades within a few metres. It joins GPUs inside a server and servers inside a rack; NVIDIA's NVL72 racks wire 72 GPUs together through a copper backplane precisely because it avoids optics.
Optical fibre carries light, keeps the signal over hundreds of metres or kilometres and packs more bandwidth per cable, but every link needs an optical transceiver at each end to convert electrical signals to light and back. Those modules cost money, draw power (several watts each, times tens of thousands of links) and fail more often than passive copper.
So the claim fibre is always better is wrong both ways: fibre wins on distance and bandwidth, copper on cost, power and reliability over the last few metres, and data centers use both.
An AI cluster is two networks.
The scale-up network (NVLink inside a server or rack) makes a few dozen GPUs behave almost like one chip, with very high bandwidth over short copper. The scale-out network (InfiniBand or Ethernet with RDMA over fibre) joins thousands of those blocks across the hall, slower per GPU but far larger. A training job is laid out so the heaviest communication stays inside the scale-up domain (compute cluster, rack).
The fabric's topology matters as much as its speed. Fat-tree and rail-optimized layouts keep any two GPUs a few switch hops apart; an oversubscribed network, cheaper to build, slows exactly the collective operations training depends on.
The bottleneck moves. A faster GPU only helps if the links keep up; otherwise the arithmetic sits idle waiting for data, the same lesson as memory bandwidth at the scale of the building.