infrastructure//compute cluster//training cluster
A training cluster is a compute cluster built to train one large model at a time, thousands to hundreds of thousands of GPUs working on the same job for weeks, and its design is dominated by the communication between those GPUs. In distributed training every GPU computes the gradients of its share of the data, then all of them exchange and average those gradients (an all-reduce) before anyone can take the next step; with the model itself split across GPUs, activations cross the network inside every step as well.
A training cluster is a compute cluster built to train one large model at a time, thousands to hundreds of thousands of GPUs working on the same job for weeks, and its design is dominated by the communication between those GPUs. In distributed training every GPU computes the gradients of its share of the data, then all of them exchange and average those gradients (an all-reduce) before anyone can take the next step; with the model itself split across GPUs, activations cross the network inside every step as well.
That lockstep makes the cluster behave like one machine with one clock. The whole job moves at the pace of the slowest GPU and the slowest link, so a training cluster is engineered around the interconnect: fast links inside each server or rack (the scale-up domain), a non-blocking fabric across the hall, and a placement of the job's parts that keeps the heaviest traffic on the shortest links. The GPUs are usually rented or owned as bare metal so no virtualization layer adds jitter.
In training, the network is part of the computer.
Doubling the GPUs only halves the time if the communication keeps up; when it does not, the extra GPUs spend their time waiting. That is why the interconnect, the topology and the collective-communication software are bought and tuned with the same care as the accelerators, and why one rack of current AI hardware is designed as a single unit (rack).
It must survive its own size. With tens of thousands of GPUs, hardware fails every few hours somewhere in the cluster; the job writes checkpoints of the model and optimizer state regularly, and the time to detect a failure, swap the node and reload the last checkpoint is lost work that grows with scale.
It is power-hungry in a peculiar way. Thousands of GPUs switching between computing and waiting at the same instant make the whole cluster's draw swing by megawatts in milliseconds, a load the power chain and sometimes the grid must absorb.
The contrast is the inference pool: training is one synchronous job measured in time to finish, serving is many independent requests measured in latency and cost per answer, and the hardware choices follow from that difference.