infrastructure//compute cluster

A compute cluster is a set of servers joined by a fast network and run by a common scheduler so that they can be used as one large computing resource, and it is the machine on which large models are trained and served and on which physics is simulated. Each server (a node) has its own processors, GPUs and memory; what makes them a cluster is the interconnect between them, the shared storage, and the software that places jobs on nodes and restarts what fails (Kubernetes, or Slurm in high-performance computing).


A compute cluster is a set of servers joined by a fast network and run by a common scheduler so that they can be used as one large computing resource, and it is the machine on which large models are trained and served and on which physics is simulated. Each server (a node) has its own processors, GPUs and memory; what makes them a cluster is the interconnect between them, the shared storage, and the software that places jobs on nodes and restarts what fails (Kubernetes, or Slurm in high-performance computing).

Clusters exist because the job is bigger than any machine (horizontal and vertical scaling). A frontier model needs more memory and arithmetic than any single server holds; a weather or fluid-dynamics simulation splits a grid over thousands of cores that exchange their boundaries every step. In both cases the result depends as much on how fast the nodes talk as on how fast each one computes.

Clusters are now usually described by what they do, because the two big AI workloads want different things:

A training cluster runs one huge job for weeks: thousands of GPUs computing in lockstep and exchanging gradients at every step, so it is built around the network and cares about the slowest node.

An inference pool serves millions of independent requests: many smaller copies of a model, each answering users, so it is built around latency per request, memory for the KV cache and cost per answer.

HPC (high-performance computing) clusters run scientific simulations (climate, materials, Monte Carlo studies, crash tests) and share much of the same hardware and network; the boundary with AI clusters is fading as simulation and learned models mix.

A cluster is only as fast as its coordination. Adding nodes helps until the time spent exchanging data, waiting for stragglers and recovering from failures grows faster than the computation each node adds; beyond that point more hardware buys nothing, which is why the network, the scheduler and the failure handling are designed as carefully as the nodes.

At scale, failure is routine: with tens of thousands of GPUs, some component fails every few hours, so a long training run checkpoints its state regularly and the cluster's software detects and replaces bad nodes without stopping the job for good.