infrastructure//compute cluster//inference pool

An inference pool is a group of accelerators that hold copies of a trained model and serve requests from users or applications, and it is where a model earns its keep after training. Each request is independent: a prompt arrives, a replica processes it (inference) and the answer streams back, while a load balancer spreads the traffic across replicas and adds or removes them as demand changes.


An inference pool is a group of accelerators that hold copies of a trained model and serve requests from users or applications, and it is where a model earns its keep after training. Each request is independent: a prompt arrives, a replica processes it (inference) and the answer streams back, while a load balancer spreads the traffic across replicas and adds or removes them as demand changes.

Its quality is measured from the user's side. Time to first token (TTFT) is how long a user waits before the answer starts; it grows with the prompt length, because the whole prompt is processed first, and with the queue in front of it. Tokens per second per user is how fast the answer then flows, set mostly by how quickly the GPU can read the model's weights from memory for each new token (memory bandwidth). Behind both sit the pool's concurrency, how many requests it serves at once, and its cost per request, which is the number the business watches (inference economics).

Serving is a trade between latency and cost per answer.

Grouping many requests on one GPU (inference batching) uses the hardware better and cuts the cost per token, but each user then waits on the others; memory for every conversation's KV cache caps how many fit at once. Every pool picks its point on that curve, and the queueing behaviour as load approaches capacity follows the same law as any server (Little's law).

It differs from a training cluster in almost every requirement. Training is one long synchronous job across thousands of GPUs and is limited by the network; serving is many short independent jobs, each replica fitting in one server or a few GPUs, limited by memory, latency targets and how well utilization follows demand.

Demand moves. Traffic follows the day and spikes on launches, so a pool either keeps spare capacity (fast answers, idle GPUs that still cost money) or scales with demand and accepts the delay of loading a model onto a new replica, which for a large model takes minutes (FinOps).

Inference can also be distributed: one large model split across the GPUs of one server, or prompt processing and token generation placed on different machines, each tuned for its own phase.