infrastructure//horizontal and vertical scaling
Horizontal and vertical scaling are the two ways of giving a system more capacity: **vertical scaling** (scale up) makes one machine bigger, with more CPU, memory or a faster GPU, while **horizontal scaling** (scale out) adds more machines that share the work. The distinction is easy to blur: adding memory, cores or disks to an existing server is vertical, however much hardware it involves; only adding servers is horizontal.
Horizontal and vertical scaling are the two ways of giving a system more capacity: vertical scaling (scale up) makes one machine bigger, with more CPU, memory or a faster GPU, while horizontal scaling (scale out) adds more machines that share the work. The distinction is easy to blur: adding memory, cores or disks to an existing server is vertical, however much hardware it involves; only adding servers is horizontal.
Vertical scaling is the simple path. Nothing in the software changes, there is no network between the parts and no coordination to design, so a database or a simulation that runs on one machine simply runs faster on a larger one. Its limits are hard ones: the largest machine on sale sets a ceiling, the price per unit of capacity rises steeply near that ceiling, and the single machine stays a single point of failure.
Horizontal scaling has no fixed ceiling and survives the loss of a machine, but it moves the difficulty into the software. The work must be split, the parts must exchange data over a network with its own latency, and state shared between machines must be kept consistent (distributed systems). Stateless web servers behind a load balancer scale out almost for free; a single relational database does not, which is why it is often the last part of a system still scaled up.
Scale up until it hurts, then scale out what must grow.
Vertical scaling is cheaper in engineering and dearer in hardware; horizontal scaling is the reverse. The point where a system must go horizontal is set by the largest affordable machine, the need to survive failures, and how much of the workload can be split without constant coordination.
AI training uses both at once. Each server is scaled up with eight or more GPUs on fast links, and thousands of such servers are scaled out over the cluster network; the speed-up is limited by the communication the split requires (interconnect, distributed training).
In the cloud, horizontal scaling is usually automatic: a group of identical instances or containers grows and shrinks with load (Kubernetes), which only works if any instance can be killed without losing anything that matters.