infrastructure//Kubernetes

Kubernetes is an open-source container orchestrator: software that runs containers across a fleet of machines, deciding where each one runs, restarting the ones that fail and adding or removing copies as load changes, and it is the standard way to operate services on a cluster today. Google released it in 2014, drawing on its internal cluster manager, and it is now maintained under the Cloud Native Computing Foundation and offered as a managed service by every large cloud provider.


Kubernetes is an open-source container orchestrator: software that runs containers across a fleet of machines, deciding where each one runs, restarting the ones that fail and adding or removing copies as load changes, and it is the standard way to operate services on a cluster today. Google released it in 2014, drawing on its internal cluster manager, and it is now maintained under the Cloud Native Computing Foundation and offered as a managed service by every large cloud provider.

Its core idea is declarative: the operator writes down the desired state (run five copies of this image, each with two CPUs and one GPU, reachable at this address) and the system works continuously to make reality match it. If a machine dies, its containers are rescheduled elsewhere; if a container crashes, it is restarted; if traffic grows, an autoscaler raises the number of copies (horizontal scaling). The unit it schedules is a pod, one or more containers that share a network address and storage, which is a different thing from a data center pod (rack).

The scheduler is the part that matters most for AI. It places each pod on a node with enough free CPU, memory and, with the right plugins, GPUs; for training it must also place the many workers of one job close together on the network and start them all at once, since a job with half its workers running only wastes the GPUs it holds. That is why training clusters often add batch-scheduling extensions or keep Slurm, the older scheduler of high-performance computing (compute cluster).

Kubernetes automates the operation of stateless, replaceable services very well, and asks a lot in return: its own control plane to run and secure, a large configuration surface, and applications designed so that any container can be killed and replaced at any moment. A small system on a few machines rarely needs it; a fleet of services that must scale and heal by itself usually does.

Containers on one node share the host's kernel, so Kubernetes' isolation between workloads is the container's (namespaces and cgroups underneath), weaker than a virtual machine's; multi-tenant clusters add policies, separate nodes or VMs for stronger walls (multi-tenancy).