OS//virtualization//container//seccomp

Seccomp (secure computing mode) is a Linux kernel feature that restricts which system calls a process may make, killing it or returning an error when it tries a forbidden one, and it is the layer that shrinks how much of the kernel a container can touch. Every action a program takes outside its own memory, opening a file, sending a packet, starting a process, loading a kernel module, is a request to the kernel through a syscall; Linux offers over three hundred of them, and most applications need a few dozen.


Seccomp (secure computing mode) is a Linux kernel feature that restricts which system calls a process may make, killing it or returning an error when it tries a forbidden one, and it is the layer that shrinks how much of the kernel a container can touch. Every action a program takes outside its own memory, opening a file, sending a packet, starting a process, loading a kernel module, is a request to the kernel through a syscall; Linux offers over three hundred of them, and most applications need a few dozen.

The original 2005 mode allowed only four calls (read, write, exit and return from a signal) and suited a process that only computes on data it is handed. The useful form today, seccomp-bpf (2012), attaches a small filter program to the process that inspects each system call and its arguments and decides: allow, deny with an error, or kill. Container runtimes apply a default profile that blocks the calls an ordinary application has no business making, such as loading kernel modules, rebooting the machine or changing the system clock.

Seccomp reduces the attack surface of the kernel.

A container shares the host's kernel, so a flaw in any system call it can reach is a possible way out. Blocking the calls a workload never needs closes most of those doors at once, cheaply, while namespaces limit what the process sees and cgroups how much it uses (attack surface).

It is one layer among several. A bug in an allowed system call is still reachable, and an overly permissive profile (or a container run as privileged, which disables it) leaves the kernel exposed; strong isolation between hostile tenants still calls for a VM or a hardened sandbox.

Writing a tight profile takes work. Blocking a call the application actually needs makes it fail in confusing ways, so profiles are usually generated by tracing what the program calls in testing.