computer science//distributed systems//Byzantine fault tolerance
Byzantine fault tolerance is the ability of a distributed system to keep agreeing correctly while some of its nodes behave arbitrarily, sending wrong, contradictory or malicious messages, and it is what a system needs when a node can be compromised, or broken in a way that does not simply stop it. A **Byzantine fault** is any behaviour outside the protocol: a server that tells two peers different things, a drone whose stuck sensor keeps broadcasting the same value, an attacker who has taken over a node. The name comes from Lamport, Shostak and Pease's 1982 paper, framed as generals who must agree on a plan while some of them are traitors.
Byzantine fault tolerance is the ability of a distributed system to keep agreeing correctly while some of its nodes behave arbitrarily, sending wrong, contradictory or malicious messages, and it is what a system needs when a node can be compromised, or broken in a way that does not simply stop it. A Byzantine fault is any behaviour outside the protocol: a server that tells two peers different things, a drone whose stuck sensor keeps broadcasting the same value, an attacker who has taken over a node. The name comes from Lamport, Shostak and Pease's 1982 paper, framed as generals who must agree on a plan while some of them are traitors.
The bound is the first thing to know. Agreement with fff arbitrary faults needs at least 3f+13f+13f+1 nodes: with fewer, a traitor can tell half the loyal nodes one thing and the other half another, and the loyal ones have no way to tell which story is false. A crash-only system such as Raft needs 2f+12f+12f+1 to survive fff failures, because a stopped node never lies. The extra fff replicas, plus more rounds of messages and signatures on all of them, are the price of trusting nobody.
In linear averaging, one node that does not listen owns the result.
A Byzantine agent in a consensus protocol only has to keep broadcasting a fixed value: everyone else drifts toward it, exactly as followers drift toward the leader in leader-follower. A malicious drone and a broken one look the same to the average.
In computing, Byzantine-tolerant replication (PBFT and its descendants) runs where participants do not trust each other, which is why permissioned blockchains use it (blockchain consensus). In flight control, redundant computers vote so that one faulty unit cannot steer the output (majority voting).
In multi-agent control the defence is resilient consensus: each agent discards the most extreme neighbour values before averaging, which works only if the communication graph is redundant enough.
Identity comes first. A node that can pretend to be many (Sybil attack) breaks any bound counted in nodes, so Byzantine tolerance presumes signed messages and managed keys (authentication); MAVLink 2 can sign its packets for that reason.
It is often confused with ordinary fault tolerance. Tolerating crashes is the common case and much cheaper, so the first design question is which failure model the system really faces, before paying for the Byzantine one.