robotics//fleet//fleet architecture

Fleet architecture is the choice of where the decisions of a multi-robot system are made: at a **centralized** server that sees everything, decides and sends orders; at each agent, **distributed**, deciding with what it knows and what its neighbours tell it; or in the **hierarchical** arrangement found almost everywhere, where the centre plans slowly and each agent reacts quickly on its own. It is the most consequential design decision of a fleet and the least glamorous one.


Fleet architecture is the choice of where the decisions of a multi-robot system are made: at a centralized server that sees everything, decides and sends orders; at each agent, distributed, deciding with what it knows and what its neighbours tell it; or in the hierarchical arrangement found almost everywhere, where the centre plans slowly and each agent reacts quickly on its own. It is the most consequential design decision of a fleet and the least glamorous one.

The trade is concrete. A centre can find the optimal plan because it sees the whole problem, debugging means reading one log with one clock, and verification is affordable; its weak points are the server and the link, each a single point of failure. Distributed agents remove that point if the design is careful, but each sees only part of the problem, so solutions are suboptimal, debugging means NNN logs with NNN clocks, and the global behaviour emerges and is hard to verify. Ten agents deciding with partial and stale information produce failures that show up at four in the morning in the customer's warehouse.

Centralize while the network and the computation allow it; distribute only what must survive without the link or react faster than the round trip to the server. Immediate safety is always local. A robot at 2 m/s that waits for a stop order from a server 100 ms away travels 20 cm before it starts braking, if the packet arrives at all; its laser scanner stops it without asking. The fast loop goes next to the actuator.

Two questions settle most cases: is there a reliable network where the fleet operates, and does the problem fit on one server? If both answers are yes, the architecture is decided. Distribution pays without infrastructure, on a network that can fail or be attacked, or when the server itself is the bottleneck.

The same principle recurs across the stack. It decides when a fleet allocates tasks by auction instead of with a central solver (task allocation), when it runs a consensus protocol instead of broadcasting a reference, when distributed estimation (distributed Kalman filter) is worth its rounds of messages over sending every measurement to one filter, and why swarm robotics stays in research. In computing it is why a single database with backups beats a replicated one until scale or availability demand otherwise (distributed systems).

The hierarchy is also a separation of timescales: motor control at hundreds of hertz on the microcontroller, the fleet layer at 1 to 10 Hz on board, planning at about one cycle per second on a local server (hierarchical control). The bandwidth and loss that push decisions outward are in communication constraints.

It applies to teams as well. By Conway's law a system copies the communication structure of the organisation that builds it, so the interface between two teams that do not talk is where the bugs will live. The principle's place among the others is in engineering patterns.