robotics//perception

Perception, in robotics, is the combination of sensing and heavy computation that turns raw data from cameras, lidars and radars into a description of the surroundings (objects, obstacles, free space, landmarks), and it is what lets a vehicle avoid a pallet, a car keep its lane or a drone find a hot spot in a field. Its output is a list of **detections**: this box is probably a person at about 12 m, that cluster of points an obstacle. Those detections are the product of a pipeline that can be wrong in three ways at once: it reports things that are not there (false positives), misses things that are (missed detections), and places real objects with an error that grows with distance.


Perception, in robotics, is the combination of sensing and heavy computation that turns raw data from cameras, lidars and radars into a description of the surroundings (objects, obstacles, free space, landmarks), and it is what lets a vehicle avoid a pallet, a car keep its lane or a drone find a hot spot in a field. Its output is a list of detections: this box is probably a person at about 12 m, that cluster of points an obstacle. Those detections are the product of a pipeline that can be wrong in three ways at once: it reports things that are not there (false positives), misses things that are (missed detections), and places real objects with an error that grows with distance.

The sensors fail in different conditions. A camera gives texture, colour and signs cheaply and is blind in the dark, in glare and in motion blur; a lidar gives centimetre geometry at night and is confused by rain, fog, black surfaces and glass; a radar sees through fog and measures velocity directly, with poor angular resolution and ghost echoes. Depth from cameras alone comes from stereo vision or from learning, and the learned detectors are mostly convolutional networks.

Fire-spotting drones with thermal cameras show both kinds of error in one flight: the sun reflected on a metal roof looks like a fire, and dense smoke hides a real one. Treating every detection as true is the origin of many failures downstream, which is why trackers carry hypotheses and confirm them over several frames (multi-target tracking, data association).

Combining detections from several sensors, each with its own failure conditions, is sensor fusion; all of them first need calibration of their internal parameters, their mounting pose and their time offset (sensor calibration).

The simplest sufficient sensor often wins. To know whether a box is on a conveyor, a photoelectric barrier costs cents and is never fooled by the lighting; a camera earns its place when the question is what the object is, where it is and what shape it has. Camera inspection and driver assistance are industry, lidar in mapping and warehouse robots is industry in those sectors, event cameras are still research.