mathematics//simulation//sim-to-real gap
The sim-to-real gap is the difference between how a controller, estimator or learned policy behaves in its simulator and how it behaves on the real machine, and managing it is a large part of the work of taking anything from a laptop to a drone, a robot arm or a plant. It comes almost always from the same places: parameters wrong by around 10 % (mass, inertia, motor constants); dynamics left out (the frame flexing, the motors' lag, the ESC's latency, thrust falling as the battery drains); sensors idealized, without the vibration that dirties a real IMU; timing, because real software has latencies and irregular periods; and contacts, friction and visual appearance when there are legs, grippers or cameras.
The sim-to-real gap is the difference between how a controller, estimator or learned policy behaves in its simulator and how it behaves on the real machine, and managing it is a large part of the work of taking anything from a laptop to a drone, a robot arm or a plant. It comes almost always from the same places: parameters wrong by around 10 % (mass, inertia, motor constants); dynamics left out (the frame flexing, the motors' lag, the ESC's latency, thrust falling as the battery drains); sensors idealized, without the vibration that dirties a real IMU; timing, because real software has latencies and irregular periods; and contacts, friction and visual appearance when there are legs, grippers or cameras.
Simulators err by omission, and the gap lives in what was omitted.
The danger is that a controller tuned in a perfect simulator learns to exploit that perfection: aggressive gains that are excellent in simulation oscillate on the real drone because the real loop carries 10 ms more delay, and a policy trained without motor lag shakes on the robot while one trained with optimistic friction slips.
The remedies form a ladder, from cheapest to dearest, and each rung is worth climbing only once the one below has been done.
1Put realistic noise, latencies and limits in the simulator2Fit its parameters to real datasystem identification3Randomize the parameters so one controller fits a familydomain randomization4Replay recorded real datalog replay5Put the real electronics in the loophardware-in-the-loop6Go from bench to tethered flight to the field
Measuring comes before randomizing. Randomizing blind, without having measured mass, inertia and delays, pays in robustness what an afternoon of bench tests would have cost: the wider the randomized ranges, the more conservative the resulting controller, the same trade as robust control.
Learned controllers feel the gap most. Reinforcement learning needs millions of attempts, affordable only in simulation, and the policy learns the simulator with its defects. Legged locomotion trained in simulation with randomization now transfers to commercial quadrupeds, and the Swift racing drone that beat world champions in 2023 combined a simulated policy with corrections learned from real flights; learning directly on an industrial plant remains rare because a chemical plant cannot be reset like an episode.
The gap is measured. A simulator compared with real logs on the same inputs shows where it is wrong and by how much; one that was never compared carries a gap of unknown size, and a validation run only inside it proves only what it already contains (simulation).
It also hides in plain modelling choices: a simulator integrated with too large a step can invent energy or instability the machine does not have, and that numerical gap looks exactly like a physical one (numerical stability).