systems engineering//verification//log replay

Log replay is the testing of new software by feeding it data recorded from the real system, as if the recording were happening live; it is used to evaluate a new estimator, detector or perception model on real sensor data before it ever runs on a vehicle, and it is the cheapest test with reality in it. A drone's flight log (a PX4 ULog, a ROS 2 bag) holds every IMU sample, GPS fix and barometer reading with their timestamps; replaying it through a modified EKF shows at once how the new filter would have tracked that flight, including the gust and the GPS glitch that happened at minute seven.


Log replay is the testing of new software by feeding it data recorded from the real system, as if the recording were happening live; it is used to evaluate a new estimator, detector or perception model on real sensor data before it ever runs on a vehicle, and it is the cheapest test with reality in it. A drone's flight log (a PX4 ULog, a ROS 2 bag) holds every IMU sample, GPS fix and barometer reading with their timestamps; replaying it through a modified EKF shows at once how the new filter would have tracked that flight, including the gust and the GPS glitch that happened at minute seven.

It is open loop, and that is its limit. The recorded data was produced with the old software in control: the vehicle went where the old controller sent it. Replaying it through a new estimator answers what would this filter have believed exactly, but replaying it through a new controller cannot answer where would the vehicle have gone, because the new commands would have changed the trajectory and therefore every later measurement. For that the loop has to be closed in software-in-the-loop, against a simulator.

Replay is the right test for anything that only observes (estimators, fault detectors, anomaly models, perception) and the wrong one for anything that acts. A growing library of real logs, especially of the flights and shifts where something went wrong, becomes a regression suite no simulator can match for realism.

Timestamps make or break it. A log whose sensor times are the reception times rather than the acquisition times replays the wrong delays, and a replay tool that ignores timestamps and feeds samples as fast as it can tests a different system (time synchronization).

It compares candidates fairly: two estimators replayed on the same thousand flights see exactly the same data, which no field test can offer, and a reference such as RTK in the log gives a ground truth to score them against.

It also validates simulators: running the simulator with the logged commands and comparing its outputs with the logged sensors measures the sim-to-real gap.

Log replay is one of the core practices of verification, beside the X-in-the-loop testing pyramid and fault injection.