control//learning-based control//domain randomization
Domain randomization is a training technique in which the parameters of a simulator are drawn at random in every episode, so that a learned policy or model must work across a whole family of simulated worlds and therefore has a better chance of working in the real one. It is the standard bridge across the sim-to-real gap for learning-based controllers: if the policy has seen masses from 0.8 to 1.5 kg, motor constants within ±15 %, delays up to 30 ms and sensor noise of several levels, the real robot looks like one more sample.
Domain randomization is a training technique in which the parameters of a simulator are drawn at random in every episode, so that a learned policy or model must work across a whole family of simulated worlds and therefore has a better chance of working in the real one. It is the standard bridge across the sim-to-real gap for learning-based controllers: if the policy has seen masses from 0.8 to 1.5 kg, motor constants within ±15 %, delays up to 30 ms and sensor noise of several levels, the real robot looks like one more sample.
What gets randomized follows what the simulator gets wrong: masses and inertias, friction, motor constants and delays, sensor noise and latency, sometimes terrain or lighting. A policy that does well on all of them cannot rely on any particular value. It is completed by two habits. Identification of the real robot centres the ranges, so that the family contains reality instead of surrounding it at random (system identification). A learned actuator model, a network trained on data from the real motors that imitates their response, replaces the simulator's idealized actuator; that was one of the keys to the legged-locomotion results.
Domain randomization is robust control by sampling.
The ranges you randomize are the uncertainty set of robust control, and the trade is the same: narrow ranges give a sharp policy that fails in the real world, wide ranges give a conservative one. The difference is the guarantee, which is empirical: the policy works on the samples it saw.
It costs conservatism. A policy that must work for every mass in the range behaves as if it did not know which one it has; adding the robot's recent history to the policy's input lets it infer the current parameters and act less cautiously, an implicit form of adaptation.
It costs compute, though less than engineering. Randomization multiplies the experience needed; GPU simulators running thousands of robots in parallel make that affordable, while choosing the ranges remains the slow, manual part.
It does not cover what the simulator cannot represent. A mechanism missing from the model (a cable that snags, a sensor that saturates in sunlight) is not randomized into existence, which is why real-world tests and a safety filter still close the loop.
The engineering ancestor is the Monte Carlo robustness test, which samples plant variants to check a fixed controller (Monte Carlo method); randomization samples them to train one.