control//learning-based control
Learning-based control is the family of controllers whose law, or the model it relies on, is learned from data or from trial and error instead of being derived from equations written by hand, and it is the frontier where machine learning meets loops that move real hardware. The members differ in how much they learn and where the optimization happens, and an engineer's question about any of them is the same: what guarantee comes with it, and what stands between it and the actuator.
Learning-based control is the family of controllers whose law, or the model it relies on, is learned from data or from trial and error instead of being derived from equations written by hand, and it is the frontier where machine learning meets loops that move real hardware. The members differ in how much they learn and where the optimization happens, and an engineer's question about any of them is the same: what guarantee comes with it, and what stands between it and the actuator.
The most ambitious member is reinforcement learning for control. The policy is the controller, the reward is a cost like the LQR's but with no restriction of form (contacts, discrete events, anything), and the optimization moves offline into training: where an MPC optimizes at every step, an RL policy only evaluates a network online, a few tenths of a millisecond on a Cortex-M7 for two hidden layers of 256 units. The price is millions of interaction steps, affordable only in simulation, so the policy learns the simulator with its defects (sim-to-real gap): no motor delay in the simulator gives a policy that oscillates on the robot, optimistic friction gives one that slips. Domain randomization is the standard answer. The big success story is legged locomotion: policies trained in simulation with randomization walk quadrupeds over rough terrain, GPU simulators train thousands of robots in parallel in minutes, and the results reach commercial robots (niche, heading to industry).
The network proposes; something simpler and verifiable disposes.
A learned controller in closed loop changes the states it will see next, so a small error leads to states it never trained on, where it errs more; and a network can have large local input-output slopes, which act as a high loop gain, the short road to instability. Deployed learners therefore run behind a safety filter and, in fleets, by staged rollout with rollback.
Less ambitious members carry less risk. A grey-box model with a learned residual lets physics give the structure and a network learn what is missing (aerodynamics, friction), with far less data and better extrapolation. A learned model plus an optimizer predicts each action's effect and lets an optimizer pick the cheapest safe one, with the local control system checking the orders (Google and DeepMind's data-centre cooling, every five minutes, not every millisecond). Learned structure with adapted coefficients (Neural-Fly, 2022) learns offline a basis for the wind's effect and adapts only a few linear coefficients in flight: slow learning of the form, fast adaptation of the numbers.
Imitation of an expert is the other way to get a policy without a reward, and it has the same compounding-error problem (imitation learning).
The real cost of RL is engineering weeks tuning the simulator, the randomization ranges and above all the reward, which is optimized exactly as written (reward hacking).
Maturity: legged locomotion is niche; drone racing (Swift beat world champions on a known course in 2023) and dexterous manipulation are research; RL in production process control is rare, because a chemical plant cannot be reset like a simulator. Foundation models for robots are research (controller design).