ML//world model
A world model is a model that represents the state of an environment and how that state evolves, especially in response to the agent's own actions, and it is used to predict consequences before acting: to plan by imagining several futures, to train a policy on imagined experience instead of costly real trials, or to notice when reality departs from expectation. Its question is always *if I do this, what happens next?*
A world model is a model that represents the state of an environment and how that state evolves, especially in response to the agent's own actions, and it is used to predict consequences before acting: to plan by imagining several futures, to train a policy on imagined experience instead of costly real trials, or to notice when reality departs from expectation. Its question is always if I do this, what happens next?
Control engineers have always worked with one, under another name. The plant model inside an MPC predicts the next seconds of a drone or a distillation column for each candidate input sequence, and the controller picks the best; that model is written from physics and fitted by system identification. Machine learning's world models do the same job with a network learned from data, often from camera images, so they can cover what no one wrote equations for: how a pile of boxes topples when pushed, how a crowd moves around a robot.
Predictions degrade with the horizon. A small error each step compounds over a long imagined rollout, as dead reckoning drifts without a fix, so planners keep their imagined horizons short and correct with fresh observations.
A planner exploits its model's mistakes. If the learned model wrongly believes a manoeuvre is safe, an optimizer will find that manoeuvre, exactly as a policy trained in a flawed simulator learns its flaws (sim-to-real gap). A world model is only as trustworthy as the region of experience it was trained on.
It is distinct from a store of facts. A knowledge graph can state that a valve feeds a tank; a world model has to predict what the level does in the next minute if the valve opens halfway. Facts can be one of its inputs, dynamics are its substance.
Whether large language models carry an implicit world model is debated: they predict text about the world very well, and they still fail at simple physical prediction in ways a model of dynamics would not. That gap is one reason embodied AI learns from interaction, where every action returns a consequence to learn from.