robotics//embodied AI

Embodied AI is artificial intelligence that perceives and acts through a body, a robot or an agent in a simulated world, and learns the relation between what it senses, what it does and what follows, and it is the branch that has to turn models into machines that walk, grasp, fly and survive in places nobody prepared for them. A model that can describe how a dog walks has learned about walking; a quadruped controller has to put each foot down at the right instant on a slope it has never seen, with motors that lag and a body that wobbles.


Embodied AI is artificial intelligence that perceives and acts through a body, a robot or an agent in a simulated world, and learns the relation between what it senses, what it does and what follows, and it is the branch that has to turn models into machines that walk, grasp, fly and survive in places nobody prepared for them. A model that can describe how a dog walks has learned about walking; a quadruped controller has to put each foot down at the right instant on a slope it has never seen, with motors that lag and a body that wobbles.

Its loop is the robot's loop with learning inside it: observe, act, receive the consequence, adapt. That shape changes what the data must be. A language model learns mostly from static data, text and images that already exist in enormous quantity. An embodied learner needs action data, trajectories of the form (st,at,rt,st+1)(s_t, a_t, r_t, s_{t+1})(st​,at​,rt​,st+1​): the state it saw, the action it took, what that earned and the state that came next. Only that kind of record contains the answer to what happens if I do this?, and it is scarce, because every real sample costs robot time, wear and sometimes a broken part.

The practical answers to that scarcity are known. Train millions of trials in simulation with physics and sensor parameters varied (domain randomization) and transfer the policy to hardware; learn from demonstrations by people or teleoperation (imitation learning); learn a predictive world model and practise inside it. Legged locomotion is the clearest success of the first route (learning-based control).

The body imposes deadlines and stakes that text does not. A decision late by a hundred milliseconds is a fall; an exploratory action is a dent or an injury, so learning on hardware runs behind safety limits and much of it happens before the robot ever moves. The gap between simulator and machine never fully closes (sim-to-real gap).

Not every reward is explicit. A robot can also learn by predicting its next observation, by imitating, or from logged interaction without a designed score, which matters because a reward for walk well is far harder to write than one for answer correctly (reward function).

Why the bodily skills are the hard part, when abstract ones fell first, is Moravec's paradox. Its place in the larger machine is in robotics: perception inward, motion planning and control outward.