Language Is Not Enough: The Researchers Trying to Give AI a Body
There is a joke among robotics researchers that goes: language models know everything about the world except how to exist in it. A large language model trained on the entire internet knows that a coffee cup is a cylinder that can be grasped around its circumference, that it is fragile, that it typically sits on flat surfaces, that it becomes hot when filled with hot liquid. What it does not know: what no amount of text training can teach: is what it feels like to pick one up. The weight distribution, the compliance of the handle, the way it pivots unpredictably if you grasp it slightly off-center. That knowledge lives in the body, not in language.
This gap: between knowing about things and knowing how to interact with them: is what the field of embodied AI is trying to close. The premise is that truly general intelligence requires not just the ability to process and generate language but the ability to learn through physical interaction with a real, three-dimensional, unpredictable world. And the researchers pursuing this premise are starting to produce results that suggest the approach may be right.
The case for embodiment in intelligence is both philosophical and practical. Philosophically, a significant current in cognitive science holds that human intelligence is not just brain computation but brain-body-world computation: that our concepts, our reasoning, and our language are all shaped by the experience of having a body that moves through space, manipulates objects, and interacts with other bodies. On this view, the disembodied intelligence of a language model is not a simplification of human intelligence but a fundamentally different kind of thing, with capabilities and limitations that differ in ways that matter.
Practically, the limitations of disembodied language models are visible in any task that requires physical common sense. Ask GPT-5 to explain how to change a tire and it will produce a clear, accurate, step-by-step explanation. Ask it to actually change the tire: to control a robot doing the task: and the gap between verbal knowledge and physical capability becomes immediately apparent. Grasping, manipulating, navigating: these require a different kind of knowledge than language models acquire from text.
Several distinct approaches to embodied AI are active in 2025 and 2026, representing different intuitions about the path to physical intelligence.
Simulation-to-real transfer: training AI systems in physics simulation environments and then deploying them in real robots: has become the dominant paradigm for robot learning at scale. Simulated environments allow training at speeds impossible in the physical world: a simulated robot can experience thousands of hours of interaction in the time it would take a physical robot to experience one. The challenge is the simulation-to-reality gap: real environments have physical properties that simulations approximate imperfectly, and models trained in simulation often fail in the real world on the exact edge cases that simulation did not accurately represent. Closing this gap is a major active research area.
Please enable JavaScript to read the full article.