Physical AI vs. Embodied AI

This is a primer on the difference between physical AI and embodied AI.

Two terms describe the same machines on the same factory floor, and the tech industry uses them almost interchangeably. They are not synonyms. One comes from decades of cognitive science. The other arrived in 2024 with a keynote and a product line. The gap between them tells you who is selling what and what they want you to expect.

Table of contents

The short answer

Embodied AI is the academic research tradition, dating to the 1980s, which holds that intelligence requires a body and emerges from sensorimotor interaction with the world. “Physical AI” is the industry umbrella term popularized from 2024 onward for AI systems that perceive and act in the physical world. The two overlap heavily and remain distinct. Which term someone reaches for, and why, is the difference that matters.

This primer traces both terms to their sources, sets out what separates them, and reads the naming itself as a move in a market: where each word comes from, the breakthrough both point at, the reality behind the demo videos, and the cultural baggage that stops anyone from seeing these machines plainly. The short version: the words are a claim about the future more than a neutral description of the present.

What do physical AI and embodied AI mean?

Both terms point at machines that sense the world and act in it: robots, autonomous vehicles, drones, and humanoids. Where they part company is in what they emphasize and who says them.

Embodied AI

Embodied AI is the older and narrower term. In research it names the intelligence of an agent that has a body, a system that perceives, reasons, and acts through physical interaction, with cognition grounded in that interaction. The word carries a specific claim, which the next section unpacks: that a body is a condition for intelligence in the first place, the ground from which cognition grows from. Google DeepMind uses embodied AI as the banner term for its Gemini Robotics program, framing the work as a research mission to build robots that help with everyday tasks.1 Academic surveys, the IEEE, and the first international standard for the field (ITU-T F.748.66, ratified in December 2025)2 all keep the term.

Physical AI

Physical AI is the broader and newer label. In industry usage it covers any AI that deals with physics, motion, sensors, and real-world execution, which is wider than embodied AI. A predictive-maintenance system watching a turbine counts as physical AI with no body in any ordinary sense. The Boston Consulting Group, one of the few to draw the line explicitly, calls embodied AI the intelligence of an agent with a body, and physical AI the umbrella for AI that acts on the physical world.3 Embodied AI, on this reading, is a subset of physical AI.

Regulators use neither word. The EU AI Act defines an “AI system” in technology-neutral language and never mentions physical AI. The US NIST AI Risk Management Framework does the same, describing an engineered or machine-based system that generates outputs influencing real or virtual environments.4 To a regulator, a humanoid robot is a machine with an embedded AI system: the machinery falls under product safety law, and the AI component under AI law. The umbrella category does not exist in the statute books.

The boundary that matters: autonomy vs. automation

The analytical line sits elsewhere: between adaptive autonomy and rule-based automation. A robot vacuum on fixed rules (if the bump sensor fires, turn forty degrees) is automation, no matter how many sensors it carries. Physical AI begins where the fixed rules end and the machine generates its own behavior from what it perceives. The dividing criterion is whether the control policy is written out in advance or learned.

Where embodied AI comes from

The argument against the disembodied mind

Embodied AI began as a rebellion inside AI itself. Through the 1970s, the dominant approach treated intelligence as symbol manipulation: represent the world in a formal model, reason over the symbols, and output a plan. Robots built this way were slow and brittle, choking on the gap between their tidy internal model and the messy world. In 1991 Rodney Brooks, then at MIT, published “Intelligence without Representation” and turned the premise around. His argument: representation is the wrong unit for building intelligent systems, and explicit world models get in the way. Better, he wrote, to use the world as its own model.5

Behavior-based robotics

Brooks had spent the 1980s building it. The subsumption architecture he developed layered simple behaviors that coupled perception directly to action, with no central planner in the middle. Instead of computing a map and then moving, the robot simply reacted, and intelligent-looking behavior emerged from body, sensors, and environment working on each other. The insect-like machines his lab built crossed ground that defeated more “intelligent” designs. The lineage reaches further than robotics, to the phenomenology of Merleau-Ponty and Heidegger, and forward through Hubert Dreyfus, who argued that human intelligence is inseparable from body and situation.

The claim inside the name

That is the intellectual content the word “embodied” carries. Intelligence grows out of a body acting on a world, out of the sensorimotor loop that ties perception to action to consequence. Embodiment is the condition for intelligence, the ground it grows from. This is why the academic term is a claim: it says something contestable about what intelligence is. The irony is worth holding onto: when today’s industry adopts “embodied AI” for a foundation model dropped into a robot body, it takes the name and quietly drops the thesis. Whether today’s simulation-trained models satisfy that argument or just wear its vocabulary is a question few product announcements pause on.

Where physical AI comes from

A materials-science term, briefly

Physical AI also has an academic birth, though a short-lived one. In 2020 the researchers Aslan Miriyev and Mirko Kovač used “physical artificial intelligence” in Nature Machine Intelligence to name something specific: robots whose intelligence lives partly in their materials and morphology, bio-inspired systems in which a soft polymer that changes shape in response to heat performs a kind of computation in the body itself.6 That meaning barely survived. Within four years the term had been picked up, emptied, and refilled.

The rebrand of 2024

The refill came from NVIDIA. From mid-2024, CEO Jensen Huang began using “physical AI” as the name for the next wave of computing. At Computex in Taipei in June 2024, he called it the next wave of AI, AI that understands the laws of physics, and AI that can work among us.7 At CES in January 2025, he escalated it to a full era, following perception AI and generative AI, and launched the Cosmos platform to serve it.7 Bio-inspired materials had nothing to do with this version. Physical AI now meant foundation models that grasp physics in simulation and then drive robots, vehicles, and factories in the real world.

A term that builds a market

The naming was a market move, and the market read it as one. NVIDIA sells the chips, the data center compute, and the simulation stack (Omniverse for digital twins, Cosmos for world models, and Isaac for robot training). Leave robotics under the older, drier “embodied AI,” and it reads as a lab discipline. Call it “physical AI,” and it becomes an investable frontier, a single ecosystem in which hardware and software are sold together. The move rhymes with “cloud computing,” a phrase that made renting space on someone else’s servers sound like weather and helped conjure the market it named.

The money followed the word

Capital arrived on cue. Physical Intelligence, a robot-foundation-model startup, raised roughly 400 million dollars at a valuation of around two billion dollars in late 2024.8 Skild AI raised about 1.4 billion dollars in January 2026 at a valuation above 14 billion dollars.8 Neither number reflects current hardware revenue. They priced a bet that the term “physical AI” was built to make legible.

What actually differs

Strip away the marketing and a clean set of distinctions remains. The table sets them side by side. The paragraphs after it cover the shared technology, then physical AI versus robotics and versus generative AI.

Dimension Embodied AI Physical AI
Origin Cognitive science and robotics, 1980s NVIDIA marketing, 2024 (a 2020 materials-science coinage, redefined)
Primary community Academic researchers, DeepMind Chip and platform vendors, consultancies, startups
Body required Yes, by definition Not always (covers bodiless systems such as process control)
Scope Narrower: agents with bodies Broader: any AI that acts on the physical world
Core claim Intelligence needs a body A new computing market exists
Typical examples Humanoids, robot learning, service robots Robots, autonomous vehicles, smart factories, drones

The breakthrough both terms point at

For all the divergence in the words, they converge on one technology: vision-language-action (VLA) models, the robotics cousins of large language models. Where a language model takes text and predicts text, a VLA model takes camera images and a spoken instruction (“clear the table”) and predicts motor commands, turning what it sees and hears into joint movement. It generates behavior on the fly instead of following a script. Physical Intelligence built π⁰ (pi-zero) on this design. Figure built Helix.9 This is the actual advance under both banners, the reason 2024 felt like a threshold.

Physical AI vs. robotics

Robotics is the engineering discipline: the mechanics, actuators, control systems, and safety design that make a machine move. Physical AI names the model layer on top, the learned policy that decides what the machine does. A robot arm from 1985 is robotics with no physical AI in it. A humanoid running a VLA model is robotics plus physical AI. The hardware is old and necessary; the intelligence layer is what is new.

Physical AI vs. generative AI

They are the same model family aimed at different outputs. Generative AI produces tokens: words, pixels, and audio. Physical AI produces actions: torques, trajectories, grasps. Under the hood, both are large models trained to predict the next thing in a sequence, but the output space differs, and that is everything. A wrong token just gives you a bad sentence. A wrong torque can tip a heavy machine onto whoever is standing next to it. The physical world does not tolerate hallucination the way a chat window does.

What the terms do

Naming as expectation management

Words about new technology are rarely neutral: they set expectations that move money and shape policy before any machine proves itself. “Physical AI” does this work deliberately. It bundles diverse activities under one heading, declares them a wave, and invites everyone to line up. The sociologist Jens Beckert calls these fictional expectations: shared images of the future that coordinate investment and action now, whether or not they come true. A term that names a market is one of the purest examples, a fictional expectation doing economic work.

Who benefits from which word

The choice of word tracks the interests of the speaker with unusual precision.

  • Vendors selling the full stack say “physical AI.” For NVIDIA, the consultancies, and hardware-and-model startups, the umbrella term is the product boundary. It lets them sell chips, simulations, and foundation models as one growth story and manufacture urgency at board level.
  • Researchers keep “embodied AI.” For academic labs and, selectively, DeepMind, the older term signals scientific lineage and a connection to the cognitive-science tradition. It signals a research lineage, a claim to scientific standing.
  • Regulators and standards bodies avoid both. They say “AI system” and “autonomous machine” because a neutral, technology-agnostic term is what a legal definition needs.

The term asserts its own reality

A term like physical AI works as a sociotechnical imaginary: a shared vision of a desirable technological future that helps summon the resources to build it. It says the future is robotic, the wave is here, and the only question is how fast you climb aboard. Where an expectation coordinates investment, the imaginary asserts a whole reality into being. Reading the vocabulary as an argument, an assertion about the future, is the first defense against being sold that future as already arrived.

The reality check

The theater of proof

Between the announcement and the deployment sits a genre of performance best described with a phrase from Bruno Latour: the theater of proof.10 A humanoid breaks a running record, dances in sync with fifteen others on state television, or serves drinks at a launch. It looks like autonomy. Often it is choreography, or a person off-camera with a controller. At Tesla’s “We, Robot” event in October 2024, Optimus robots poured drinks and chatted with guests. Observers, and later the company, confirmed that the interactions were teleoperated, run by human operators in the background.11 The walking was automated. The conversation was a puppet show.

Teleoperation has a legitimate use

Remote control has a legitimate purpose. Modern VLA models learn from demonstration, and the cheapest source is a human piloting the robot through the motion. Teleoperation is the front end of the training pipeline, and the camera-ready demo is a byproduct. It becomes deception only when a staged video is sold as autonomy it does not have. The honest signal is an uncut run: command entered, real-time clock, visible fumbles, autonomous recovery. Cut footage proves nothing.

Where the real work happens

The genuine deployments convince precisely because they are unspectacular. The Digit robot from Agility Robotics moves totes onto conveyors in a GXO warehouse and has passed 100,000 totes in live operation, though its sustained rate runs well below the roughly 66 totes per hour it hits in a trade-show demo.12 Figure’s humanoid spent about ten months on a body-shop line at BMW’s Spartanburg plant, placing sheet-metal parts into welding fixtures for more than 30,000 vehicles before it was retired.13 Read carefully; that is a narrow achievement: two robots on one task in a defined pilot.

The two bottlenecks

Two limits keep recurring.

  • The interaction gap. Films trained everyone to expect flawless machines. Real ones hesitate, take a few seconds to plan, and sometimes grab at nothing. Human-robot trust research finds that a single visible fault drops trust sharply and that trust falls faster than it builds, which puts a psychological ceiling on adoption that no press release mentions.14
  • The data bottleneck. Language models had the internet to train on. Robots have no equivalent corpus of physical experience, and Rodney Brooks argues that watching video cannot supply it, because dexterity depends on force and touch that a camera never records.15 The shortage of real-world manipulation data, more than raw compute, is the binding constraint.

Three regional bets

The contest plays out along three regional lines. The United States leads on foundation models and concentrates capital, with a protectionist push to keep the hardware supply chain at home. China writes embodied intelligence into its fifteenth Five-Year Plan, drives mass production through supply chains built for electric vehicles, and stages spectacles (dancing troupes and a robot half-marathon) as demonstrations of reach.16 Europe takes the regulated path: quieter on stage, present in industrial pilots, and organized around risk assessment under the AI Act. Different bets, one field.

The cultural baggage

Almost no one has worked alongside an advanced humanoid robot. Into that experience gap the mind pours everything it already carries, and what it carries is centuries of stories. We meet them through inherited images, and the images are older than the microchip.

The long memory of the artificial being

The fear of a created thing slipping its maker’s control runs deep. It is in the Golem of sixteenth-century Prague, clay that obeys until it does not. It is in Mary Shelley’s Frankenstein, where the fault lies with the creator who abandons his responsibility, and the creature pays for it. It surfaces in Fritz Lang’s Metropolis (1927), whose robot Maria is the first iconic robot on screen and a female-coded one, fusing the fear of autonomous technology with the fear of female sexuality, an echo still heard when AI assistants are given women’s voices.

The Asimov misreading

No misremembering matters more than Isaac Asimov’s. His Three Laws of Robotics, from 1950, get cited in today’s AI-alignment debates as though a science-fiction writer had drafted a working safety protocol seventy years early, ready to be typed into code. Read the stories, and the opposite is true. Asimov built the Laws precisely to break them. Every tale turns on the Laws colliding, generating loopholes, and producing outcomes their designers never intended. They are literature about the inevitable failure of formal rules to govern behavior in a messy world.17 Citing them as a solved control problem inverts the point of the fiction.

The uncanny valley is real

One piece of the cultural response has hard evidence behind it. In 1970 the roboticist Masahiro Mori proposed the uncanny valley: as a robot grows more humanlike, our affinity for it rises, then plunges into revulsion just short of the real thing before recovering.18 For decades this was a striking anecdote. A 2022 meta-analysis by Diel, Weigelt, and MacDorman settled it, pooling 72 studies and 247 effect sizes for a large effect (Hedges g of 1.01).19 The revulsion is a strong, measurable reaction, strongest when a near-human face carries a small wrongness the brain reads as disease.

The whiteness of the picture

How AI is pictured is as loaded as how it is imagined. Stock photography and press imagery converge on a narrow vocabulary: a glossy white android, a glowing blue brain, a Terminator skull, a Michelangelo hand reaching toward a spark. Stephen Cave and Kanta Dihal named part of this the whiteness of AI: the reflex of depicting intelligent machines as white encodes old assumptions about power and perfection.20 The Better Images of AI initiative, launched in 2021, exists to break the cliché, offering pictures that show human labor and real systems in place of chrome messiahs.21 The stock image carries a load: it sets an expectation of flawlessness a stumbling real robot cannot meet, deepening the fall into the uncanny valley.

The Japan myth

The most persistent cultural story is that Japan, steeped in Shinto animism, simply loves robots where the West fears them. The evidence does not support it. On the Negative Attitudes toward Robots Scale, Japanese respondents report equal or higher explicit anxiety than Western samples, especially about the social impact of robots, which turns the cliché on its head.22 Implicit-association tests point the same way: at the millisecond level, people across cultures prefer humans to machines, with no special Japanese warmth. The affinity story is better explained by industrial policy than by spirituality: Japan and China face demographic decline and labor shortage, resist large-scale immigration, and need automation to hold their economies together. Repeating the friendly-robot narrative in the West as cultural essence is a form of techno-orientalism.

Design as damage control

The companies have read this research. Agility’s Digit, one of the more successful industrial humanoids, has no face at all, only a sensor block where a head would be, so the brain files it as a tool, and the uncanny valley never opens. 1X dressed its home robot Neo in a soft knit suit, trading the cold Terminator surface for something domestic and vulnerable. These are deliberate, empirically informed attempts to route around inherited fear.

The lag runs backward

William F. Ogburn’s cultural lag model, from 1922, fits most technologies: invention races ahead while laws, norms, and institutions scramble to catch up.23 With the intelligent machine the sequence ran the other way. The images stood ready centuries before the hardware, the fear arrived pre-loaded, and the imaginary even steered what got built, as the faceless Digit and soft-suited Neo just showed. This garden explores that reversal as cultural lead: culture leading, technology following its script.

All of it is the present talking about the future, where the practitioner reading picks up.

A practitioner’s view

I work as a critical futures researcher, and the reason the terminology holds my attention is that “physical AI” is, at bottom, a present future: a claim made in the present about a future that has not arrived, dressed as a report on the present tense. So I analyze it the way I analyze any statement about the future, by asking what it is doing here and now rather than whether it will come true.

The distinction I lean on comes from the sociologist Niklas Luhmann, who separated present futures from future presents. A future present is the actual state of the world on some later date, the messy factory floor of 2032 as it will really be. A present future is the image of that moment held today in a keynote or a pitch deck. The two differ, and the whole task is to keep them apart.

Two things follow.

The term “physical AI” is one such present future. Where an earlier section read it as market-making, the futures point is narrower: when NVIDIA says the era is here, it uses the present tense to make that present future feel like a present fact, and naming the wave is part of building it.

Our reactions to these machines run through inherited images of the future as well. The Golem, Frankenstein, the Terminator, the tidy Asimov Laws are present futures too, pictures we carry that shape what we expect before any direct experience. The cultural section of this primer maps that lens, which is part of the technical story.

So the work is to separate the staged future from the one that actually arrives. The dancing robots and the viral demos are present futures, engineered for effect. The 100,000 totes moved in a warehouse and the sheet-metal parts placed at BMW are future presents leaking through, unglamorous and real.

For all that, physical AI has substance. The VLA breakthrough is genuine, and some of these machines do real work. But the word arrived years ahead of the reality it names, which is the normal condition of a technology sold as a future. My job, and the reason I keep foresight and forecasting apart, is to tell you which one you are looking at. Tell me how you talk about physical AI, and I will tell you what you expect from the present.

For the futures-studies frame that runs under this primer, start with Foresight vs. Forecasting and the distinction between present futures and future presents. On why images of the future are never neutral, see No future is neutral, Sociotechnical Imaginaries, and Fictional Expectations. For methods that test a strategy against several futures rather than one forecast, see Scenario Planning and Windtunneling. On the governance side, see Anticipatory Governance.

For the primary sources, Rodney Brooks wrote “Intelligence without Representation,” the origin text for embodied AI, and his 2025 essay on humanoid dexterity is, to my reading, the sharpest technical corrective to the current hype.515 For the cultural argument, Cave and Dihal on the whiteness of AI and the 2022 uncanny-valley meta-analysis are the load-bearing references.2019

  1. Google DeepMind (2025). “Gemini Robotics brings AI into the physical world.” (DeepMind

  2. ITU-T Recommendation F.748.66 (12/2025). “Requirements and framework for embodied artificial intelligence systems.” Approved 14 December 2025. (ITU

  3. Boston Consulting Group (2026). “What Is Physical AI?” (BCG) Distinguishes embodied AI from physical AI, the broader umbrella. 

  4. EU AI Act, Article 3(1). (artificialintelligenceact.eu) Defines “AI system” technology-neutrally, with no category “physical AI.” See also NIST (2023). AI Risk Management Framework (AI RMF 1.0). (NIST) which also avoids the term. 

  5. Brooks, R. A. (1991). “Intelligence without Representation.” Artificial Intelligence, 47(1-3), 139-159. (DOI) The founding argument that intelligence emerges from bodily interaction and that “the world is its own best model.”  2

  6. Miriyev, A., & Kovač, M. (2020). “Skills for physical artificial intelligence.” Nature Machine Intelligence, 2, 658-660. (DOI) The academic coinage of “physical artificial intelligence,” on bio-inspired materials and morphological computation. 

  7. NVIDIA (2024). Jensen Huang COMPUTEX 2024 keynote. (NVIDIA) Source of Huang’s “the next wave of AI is physical AI” (Computex, June 2024). And NVIDIA (2025), CES keynote. (NVIDIA) Where Huang frames physical AI as the next era and launches Cosmos (January 2025).  2

  8. Reuters (2024). “Robot AI startup Physical Intelligence raises $400 mln from Bezos, OpenAI.” (Reuters) And TechCrunch (2026). “Robotic software maker Skild AI hits $14B valuation.” (TechCrunch 2

  9. Physical Intelligence (2024). “π0: A Vision-Language-Action Flow Model for General Robot Control.” (Physical Intelligence) And Figure (2025). “Helix: A Vision-Language-Action Model for Generalist Humanoid Control.” (Figure

  10. The phrase is Bruno Latour’s “théâtre de la preuve,” from his account of Louis Pasteur staging public demonstrations to make proof persuasive. Latour, B. (1988). The Pasteurization of France. Harvard University Press. (Wikipedia

  11. Cybernews (2024). “Tesla Optimus bots actually controlled by humans during the We, Robot event.” (Cybernews) The October 2024 interactions were teleoperated. 

  12. Agility Robotics (2024). “Digit Deployed at GXO in Historic Humanoid RaaS Agreement.” (Agility Robotics) The ~66 totes/hour figure is a demo number; sustained real-world throughput runs materially lower. 

  13. BMW Group (2025). “Successful test of humanoid robots at BMW Group Plant Spartanburg.” (BMW Group) The Figure 02 pilot placed sheet-metal parts over roughly ten months, supporting 30,000+ vehicles. 

  14. Salem, M., Lakatos, G., Amirabdollahian, F., & Dautenhahn, K. (2015). “Would You Trust a (Faulty) Robot? Effects of Error, Task Type and Personality on Human-Robot Cooperation and Trust.” Proceedings of HRI 2015, 141-148. (DOI) A faulty robot lowered rated trust. See also Hancock, P. A., et al. (2011). “A Meta-Analysis of Factors Affecting Trust in Human-Robot Interaction.” Human Factors, 53(5), 517-527. (DOI

  15. Brooks, R. A. (2025). “Why Today’s Humanoids Won’t Learn Dexterity.” (rodneybrooks.com) Argues video imitation cannot teach dexterity, which depends on force and touch data a camera never captures.  2

  16. MERICS (2026). “Embodied AI: China’s ambitious path to transform its robotics industry.” (MERICS

  17. Vinsel, L. / Brookings (2015). “Isaac Asimov’s Laws of Robotics Are Wrong.” (Brookings) On how Asimov’s stories dramatize the Laws failing rather than working as a safeguard. Primary source: Asimov, I. (1950). I, Robot

  18. Mori, M. (1970/2012). “The Uncanny Valley.” Translated by K. F. MacDorman and N. Kageki, IEEE Robotics & Automation Magazine, 19(2), 98-100. (DOI

  19. Diel, A., Weigelt, S., & MacDorman, K. F. (2022). “A Meta-analysis of the Uncanny Valley’s Independent and Dependent Variables.” ACM Transactions on Human-Robot Interaction, 11(1), 1-33. (DOI) 72 studies, 247 effect sizes, Hedges g = 1.01 [0.80, 1.22].  2

  20. Cave, S., & Dihal, K. (2020). “The Whiteness of AI.” Philosophy & Technology, 33(4), 685-703. (DOI) On the reflex of depicting AI as White.  2

  21. Better Images of AI. (betterimagesofai.org) A non-profit launched in 2021 (We and AI, with the Leverhulme Centre for the Future of Intelligence). 

  22. MacDorman, K. F., Vasudevan, S. K., & Ho, C.-C. (2009). “Does Japan really have robot mania? Comparing attitudes by implicit and explicit measures.” AI & Society, 23(4), 485-510. (DOI) Finds Japanese explicit attitudes no more positive than US, with no implicit difference: both groups prefer humans to robots. 

  23. Ogburn, W. F. (1922). Social Change with Respect to Culture and Original Nature. New York: B. W. Huebsch. (archive.org) The origin of cultural lag: material culture outpaces the norms meant to govern it. 

No notes link to this note yet.


Note Graph

ESC