Research guide / Concepts

Physical AI vs Embodied AI

The terms overlap, but they are more useful when attached to a specific research question or operating system. Compare embodiments, interfaces and evidence instead of treating terminology as a capability ranking.

Source contextPrimary specReviewed
Sources 5
google-deepmind/open_x_embodiment official repository

google-deepmind · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

RT-2 research project

Google DeepMind · Publication date not disclosed · Source accessed 2026-09-28

mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

openvla/openvla official repository

openvla · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Gemini Robotics model family

Google DeepMind · Publication date not disclosed · Source accessed 2026-09-28

Reviewed recordLast reviewed 5 sources ↗3 sourced fields
Identity boundary. Original PhysicalAI.best editorial guide. Source links support factual examples; evaluation questions and decision rules are editorial guidance, not upstream endorsements.

Use a working distinction

In this guide, embodied AI emphasizes an agent whose observations and actions depend on its body and environment. Physical AI emphasizes the route from learned intelligence to physical-world systems and operation. These are editorial lenses with substantial overlap. Neither phrase alone establishes that a system has hardware, generalizes to a new body, runs without assistance or can be purchased. The record should state those properties directly.

Source contextPrimary specReviewed
Sources 2
RT-2 research project

Google DeepMind · Publication date not disclosed · Source accessed 2026-09-28

google-deepmind/open_x_embodiment official repository

google-deepmind · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Embodiment is a concrete interface

Describe the cameras, state inputs, action representation, actuators and workspace. A policy trained for one manipulator may require different data or control integration for another. Open X-Embodiment provides a useful reference for thinking about multiple robot datasets; its existence does not make all robots interchangeable. Ask which observations and actions the proposed implementation expects and whether those match the actual equipment.

Source contextPrimary specReviewed
Sources 2
google-deepmind/open_x_embodiment official repository

google-deepmind · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

openvla/openvla official repository

openvla · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

A simulated body still needs its environment label

An embodied task can be evaluated in simulation. CALVIN provides a language-conditioned manipulation environment and sequence protocol, which makes it valuable research evidence. Labeling the agent embodied does not turn those scores into measurements from physical hardware. Keep the simulator, scene split and robot setup alongside the result before comparing it with another implementation.

Source contextPrimary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Do not infer a development method from a label

The same marketing term can cover different model architectures and integration approaches. RT-2 explores action-token prediction, while other systems expose separate planning and action components. Review the actual input/output interface, adaptation requirements and access conditions. A useful technical comparison begins at these implementation choices rather than arguing that one vocabulary term is inherently more advanced.

Source contextPrimary specReviewed
Sources 2
RT-2 research project

Google DeepMind · Publication date not disclosed · Source accessed 2026-09-28

Gemini Robotics model family

Google DeepMind · Publication date not disclosed · Source accessed 2026-09-28

Choose the questions that match the decision

A researcher may need reproducible training and transfer experiments; a buyer may need reliable task completion and service support. Record both when they exist, without substituting one for the other. For a proposed system, ask what embodiment was tested, what changed at deployment, what human help remains and what evidence covers the target environment. These questions stay useful even when terminology changes.

Source contextPrimary specReviewed
Sources 2
openvla/openvla official repository

openvla · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Additional disclosed fields

Cross-embodiment example
Open X-Embodiment organizes robot datasets from multiple embodiments
Primary specReviewed
Sources 1
google-deepmind/open_x_embodiment official repository

google-deepmind · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Language-action example
RT-2 studies action tokens in vision-language robot control
Primary specReviewed
Sources 1
RT-2 research project

Google DeepMind · Publication date not disclosed · Source accessed 2026-09-28

Evaluation example
CALVIN evaluates language-conditioned manipulation sequences in simulation
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

What this evidence does not establish

  • The terminology is not used as a maturity scale.
  • Cross-embodiment research does not establish plug-and-play transfer to arbitrary hardware.

Relationships & deployments

Related reading is an editorial crosslink. Sourced connections describe relationships reported in the cited material. A link to a versioned profile does not establish compatibility with that version unless the connection note explicitly identifies it.

Record history & verified changes

A research review records when we checked a source. It does not mark a product launch or a new deployment.

Initial reviewed record. No subsequent field change has been recorded.

Inspect the evidence

Sources & evidence

md-openx
google-deepmind/open_x_embodiment official repository

google-deepmind · Repository · Publication date not disclosed · Source accessed 2026-09-28

License: Apache-2.0 repository software; CC BY 4.0 other repository content; constituent dataset licenses separate. Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

md-gemini-models
Gemini Robotics model family

Google DeepMind · Official spec · Publication date not disclosed · Source accessed 2026-09-28

Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.