Models
Inspect model families, revisions, access conditions and evaluation context before choosing a policy.
ACT
ACT predicts chunks of robot actions from demonstrations, with the original project focused on fine-grained bimanual manipulation.
Action ControlNet (ACNet)
ACNet is a research adapter for reducing discontinuities when robot policies predict action chunks asynchronously from aging observations.
Diffusion Policy
Diffusion Policy formulates visuomotor action generation as conditional denoising, with reference experiments and training artifacts available for inspection.
Field Foundation Models
Field AI’s model family underpins an autonomy software offering for robot operation in variable field environments.
Figure Helix
Figure’s vision-language-action model family for its humanoids, with an original upper-body release and a separately announced whole-body Helix 02 generation.
Gemini Robotics 2
Gemini Robotics 2 is the vision-language-action member of Google DeepMind’s Robotics 2 release, with access and control responsibilities distinct from the other family members.
Gemini Robotics ER 2
Gemini Robotics ER 2 is the embodied reasoning vlm member of Google DeepMind’s Robotics 2 release, with access and control responsibilities distinct from the other family members.
Gemini Robotics On-Device 2
Gemini Robotics On-Device 2 is the on-device vision-language-action member of Google DeepMind’s Robotics 2 release, with access and control responsibilities distinct from the other family members.
GEN-1
Generalist AI’s GEN-1 is a multimodal robotic model and inference system described around real-time actions and adaptation to specific robot tasks.
NVIDIA Cosmos Predict
NVIDIA Cosmos Predict is a specialized Cosmos release with a defined video world model role; the inspected repository now directs future development to Cosmos 3.
NVIDIA Cosmos Reason
NVIDIA Cosmos Reason is a specialized Cosmos release with a defined physical reasoning vlm role; the inspected repository now directs future development to Cosmos 3.
NVIDIA Cosmos Transfer
NVIDIA Cosmos Transfer is a specialized Cosmos release with a defined controllable world generation role; the inspected repository now directs future development to Cosmos 3.
NVIDIA GR00T N2
A previewed world-action model in NVIDIA’s robotics roadmap, identified separately from the downloadable GR00T N1.7 release.
NVIDIA Isaac GR00T N1.7
An open robot foundation model release with multimodal reasoning, a continuous-action diffusion head and documented fine-tuning paths for supported embodiments.
Octo
Octo is a generalist robot policy designed for adaptation to different observation and action spaces, with language or goal-image conditioning.
OpenHelix
OpenHelix is an independent research implementation exploring a dual-system VLA architecture inspired by Figure’s Helix work.
OpenVLA
OpenVLA is a research VLA with public weights and adaptation tooling, connecting visual observations and language instructions to manipulation actions.
OpenVLA-OFT
OpenVLA-OFT is an optimization and fine-tuning recipe that changes action decoding and representation for OpenVLA, with a separately documented bimanual extension.
Physical Intelligence π0
A flow-matching VLA that starts from a pretrained vision-language model and learns continuous robot control across multiple robot configurations.
Physical Intelligence π0.5
A generalist VLA that combines heterogeneous robot data, semantic prediction and web-derived supervision to study manipulation beyond its training locations.
Physical Intelligence π0.6
A hierarchical VLA described in Physical Intelligence’s model card, with semantic subtask prediction and an action expert built around a Gemma 3 backbone.
Physical Intelligence π0.7
A steerable generalist robotic foundation model documented in a named Physical Intelligence research paper.
RT-1
RT-1 connects camera observations and language commands to robot actions through a transformer policy and compressed visual tokens.
RT-2
RT-2 studies transferring pretrained vision-language representations into robot control by treating robot action tokens as another output language.
Skild Brain
Skild AI’s general-purpose robotics model family, organized around high-level behavior and lower-level control across different robot forms.
SmolVLA
SmolVLA is a compact base VLA exposed through LeRobot, with continuous action generation and a documented path to task-specific fine-tuning.
TS-Mask VLA
TS-Mask VLA is a manipulation research framework that models temporal and spatial structure in discrete action sequences.
VLA-Adapter
VLA-Adapter studies a lightweight connection between a visual-language backbone and action generation, with a public implementation and adaptation examples.
Wayve AI Driver
Wayve’s AI Driver is an embedded driving software offering designed to be licensed to vehicle manufacturers and integrated with vehicle sensors and compute.
Wayve GAIA
GAIA is Wayve’s world-model family for driving simulation; the reviewed GAIA-4 release constructs closed-loop scenarios from recorded sensor data.
X-VLA
X-VLA uses embodiment-related soft prompts to adapt a flow-based action policy across differing robot data sources.