Benchmark / simulation

LIBERO

LIBERO separates manipulation evaluation into spatial, object, goal and longer-horizon task suites, allowing transfer behavior to be examined along different axes.

Source contextPrimary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Reviewed recordLast reviewed 6 sources ↗6 sourced fields

What the protocol tests

Task definitions and initial states are released with demonstration data. The commonly reported Long suite should be identified explicitly rather than confused with a training collection or an all-task aggregate.

Source contextPrimary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Reading results

A valid result retains training split, camera and state inputs, checkpoint selection, episode budget and environment revision. Fine-tuned results and zero-shot results answer different questions.

Source contextPrimary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Methodology and reuse limits

Simulation task success does not establish physical reliability. Different suite protocols must not be combined into a global robot score.

Source contextPrimary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Evaluation protocol

Exact version / scope
Original LIBERO suites
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Environment
simulation
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Metric definition
Task success rate by suite
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Task scope
Not publicly disclosed
Embodiment
Simulated Franka arm
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Access & reuse

Access model
Public research documentation and evaluation resources
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

License
Not publicly disclosed

Additional disclosed fields

Category
Robot evaluation benchmark
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Reported benchmark results

Context before scores. Results from different benchmarks are not directly comparable. A simulation result does not establish real-world reliability, safety or commercial availability.
Model / revisionBenchmark / metricReported resultEnvironmentEvidence & scope
OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Spatial97.6Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Object98.4Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Goal97.9Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Long94.5Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

What this evidence does not establish

  • Simulation task success does not establish physical reliability. Different suite protocols must not be combined into a global robot score.

Relationships & deployments

Related reading is an editorial crosslink. Sourced connections describe relationships reported in the cited material. A link to a versioned profile does not establish compatibility with that version unless the connection note explicitly identifies it.

uses
LIBERO-Plus → LIBERO
Sources 1
sylvestf/LIBERO-plus official repository

sylvestf · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Connection reviewed 2026-09-28

uses
LIBERO-Pro → LIBERO
Sources 1
Zxy-MLlab/LIBERO-PRO official repository

Zxy-MLlab · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Connection reviewed 2026-09-28

uses
RoboCerebra → LIBERO
Sources 1
qiuboxiang/RoboCerebra official repository

qiuboxiang · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Connection reviewed 2026-09-28

evaluated on
OpenVLA-OFT → LIBERO
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Connection reviewed 2026-09-28

evaluated on
TS-Mask VLA → LIBERO
Sources 1
TS-Mask VLA original paper

Shengzhuo Yang and coauthors · Published 2026-07-10 · Source accessed 2026-09-28

Connection reviewed 2026-09-28

Curated comparisons

Record history & verified changes

A research review records when we checked a source. It does not mark a product launch or a new deployment.

Initial reviewed record. No subsequent field change has been recorded.

Inspect the evidence

Sources & evidence

md-libero
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Repository · Publication date not disclosed · Source accessed 2026-09-28

License: MIT code; CC BY 4.0 dataset. Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

md-liberoplus
sylvestf/LIBERO-plus official repository

sylvestf · Repository · Publication date not disclosed · Source accessed 2026-09-28

Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

md-liberopro
Zxy-MLlab/LIBERO-PRO official repository

Zxy-MLlab · Repository · Publication date not disclosed · Source accessed 2026-09-28

License: MIT code. Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

md-robocerebra
qiuboxiang/RoboCerebra official repository

qiuboxiang · Repository · Publication date not disclosed · Source accessed 2026-09-28

License: Apache-2.0 code. Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

md-tsmask
TS-Mask VLA original paper

Shengzhuo Yang and coauthors · Paper · Published 2026-07-10 · Source accessed 2026-09-28

Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.