LIBERO separates manipulation evaluation into spatial, object, goal and longer-horizon task suites, allowing transfer behavior to be examined along different axes.
Task definitions and initial states are released with demonstration data. The commonly reported Long suite should be identified explicitly rather than confused with a training collection or an all-task aggregate.
Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28
Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.
Reading results
A valid result retains training split, camera and state inputs, checkpoint selection, episode budget and environment revision. Fine-tuned results and zero-shot results answer different questions.
Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28
Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.
Reported benchmark results
Context before scores. Results from different benchmarks are not directly comparable. A simulation result does not establish real-world reliability, safety or commercial availability.
Model / revision
Benchmark / metric
Reported result
Environment
Evidence & scope
OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioception
Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28
Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.
Methodology & comparability
Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.
Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28
Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.
Methodology & comparability
Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.
Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28
Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.
Methodology & comparability
Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.
Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28
Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.
Methodology & comparability
Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.
Related reading is an editorial crosslink. Sourced connections describe relationships reported in the cited material. A link to a versioned profile does not establish compatibility with that version unless the connection note explicitly identifies it.