Curated comparison · Evaluation protocol design

LIBERO vs CALVIN for VLA evaluation

LIBERO and CALVIN organize robot-learning evaluation around different task and sequence protocols. Choose the benchmark that answers your research question, keeping their metrics and environments separate.

Source contextPrimary specReviewed
Sources 2
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Reviewed 7 decision dimensionsNo universal winner
Version and date boundaries matter. The exact revision below defines each column. A deployment of an earlier revision is not evidence that a newer revision is in production. Missing fields remain unknown.

On small screens, swipe the table horizontally. Each row carries its own source and review date.

Decision dimensionLIBERO Exact revision / scope
Original LIBERO suites
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

CALVIN Exact revision / scope
CALVIN language-conditioned long-horizon evaluation
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Exact version

Keep generations, checkpoints and software revisions separate.

Original LIBERO suites
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

CALVIN language-conditioned long-horizon evaluation
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

System or model type

Choose alternatives that address the same decision.

Robot evaluation benchmark
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Robot evaluation benchmark
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Environment

Simulation and real robots establish different evidence.

simulation
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

simulation
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Metric

Only matching definitions and protocols are directly comparable.

Task success rate by suite
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Average completed task-chain length; per-chain success
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Embodiment

Hardware and control boundaries are part of the evaluation.

Simulated Franka arm
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Manipulation environment supplied by CALVIN
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Access conditions

Availability of a paper is distinct from access to a model or service.

Public research documentation and evaluation resources
Primary specReviewed
Sources 1
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Public research documentation and evaluation resources
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

License / reuse

Code, checkpoints, datasets and assets can have different terms.

Not publicly disclosedNot publicly disclosed

Choose based on your constraints.

These are evaluation paths, not a ranking. Verify the current terms and run a task-specific trial before committing.

You want to examine transfer across task variations

Read LIBERO’s suite definitions, training data and evaluation settings. Report the selected suite and adaptation protocol rather than an ambiguous all-benchmark number.

You want to study language-conditioned task sequences

Read CALVIN’s sequence construction, train-test setup and metric definition. Do not equate sequence performance with another benchmark’s independent task success rate.

Evidence gaps to resolve

  • Published fields describe the identified revision only; configuration and source dates may differ.
  • No matched independent operating trial is asserted by this comparison. Missing price, reliability and intervention figures remain unknown.
  • Benchmark results require the same protocol, inputs and revision context. Cross-benchmark totals are not calculated.
Inspect the evidence

Sources & evidence

md-libero
Lifelong-Robot-Learning/LIBERO official repository

Lifelong-Robot-Learning · Repository · Publication date not disclosed · Source accessed 2026-09-28

License: MIT code; CC BY 4.0 dataset. Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.