Benchmark / simulation

CALVIN

CALVIN tests whether language-conditioned policies can compose manipulation skills over long sequences in simulated environments.

Source contextPrimary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Reviewed recordLast reviewed 3 sources ↗6 sourced fields

What the protocol tests

The dataset offers different environment splits, including single-environment and cross-environment training. Long-horizon multitask language control is a separate evaluation from isolated single-task execution.

Source contextPrimary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Reading results

The sequence protocol resets the scene at the beginning of a sequence, then measures continued task completion. Repository changes have corrected task-success criteria, making environment revision part of the result identity.

Source contextPrimary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Methodology and reuse limits

A score without its train/test split, chain protocol and environment revision is not comparable.

Source contextPrimary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Evaluation protocol

Exact version / scope
CALVIN language-conditioned long-horizon evaluation
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Environment
simulation
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Metric definition
Average completed task-chain length; per-chain success
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Task scope
Not publicly disclosed
Embodiment
Manipulation environment supplied by CALVIN
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Access & reuse

Access model
Public research documentation and evaluation resources
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

License
Not publicly disclosed

Additional disclosed fields

Category
Robot evaluation benchmark
Primary specReviewed
Sources 1
mees/calvin official repository

mees · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Reported benchmark results

Context before scores. Results from different benchmarks are not directly comparable. A simulation result does not establish real-world reliability, safety or commercial availability.
No result meets our complete revision and methodology requirements for this record yet. Inspect the original benchmark documentation before comparing published scores.

What this evidence does not establish

  • A score without its train/test split, chain protocol and environment revision is not comparable.

Relationships & deployments

Related reading is an editorial crosslink. Sourced connections describe relationships reported in the cited material. A link to a versioned profile does not establish compatibility with that version unless the connection note explicitly identifies it.

evaluated on
OpenHelix → CALVIN
Sources 1
OpenHelix-Team/OpenHelix official repository

OpenHelix-Team · Publication date not disclosed · Source accessed 2026-09-28

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

Connection reviewed 2026-09-28

evaluated on
TS-Mask VLA → CALVIN
Sources 1
TS-Mask VLA original paper

Shengzhuo Yang and coauthors · Published 2026-07-10 · Source accessed 2026-09-28

Connection reviewed 2026-09-28

Curated comparisons

Record history & verified changes

A research review records when we checked a source. It does not mark a product launch or a new deployment.

Initial reviewed record. No subsequent field change has been recorded.

Inspect the evidence

Sources & evidence

md-openhelix
OpenHelix-Team/OpenHelix official repository

OpenHelix-Team · Repository · Publication date not disclosed · Source accessed 2026-09-28

License: MIT code. Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.

Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.

md-tsmask
TS-Mask VLA original paper

Shengzhuo Yang and coauthors · Paper · Published 2026-07-10 · Source accessed 2026-09-28

Factual summary and attribution only; no upstream prose, images, weights, or dataset redistributed.