The evidence-linked database

Benchmarks

Understand what a benchmark measures, how it is run, and what its results cannot establish.

18 reviewed entriesEvidence before volume ↗

Reported results, in context

Context before scores. Results from different benchmarks are not directly comparable. A simulation result does not establish real-world reliability, safety or commercial availability.
Model / revisionBenchmark / metricReported resultEnvironmentEvidence & scope
OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Spatial97.6Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Object98.4Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Goal97.9Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioceptionLIBEROTask success rate (%) · LIBERO-Long94.5Reported 2025-02-27simulationSimulated Franka Emika PandaPrimary releaseVerified 2026-09-28
Sources 1
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1

Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28

Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied.

Methodology & comparability

Per-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate.

Comparison group: oft-2502.19645v1-table1-wrist-proprio

Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed.

Verified 2026-09-28

18 entries with reviewed evidence

Alphabetical · no paid ranking
CABenchmarks

CALVIN

CALVIN tests whether language-conditioned policies can compose manipulation skills over long sequences in simulated environments.

Benchmarksimulation
Reviewed 2026-09-28
KIBenchmarks

Kinetix

Kinetix is a JAX-based physics environment for studying generalization across procedurally generated control problems.

Benchmarksimulation
Reviewed 2026-09-28
LIBenchmarks

LIBERO

LIBERO separates manipulation evaluation into spatial, object, goal and longer-horizon task suites, allowing transfer behavior to be examined along different axes.

Benchmarksimulation
Reviewed 2026-09-28
LIBenchmarks

LIBERO-Plus

LIBERO-Plus examines how policies respond to changes in cameras, initialization, instructions, lighting, backgrounds, noise and layout.

Benchmarksimulation
Reviewed 2026-09-28
LIBenchmarks

LIBERO-Pro

LIBERO-Pro extends LIBERO evaluation to probe robustness beyond memorized initial states, objects and instructions.

Benchmarksimulation
Reviewed 2026-09-28
MABenchmarks

ManiSkill2

ManiSkill2 evaluates manipulation with varied objects and task families using SAPIEN, including rigid and deformable interactions.

Benchmarksimulation
Reviewed 2026-09-28
MIBenchmarks

MIKASA-Robo

MIKASA-Robo targets manipulation that depends on remembering information through occlusion, delay or multistage interaction.

Benchmarksimulation
Reviewed 2026-09-28
MOBenchmarks

MolmoSpaces Bench

MolmoSpaces Bench uses a versioned ecosystem of scenes, objects and robot assets to test manipulation in varied simulated environments.

Benchmarksimulation
Reviewed 2026-09-28
RLBenchmarks

RLBench

RLBench is a manipulation benchmark built on CoppeliaSim and PyRep, with task variations and interfaces for imitation and reinforcement learning.

Benchmarksimulation
Reviewed 2026-09-28
ROBenchmarks

RoboArena

RoboArena compares generalist policies through decentralized, blinded pairs of physical robot trials across different scenes and tasks.

Benchmarkreal_robot
Reviewed 2026-09-28
ROBenchmarks

RoboCasa

RoboCasa’s original release provides kitchen manipulation environments for studying everyday robot tasks and environment variation.

Benchmarksimulation
Reviewed 2026-09-28
ROBenchmarks

RoboCasa365

RoboCasa365 expands household manipulation evaluation with a separately published task collection and diverse simulated kitchens.

Benchmarksimulation
Reviewed 2026-09-28
ROBenchmarks

RoboCerebra

RoboCerebra evaluates long-horizon manipulation with an emphasis on higher-level planning, observation and memory.

Benchmarksimulation
Reviewed 2026-09-28
ROBenchmarks

RoboDojo

RoboDojo combines simulated and real-robot evaluation across generalization, memory, precision and longer task horizons.

Benchmarkmixed
Reviewed 2026-09-28
ROBenchmarks

RoboMME

RoboMME evaluates robot policies on tasks that require temporal, spatial, object and procedural memory.

Benchmarksimulation
Reviewed 2026-09-28
ROBenchmarks

RoboTwin 2.0

RoboTwin 2.0 combines a bimanual manipulation benchmark with a synthetic demonstration-generation pipeline and structured domain randomization.

Benchmarksimulation
Reviewed 2026-09-28
SIBenchmarks

SimplerEnv

SimplerEnv uses simulation to study manipulation policies in environments aligned with real robot evaluation setups.

Benchmarksimulation
Reviewed 2026-09-28
VLBenchmarks

VLABench

VLABench focuses on language-conditioned manipulation with longer task structure and semantic generalization challenges.

Benchmarksimulation
Reviewed 2026-09-28