Benchmarks
Understand what a benchmark measures, how it is run, and what its results cannot establish.
Reported results, in context
| Model / revision | Benchmark / metric | Reported result | Environment | Evidence & scope |
|---|---|---|---|---|
| OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioception | LIBEROTask success rate (%) · LIBERO-Spatial | 97.6Reported 2025-02-27 | simulationSimulated Franka Emika Panda | Primary releaseVerified 2026-09-28Sources 1Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1 Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28 Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied. Methodology & comparabilityPer-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate. Comparison group: oft-2502.19645v1-table1-wrist-proprio Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed. Verified 2026-09-28 |
| OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioception | LIBEROTask success rate (%) · LIBERO-Object | 98.4Reported 2025-02-27 | simulationSimulated Franka Emika Panda | Primary releaseVerified 2026-09-28Sources 1Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1 Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28 Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied. Methodology & comparabilityPer-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate. Comparison group: oft-2502.19645v1-table1-wrist-proprio Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed. Verified 2026-09-28 |
| OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioception | LIBEROTask success rate (%) · LIBERO-Goal | 97.9Reported 2025-02-27 | simulationSimulated Franka Emika Panda | Primary releaseVerified 2026-09-28Sources 1Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1 Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28 Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied. Methodology & comparabilityPer-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate. Comparison group: oft-2502.19645v1-table1-wrist-proprio Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed. Verified 2026-09-28 |
| OpenVLA-OFTOpenVLA-OFT, arXiv:2502.19645v1 Table I final row; wrist camera + proprioception | LIBEROTask success rate (%) · LIBERO-Long | 94.5Reported 2025-02-27 | simulationSimulated Franka Emika Panda | Primary releaseVerified 2026-09-28Sources 1Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success, v1 Moo Jin Kim, Chelsea Finn and Percy Liang · Published 2025-02-27 · Source accessed 2026-09-28 Table I final row and evaluation protocol inspected. Factual per-suite results only; no leaderboard copied. Methodology & comparabilityPer-suite fine-tuning; best checkpoint selected from periodic evaluations; 500 trials per suite. Author-reported simulation, not an independent reproduction or physical reliability estimate. Comparison group: oft-2502.19645v1-table1-wrist-proprio Paper explicitly CC BY 4.0. Attributed factual result only; no dataset or leaderboard redistributed. Verified 2026-09-28 |
CALVIN
CALVIN tests whether language-conditioned policies can compose manipulation skills over long sequences in simulated environments.
Kinetix
Kinetix is a JAX-based physics environment for studying generalization across procedurally generated control problems.
LIBERO
LIBERO separates manipulation evaluation into spatial, object, goal and longer-horizon task suites, allowing transfer behavior to be examined along different axes.
LIBERO-Plus
LIBERO-Plus examines how policies respond to changes in cameras, initialization, instructions, lighting, backgrounds, noise and layout.
LIBERO-Pro
LIBERO-Pro extends LIBERO evaluation to probe robustness beyond memorized initial states, objects and instructions.
ManiSkill2
ManiSkill2 evaluates manipulation with varied objects and task families using SAPIEN, including rigid and deformable interactions.
MIKASA-Robo
MIKASA-Robo targets manipulation that depends on remembering information through occlusion, delay or multistage interaction.
MolmoSpaces Bench
MolmoSpaces Bench uses a versioned ecosystem of scenes, objects and robot assets to test manipulation in varied simulated environments.
RLBench
RLBench is a manipulation benchmark built on CoppeliaSim and PyRep, with task variations and interfaces for imitation and reinforcement learning.
RoboArena
RoboArena compares generalist policies through decentralized, blinded pairs of physical robot trials across different scenes and tasks.
RoboCasa
RoboCasa’s original release provides kitchen manipulation environments for studying everyday robot tasks and environment variation.
RoboCasa365
RoboCasa365 expands household manipulation evaluation with a separately published task collection and diverse simulated kitchens.
RoboCerebra
RoboCerebra evaluates long-horizon manipulation with an emphasis on higher-level planning, observation and memory.
RoboDojo
RoboDojo combines simulated and real-robot evaluation across generalization, memory, precision and longer task horizons.
RoboMME
RoboMME evaluates robot policies on tasks that require temporal, spatial, object and procedural memory.
RoboTwin 2.0
RoboTwin 2.0 combines a bimanual manipulation benchmark with a synthetic demonstration-generation pipeline and structured domain randomization.
SimplerEnv
SimplerEnv uses simulation to study manipulation policies in environments aligned with real robot evaluation setups.
VLABench
VLABench focuses on language-conditioned manipulation with longer task structure and semantic generalization challenges.