The evidence ledger

Benchmark result intelligence

Paper-reported results with protocol, original table, upstream revision and comparability limits. No score combines different benchmarks.

Compare protocols before numbers.

AllenAI VLA Evaluation Harness · Apache-2.0 metadata · pinned upstream 6cc3e1bd9fc4. These 557 eligible rows retain paper-variant identity, protocol and table provenance. Exact checkpoint hashes are not established. Only rows marked “Original table spot-checked” have a direct table check.

AI-assisted upstream extraction can contain errors. Simulation results do not establish real-robot capability. Rows are grouped by benchmark and name, never ranked across protocols. The four original V1 results remain in the benchmark catalog.

External-only leaderboards · open the original

The upstream policy excludes these leaderboards from mirrored result tables.

557 matching result rows · Page 1 of 47
calvin · simulation

3D Diffuser Actor

3D Diffuser Actor · 2402.10885 · Table 4

3.35subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table 4
Comparability group
calvin:b066157c830f3801
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

ABC→D zero-shot generalization setup with 1000 evaluation chains and avg_len metric. Language-annotated training data.

Protocol evidence retained in extraction snapshot SHA-256 95a052e2941efa1b3292f888fd8a3a31fd27d9d65c087cff8313b483cce5c2db; consult the original paper and table.

ABC→D zero-shot generalization setup with 1000 evaluation chains and avg_len metric. Language-annotated training data. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
93.8
2_tasks
80.3
3_tasks
66.2
4_tasks
53.3
5_tasks
41.2
Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2402.10885] 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations (https://arxiv.org/abs/2402.10885) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

3D Diffuser Actor (from EL3DD)

3D Diffuser Actor · 2511.13312 · Table 1

3.34subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table 1
Comparability group
calvin:95d9db8e60b1957d
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

The baseline corresponds to the 3DDA model. ABC→D split.

Protocol evidence retained in extraction snapshot SHA-256 59d84dd3ac10079c4405792fc7490b3691059b22e8d86002585021021d786ac1; consult the original paper and table.

The baseline corresponds to the 3DDA model. ABC→D split. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
93.7
2_tasks
80.1
3_tasks
66.1
4_tasks
53.2
5_tasks
41
Open reporting paper ↗

Open model paper ↗

Inspect 3 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2402.10885] 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations (https://arxiv.org/abs/2402.10885) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

[2511.13312] EL3DD: Extended Latent 3D Diffusion for Language Conditioned Multitask Manipulation (https://arxiv.org/abs/2511.13312) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

3D Diffuser Actor + GraspCorrect

3D Diffuser Actor + GraspCorrect · 2503.15035 · Table 2

3.9subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table 2
Comparability group
calvin:6f111082b1676d3b
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

ABC→D, 100 scenarios, avg_len metric.

Protocol evidence retained in extraction snapshot SHA-256 535ad865710c0d81a3e1654377d592754954fc0c8ee440ae280724fce8e88830; consult the original paper and table.

ABC→D, 100 scenarios, avg_len metric. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
97
2_tasks
87
3_tasks
77
4_tasks
69
5_tasks
54
Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2503.15035] GraspCorrect: Robotic Grasp Correction via Vision-Language Model-Guided Feedback (https://arxiv.org/abs/2503.15035) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

3D Foresight

3D Foresight · 2502.10028 · Table I

4.23subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table I
Comparability group
calvin:621e893c0fd1365e
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

ABC→D split, 1000 chains, avg_len metric.

Protocol evidence retained in extraction snapshot SHA-256 3a28a6cf7f7385697582121802106187f67fc54d30389611ab4adff3dc3831fd; consult the original paper and table.

ABC→D split, 1000 chains, avg_len metric. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
96.9
2_tasks
92
3_tasks
85.7
4_tasks
78.8
5_tasks
71.3
Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2502.10028] 3D Dynamics-Aware Manipulation: Endowing Manipulation Policies with 3D Foresight (https://arxiv.org/abs/2502.10028) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

ADPro

ADPro · 2508.06266 · Table II

3.64subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table II
Comparability group
calvin:9fe7b4f17321c1ed
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

ABC->D zero-shot evaluation, 3 seeds. Test-time adaptation on 3D Diffuser Actor. 25% fewer inference steps.

Protocol evidence retained in extraction snapshot SHA-256 241049418bb7f6eb89c00def34daccfce8686bd5b45cb3c26ea69a0222e5060d; consult the original paper and table.

ABC->D zero-shot evaluation, 3 seeds. Test-time adaptation on 3D Diffuser Actor. 25% fewer inference steps. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
94.7
2_tasks
83
3_tasks
73.6
4_tasks
61.4
5_tasks
51.1
Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2508.06266] ADPro: a Test-time Adaptive Diffusion Policy via Manifold-constrained Denoising and Task-aware Initialization for Robotic Manipulation (https://arxiv.org/abs/2508.06266) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

Anchor-Align VLA

Anchor-Align VLA · 2607.13429 · Table 2

4.5subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table 2
Comparability group
calvin:7657502c303d62e5
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.8

CALVIN ABC->D split, average rollout length (Len, 0-5) over 5-instruction chains.

Protocol evidence retained in extraction snapshot SHA-256 b6fbfc87d45e179eab5804bac3397aa6582ad2a1da99c4914542556e39eb5f5a; consult the original paper and table.

CALVIN ABC→D, average rollout length (Len, 0-5) over 5-instruction chains; this paper's method. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
99.1
2_tasks
95.8
3_tasks
90.6
4_tasks
84.7
5_tasks
77.9
Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2607.13429] Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment (https://arxiv.org/abs/2607.13429) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

AR-VRM (ABC->D)

AR-VRM (ABC->D) · 2508.07626 · Table 1

3.29subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table 1
Comparability group
calvin:59d6e3b4e1e28921
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

ABC->D split evaluation. Standard CALVIN evaluation.

Protocol evidence retained in extraction snapshot SHA-256 329fe04db27200ef9d2589045cdb888a5b431cf4fdbfa296c675b0207a0f545b; consult the original paper and table.

ABC->D split evaluation. Standard CALVIN evaluation. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
90.1
2_tasks
75.9
3_tasks
64.2
4_tasks
53.1
5_tasks
46.1
Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2508.07626] AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning (https://arxiv.org/abs/2508.07626) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

Astra

Astra · 2408.01147 · Table 3

3.29subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table 3
Comparability group
calvin:ac01cdf95f460b59
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

ABC→D split. Uses 1% language-annotated subset for training. Trained from scratch for 20 epochs.

Protocol evidence retained in extraction snapshot SHA-256 1bc282501e4588df963baf848f77407eab934325e723cf2ba55dba649b347c1d; consult the original paper and table.

ABC→D split. Uses 1% language-annotated subset for training. Trained from scratch for 20 epochs. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
89.7
2_tasks
79.2
3_tasks
65.8
4_tasks
52.4
5_tasks
42.3
Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2408.01147] Astra: Efficient Transformer Architecture and Contrastive Dynamics Learning for Embodied Instruction Following (https://arxiv.org/abs/2408.01147) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

AVA-VLA

AVA-VLA · 2511.18960 · Table 2

4.65subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table 2
Comparability group
calvin:402d94807bc6ffca
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

ABC→D split, standard CALVIN evaluation.

Protocol evidence retained in extraction snapshot SHA-256 0b61c14c6e36eb9503be64e84239078cb7d1508d3183ea359e356803a51eadf6; consult the original paper and table.

ABC→D split, standard CALVIN evaluation. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
99.6
2_tasks
97.6
3_tasks
94.1
4_tasks
89.9
5_tasks
84.1
Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2511.18960] AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention (https://arxiv.org/abs/2511.18960) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

BEAST-F

BEAST-F · 2506.06072 · Table 1

4.42subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table 1
Comparability group
calvin:d2606600595b2323
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

ABC->D split, 1000 chains. No robot data pretraining. Uses Florence-2 VLM.

Protocol evidence retained in extraction snapshot SHA-256 9de802a0fcc3b1c665925a0c92be59b00f925fb891b0c738092e9c202a399028; consult the original paper and table.

ABC->D split, 1000 chains. No robot data pretraining. Uses Florence-2 VLM. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
99.8
2_tasks
96.5
3_tasks
89.3
4_tasks
82.7
5_tasks
74.4
Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2506.06072] BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning (https://arxiv.org/abs/2506.06072) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

BehaviorVLA

BehaviorVLA · 2605.22671 · Table 4

4.36subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table 4
Comparability group
calvin:a5a9b77b21d7ecaf
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

ABC->D setting, 1000 instruction chains of length 5, 500 rollouts. Standard CALVIN protocol.

Protocol evidence retained in extraction snapshot SHA-256 ef6db1e19b18e7521c2e7f64682c85ea7ff44360e5db227aacfb646a7758d442; consult the original paper and table.

ABC->D setting, 1000 instruction chains of length 5, 500 rollouts. Standard CALVIN protocol. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Reported component scores

Component values retain the upstream protocol scale. CALVIN chain completion components use percentages; its aggregate is completed subtasks (0–5).

1_task
96
2_tasks
92
3_tasks
87.3
4_tasks
82.9
5_tasks
77.3
Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2605.22671] From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model (https://arxiv.org/abs/2605.22671) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.

calvin · simulation

Being-H0.7

Being-H0.7 · 2605.00078 · Table 1

4.48subtasksavg_len
Upstream extractionPaper reported · no independent reproductionnot directly comparable
Protocol, scores & provenance
Model revision scope
paper_variant
Checkpoint hash
Not publicly disclosed
Protocol
calvin@6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Environment
simulation
Embodiment
Not publicly disclosed for this result in the curated row
Training regime
finetuned
Paper table
Table 1
Comparability group
calvin:4c80a776acba8637
Reproduction status
PAPER_REPORTED
Upstream commit
6cc3e1bd9fc4bed26fb0a3aabd342c5fed11ca09
Upstream data date
2026-08-10
Imported / checked
2026-09-28
Upstream curator
opus 4.6

CALVIN* denotes ABC→D split. Reports avg_len metric.

Protocol evidence retained in extraction snapshot SHA-256 4f488803411d6f000a125cc3c2225449ad0450c0cc73c1e766c05e2d82b7ce2a; consult the original paper and table.

CALVIN* denotes ABC→D split. Reports avg_len metric. PhysicalAI.best has not reproduced this experiment. Identity is the paper-defined variant; checkpoint hash, evaluation seeds and exact weight artifact are not disclosed in this upstream row. Upstream aggregate/suite metric units must not be mixed.

Open reporting paper ↗
Inspect 2 sources
AllenAI VLA benchmark data — pinned curated snapshot ↗

Allen Institute for AI / VLA evaluation harness contributors · Accessed 2026-09-28 · direct primary review

Apache-2.0 (repository); original papers retain their own terms · Normalized licensed repository data with attribution; no mirror of external-only leaderboards or model/dataset artifacts.

Primary for the upstream snapshot, not an independent verification of every original paper or score. Upstream curation is AI-assisted.

[2605.00078] Being-H0.7: A Latent World-Action Model from Egocentric Videos (https://arxiv.org/abs/2605.00078) ↗

Original paper authors (linked by upstream) · Accessed 2026-09-28 · upstream attributed

Attribution link and factual numeric reference only; paper text, figures and artifacts are not republished.

Paper URL and extracted context were inspected in the pinned upstream data. Document title/identity and availability were checked separately. Numerical tables have not been independently rechecked except where a result explicitly says so; this is not a reproduction claim.