The paper separates planning and observation measures from end-task success. A planner can make a plausible sequence while execution still fails; these metrics should not be collapsed into one robot ability score.
NeurIPS / RoboCerebra authors · Publication date not disclosed · Source accessed 2026-09-28
Reading results
The current code uses an OpenVLA-style evaluation stack, a local LIBERO dependency and explicit dataset preparation. Policy and environment revisions must both be recorded for reproduction.
qiuboxiang · Publication date not disclosed · Source accessed 2026-09-28
Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.
Reported benchmark results
Context before scores. Results from different benchmarks are not directly comparable. A simulation result does not establish real-world reliability, safety or commercial availability.
No result meets our complete revision and methodology requirements for this record yet. Inspect the original benchmark documentation before comparing published scores.
What this evidence does not establish
Research planning metrics are not evidence of customer deployment. Code licensing does not automatically establish dataset licensing.
Related reading is an editorial crosslink. Sourced connections describe relationships reported in the cited material. A link to a versioned profile does not establish compatibility with that version unless the connection note explicitly identifies it.