The robustness suite deliberately changes conditions that original LIBERO evaluations may keep fixed. Its protocol is useful for understanding sensitivity rather than replacing the original task benchmark.
sylvestf · Publication date not disclosed · Source accessed 2026-09-28
Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.
Reading results
Evaluation trial counts differ from common original LIBERO recipes. A robustness result needs its perturbation level, task set and trial protocol; a model’s original LIBERO score cannot be copied into this suite.
sylvestf · Publication date not disclosed · Source accessed 2026-09-28
Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.
Reported benchmark results
Context before scores. Results from different benchmarks are not directly comparable. A simulation result does not establish real-world reliability, safety or commercial availability.
No result meets our complete revision and methodology requirements for this record yet. Inspect the original benchmark documentation before comparing published scores.
What this evidence does not establish
Repository availability is not proof of a dataset reuse license; only attributed metadata is retained here.
Related reading is an editorial crosslink. Sourced connections describe relationships reported in the cited material. A link to a versioned profile does not establish compatibility with that version unless the connection note explicitly identifies it.