The dataset offers different environment splits, including single-environment and cross-environment training. Long-horizon multitask language control is a separate evaluation from isolated single-task execution.
mees · Publication date not disclosed · Source accessed 2026-09-28
Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.
Reading results
The sequence protocol resets the scene at the beginning of a sequence, then measures continued task completion. Repository changes have corrected task-success criteria, making environment revision part of the result identity.
mees · Publication date not disclosed · Source accessed 2026-09-28
Official README inspected live. Repository code terms do not automatically cover model weights, data, or third-party assets.
Reported benchmark results
Context before scores. Results from different benchmarks are not directly comparable. A simulation result does not establish real-world reliability, safety or commercial availability.
No result meets our complete revision and methodology requirements for this record yet. Inspect the original benchmark documentation before comparing published scores.
What this evidence does not establish
A score without its train/test split, chain protocol and environment revision is not comparable.
Related reading is an editorial crosslink. Sourced connections describe relationships reported in the cited material. A link to a versioned profile does not establish compatibility with that version unless the connection note explicitly identifies it.