Survey 1 · decision brief · primary-source first

Datasets and benchmarks for language-space RL in robotic manipulation

What to use when a high-level VLM repeatedly chooses a language subtask, a low-level VLA executes it, and learning must reward useful decomposition, memory, uncertainty handling, and recovery—not merely end-to-end imitation.

Search date: 3 Sep 2026Related-work depth: 110 decision candidatesFocus: bimanual / YAM-likeOffline & off-policy RL

Executive answer

The most coherent first project is not a single dataset. Use RoboCasa365 target composites for the first high-level RL proof of concept, initialize the planner/critic with RoboVQA plus RoboCasa’s own subtask streams, and develop the low-level bimanual VLA and real protocol on ABC-130K / the open YAM station. Keep RoboTwin 2.0 as the bimanual transfer and robustness check. This division preserves the scientific question—learning decisions in language space—without making low-level control or simulator construction the bottleneck.

#1 Sim

RoboCasa365

365 kitchen tasks, 2,500 scenes, 600+ human and 1,600+ synthetic hours. The July 2026 target-data update adds timestep-aligned subtask index, atomic skill, stage, and natural-language labels—the closest ready-made transition structure for Q(s, language-subtask).1↗

#1 Real

ABC-130K

130,703 YAM episodes / 3,590.7 hours; 42,980 episodes have separate, time-aligned subtask annotations. Hardware, MuJoCo model, code, and Apache-2.0 data make it the cleanest embodiment bridge.2↗

#1 Brain

RoboVQA

238 hours of long-horizon video-text supervision explicitly covering planning, success, affordance, past description, and future prediction. It matches the planner/critic outputs unusually well, but is not action-policy data.3↗

Critical design point. Offline successful demonstrations alone cannot identify whether unchosen language actions were bad. To learn a useful high-level Q-function, add either (a) simulator rollouts for alternative subtasks, (b) logged failures/interventions from RoboMIND/AgiBot/your robot, or (c) conservative value learning with explicit support constraints. Treat “next annotated subtask” as behavior cloning supervision, not automatically as a reward label.

1Recommended minimum stack

Phase A: RoboCasa365’s 50 target tasks, beginning with 16 composite-seen and 16 composite-unseen tasks; expose atomic skills as the language action set and store every high-level transition. Phase B: reproduce on 8–12 RoboTwin tasks chosen for coordination and tool use. Phase C: collect 2 small YAM task families—constrained kitting and bag/box packing—with failures and corrective interventions, seeded from ABC. This is more decision-useful than pretraining on every large corpus.

What the project is actually testing

The proposed system is a semi-Markov hierarchy: at decision time t, a VLM observes images plus history and proposes a short language command; a VLA executes it for a variable duration; the environment then returns progress, failure, and a new observation. The research claim lives at the boundary between subtasks.

Needed supervision

Aligned (state/history, chosen subtask, outcome, next state); temporal subtask segments; terminal and partial success; negative or recovery trajectories; enough language variation to separate semantics from string memorization.

Needed tasks

Branching plans, uncertain state, mistakes that change the next best action, delayed dependencies, or goals requiring semantic interpretation. Fixed five-step scripts test horizon, but weakly test closed-loop reasoning.

Needed generalization

Hold out compositions, object categories, rules, scene layouts, and perturbations—not random episodes from the same task. Report seen-skill/unseen-composition separately from truly unseen skills.

Inclusion test

A candidate enters the main comparison if it materially supports at least one of: (1) high-level planner/critic supervision; (2) runnable long-horizon or uncertainty-rich evaluation; (3) bimanual/VLA skill learning; or (4) a credible sim-to-real route. Resources with only task-level captions, only successful short skills, or only mobile/humanoid embodiments remain useful but receive a lower fit rating.

Decision matrix

realsimtraining databenchmarkhigh fitconditionallow fit
Source-native counts are preserved; “tasks,” “skills,” and episodes are not interchangeable. Hours and labels describe the downloadable release as documented on the search date.
Candidate / roleWorld & embodimentScaleLanguage / subtask labelsWhy it is a good task sourceOffline / high-level RL readinessAvailability / licenseMain limitation
RoboCasa365
SimDataBenchmark
MuJoCo; diverse mobile single-arm/humanoid-style embodiments, not native YAM365 tasks (65 atomic, 300 composite), 2,500 scenes; 600+ h human + 1,600+ h synthetic demos. Target set: 50 tasks / 25k human demos; composites reach 15+ subtasks.1↗,4↗Composite target data now labels every timestep with subtask index, atomic-skill name, stage (pick/place/navigate), and natural-language instruction.Semantic household goals, long compositions, stateful appliances, unseen composite split, controlled scene variation.Best Runnable success predicates and exact segment boundaries make high-level transitions easy to construct; generate counterfactual/off-policy rollouts serially.Code/data public; verify licenses for individual assets/textures before redistribution.Kitchen-only; most data are successful demonstrations; embodiment and contact physics differ from a fixed bimanual YAM.
ABC-130K / ABC Sim
RealSimData
Two 6-DoF YAM arms, parallel grippers; matching MuJoCo station130,703 real episodes / 3,590.7 h; 195 tasks in seven primitive families. 42,980 annotated episodes. ABC Sim adds ~400 h / 20 tasks.2↗,5↗Episode-level task name; optional separate annotation.mcap with timestamped segment labels. Text is an executable subtask segment, not a VQA rationale.Bimanual folding, handover, insertion, tool use, assembly; durations ~7–469 s; public examples include multi-shirt folding/stacking, box folding, and bag packing.Strong data Excellent VLA initialization and embodiment match. High-level RL still needs rewards, negatives, and an environment wrapper.Apache-2.0; gated click-through; data ~22.6 TB; code, MJCF and hardware released.Only one-third of episodes annotated; primarily behavior-cloning data; no unified runnable real benchmark across all 195 tasks.
RoboVQA
Real videoBrain data
Robot, human, human-with-tool across three office buildings238 h; 829,502 video-text pairs; ~90k medium-horizon and ~5k long-horizon instructions.3↗Six QA families: next-step planning, success detection, discriminative/generative affordance, past description, future prediction; hindsight temporal segmentation.Direct supervision for closed-loop next-subtask choice, memory, and critic-like success judgment.Best VLM seed Train planner and outcome heads; pair with robot-domain states before using as a Q-function.Public Google Cloud release; research paper/project page. Verify bucket terms.Many evaluations use prerecorded teleoperation; actions are text answers, not low-level executable trajectories; intervention labels are acknowledged as imperfect.
RoboMIND 2.0
RealTwinData
Six heterogeneous dual-arm/mobile embodiments; digital twins310k real dual-arm trajectories / 1,000+ h / 739 tasks; 20k simulation trajectories; tactile and mobile subsets.6↗Task language and frame-level fine-grained instruction updates; predecessor documents 5k failure demonstrations with detailed causes.7↗Failure causes, longer dual-arm episodes, tactile contact, and matched simulation are unusually relevant to recovery and sim-to-real.Strong critic data Candidate for success/failure and recovery pretraining; audit exact V2 failure-label carryover before depending on it.Project links downloads through ModelScope; release terms and per-subset licensing require inspection.Embodiment heterogeneity and evolving versions complicate normalization; public metadata do not provide a single benchmark protocol as clean as RoboCasa/CALVIN.
RoboTwin 2.0
SimDataBenchmark
SAPIEN; ALOHA-AgileX plus four other dual-arm configurations50 bimanual tasks; 50 clean demos/task in standard LeRobot protocol; a public unified mirror reports 100k+ trajectories overall.8↗,9↗Task-level natural-language descriptions with configurable paraphrase count; trajectories are not generally segmented into compositional subtasks.True two-arm coordination, tool use, handovers, articulated objects; clutter, lighting, camera, table-height and language randomization.Good transfer test Sparse terminal success supports offline/off-policy RL, but the high-level action vocabulary must be authored.MIT code; downloadable dataset/mirror. Confirm asset-level terms.Most tasks are short/fixed rather than branching; scripted experts bias trajectories; some upstream task bugs have been documented.
CALVIN
SimDataBenchmark
PyBullet tabletop; one Franka Panda34 atomic tasks; 24 h teleoperated play (6 h × 4 environments); evaluation uses 1,000 chains of five language goals; only ~1% / ~24k windows are language-annotated.10↗,11↗~400 crowdsourced phrasings mapped to task IDs and temporal start/end indexes; annotations describe atomic skills such as opening a drawer or moving a block.Continuous play, sequential evaluation, explicit unseen-environment split, cheap and mature.Fast baseline Excellent smoke test for memory and hierarchical logging; can synthesize alternate subtask rollouts.Open-source code and downloadable splits (166–656 GB); research use, inspect repository/submodule licenses.Single scene grammar, single arm, five-task chains are often fixed and offer little semantic uncertainty; old Python/PyBullet stack.
BEHAVIOR-1K / 2026 Challenge
SimDataBenchmark
OmniGibson; mobile manipulators; rigid, articulated, deformable, fluid objects1,000 full activities / 50 scenes / 10k+ objects. Challenge data: 20k human teleop demos over 100 tasks, average 351.5 s and 27.1 labeled skills/episode.12↗,13↗Episode language plus 270,600 segments from 31 skill types (open/close, pour, chop, wipe, attach, etc.). Goals are first-order-logic BDDL predicates.Best uncertainty/memory stress test: stateful household processes, long horizons, deformables and many valid orderings.High scientific fit Skill labels and progress predicates enable high-level RL; computationally heavy and mobile.Code/tasks public; asset dataset partly encrypted due to third-party licensing; 2026 demos ~3.27 TB.Direct conflict with preferred tabletop/YAM embodiment; simulator complexity can dominate the research; sim-to-real is expensive.
VLABench
SimGenerated dataBenchmark
MuJoCo tabletop; single-arm manipulator100 categories: 60 primitive + 40 composite; 2,000+ objects; automated expert trajectory generation, but no stable canonical episode/hour total advertised.14↗Natural instructions include implicit intent, world knowledge, spatial/physical constraints; composite task language is task-level rather than a released universal subtask segmentation.Excellent semantic/OOD evaluation: ordering books by publication year, commonsense object selection, sorting with implicit rules.Planner benchmark Strong test of VLM decomposition; author subtask action schemas and reward checkpoints for RL.Open code/data-generation repository; verify code and asset licenses separately.Single arm; synthetic expert data; reasoning benchmark can conflate perception/world knowledge with control; scale not reported in one auditable unit.
AgiBot World Beta
RealData
Mobile dual-arm humanoid; grippers/dexterous hands/tactile; five domains1,001,552 trajectories / 2,976.4 h / 217 tasks / 87 skills / 106 scenes.15↗Task, scene, item, and subtask-segmented skill annotations; imperfect episodes include annotated error states and recovery.Longer real tasks (typically 30–60 s), real operational scenes, explicit recovery, many atomic skills.Strong semantics Valuable planner/critic corpus, but action-policy transfer to YAM is indirect.Public Hugging Face release, CC BY-NC-SA 4.0.Non-commercial/share-alike; mobile humanoid embodiment and high-dimensional hands diverge from the preferred platform.
Language Table
RealSimDataBenchmark
Planar x-y pushing of colored blocks, one low-cost arm442,226 real + 181,020 human-sim episodes; 206k unique language strings reported across the mix.16↗Every trajectory has a low-level natural-language command; relabeling yields huge paraphrase diversity and interactive commands can change online.Exceptional language grounding and correction-by-new-command; clean test of paraphrase and relational generalization.Language pretrain Useful for instruction encoder/low-level interface, weak for value learning.Apache-2.0 repository and public TFDS/GCS data.2D pushing, no grasping or bimanual control, short skills, little plan branching—do not use as the main benchmark.

Training resources: what each can and cannot teach

Planner / memory: RoboVQA

Use
Supervised warm-start for next-subtask generation, “did it work?”, affordance filtering, history summary, and future prediction.
Sample target
Given goal + last frames + prior subtasks, choose the next executable instruction and predict whether the previous instruction succeeded.
Do not infer
A RoboVQA score is not robot return; much footage is teleoperated and the textual candidate set may include actions unsupported by the target VLA.

Recommendation: multitask fine-tune the same visual backbone with separate planner, success, and history heads; calibrate them again on target-simulator rollouts.

Grounded planning: ShareRobot

Scale
51,403 curated OXE instances, 23 source datasets, 12 embodiments, 107 atomic task types, and 1,028,060 QA pairs; examples include 10,290 high-level descriptions and 28,181 low-level instructions.17↗
Labels
Task planning, object affordance, and end-effector trajectory QA; low-level phrases such as “grasp the ketchup bottle.”
Best use
Auxiliary VLM grounding and affordance supervision.

Limitation: QA is generated from selected OXE frames with template/VLM assistance; it adds labels, not new physical experience, and can transmit generator bias.

Low-level bimanual VLA: ABC-130K

Use
Pretrain the executor on the exact dual-YAM geometry; sample annotated segments as short-horizon language-conditioned trajectories.
Granularity
Task string plus timestamped subtask segments stored separately from the episode, sharing one absolute nanosecond clock.
Practicality
Train on a selected task/object shard first; the full release is 22.6 TB and uses independent camera/state clocks.

Recommendation: prioritize annotated tasks overlapping the final real domains; add paraphrases while retaining a canonical skill ID.

Outcome and recovery: RoboMIND / AgiBot

Use
Success/failure classification, failure-cause prediction, retry-versus-switch decisions, and recovery-sequence imitation.
Strength
Both explicitly retain imperfect data; this is rare in large public robot corpora.
Risk
Outcome prevalence, intervention policy, embodiment, and annotation schema can leak shortcuts.

Recommendation: use as auxiliary pretraining, then collect target-domain failures. A critic trained only on another robot’s mistakes is not a calibrated target-domain value estimator.

Large generic corpora are secondary here. DROID (76k demonstrations / 350 h), BridgeData V2 (60,096 trajectories, 13 skills), and Open X-Embodiment offer useful visual-action initialization, but mostly episode-level task language and successful short skills. BridgeData V2 did evaluate offline RL and every trajectory has a natural-language task annotation,18↗ yet none is as aligned to bimanual subtask-level high-level RL as ABC/RoboMIND/AgiBot. Use them only if the executor underperforms on basic grounding.

Simulation shortlist and experiment order

1RoboCasa365: main scientific result

QuestionCan offline/off-policy learning improve language-subtask selection beyond supervised next-step prediction?Start small8 atomic skills + 8 composite tasks in a few kitchens; use the released timestep labels to build option boundaries.ScaleComposite-seen → composite-unseen → held-out scenes/objects/instructions. Report progress predicates, full success, recovery rate, intervention rate, and subtask efficiency.Why firstThe data and task graph already expose the abstraction the project needs. No other reviewed simulator combines this annotation density with comparable long-horizon scale.

2RoboTwin 2.0: embodiment and robustness check

QuestionDoes the learned high-level mechanism survive bimanual timing, contact, and visual/language perturbations?Choose tasksHandover, lift pot, stack bowls, hammer, scan object, stamp/seal, turn switch, and a multi-object sorting task.Required workDefine a compact canonical language action set and instrument intermediate predicates; stock tasks do not expose rich hierarchical labels.Why secondIt matches the bimanual form factor and has strong domain randomization, but is weaker than RoboCasa for branching/memory.

3CALVIN: rapid ablation harness

QuestionDo memory, closed-loop replanning, and the value head help at all?ProtocolRe-use 1,000 five-goal chains; create perturbations after goals 1–3 (move a block, undo a drawer, corrupt one observation) so the next action cannot be a fixed script.Why thirdFast and standardized, but success on unperturbed CALVIN does not establish semantic reasoning or bimanual transfer.

Do not start with BEHAVIOR-1K. It is scientifically attractive and its 2026 labels are exceptionally rich, but GPU/simulator complexity, mobile control, long 5.9-minute episodes, and asset licensing make it a poor first implementation. Promote it only after the language-option learner is stable.

Two feasible real-world task sets

A. Rule-conditioned kitting & light assembly

Recommended first A fixed dual-YAM workcell with bins, trays, labeled parts, fasteners, simple tools, and distractors.

  1. Read a natural-language order or infer a rule from a reference card.
  2. Inspect partially occluded bins; move blockers if needed.
  3. Select compatible parts and place them in order.
  4. Perform one bimanual assembly/insertion or tool action.
  5. Verify orientation/count; correct a wrong, dropped, or missing part; close/package.

Why reasoning is necessary: the next subtask depends on observations and order constraints; unseen orders compose trained skills; memory tracks counts and already-inspected bins; controlled faults create recovery data.

Proposed split: train on 6 part families × 4 order rules; test held-out rule–part combinations, novel distractors, missing-part cases, and one unseen tool. Collect 50–100 expert episodes/task plus policy rollouts and interventions. Reward partial predicates (correct part, correct slot, assembly seated) and terminal order correctness.

B. Bag / box packing with deformables

Second, harder A dual-YAM station with a backpack or folding carton, heterogeneous items, packing constraints, and optional zipper/lid.

  1. Open and stabilize container bimanually.
  2. Inspect item set against a checklist; decide packing order by fragility/size.
  3. Reorient and insert items, maintaining access to later items.
  4. Detect protrusion, drop, wrong item, jam, or failed closure.
  5. Repack or retry; close the zipper/lid and verify.

Why reasoning is necessary: order is state-dependent, closure hides prior state, memory matters, and failed insertion changes the best next move. ABC already demonstrates box folding, DAgger-improved box closing, and end-to-end student-bag packing, reducing feasibility risk.5↗

Proposed split: train on item sets and constraints separately; test unseen combinations, swapped container, occluded checklist, and injected packing mistakes. Use progress reward but reserve strict closure/checklist correctness as terminal reward.

Why not laundry first? Garment folding is visually compelling and well represented in ABC, but deformable-state estimation can overwhelm the language-space question. Packing preserves bimanual/deformable difficulty while giving clearer symbolic progress and cheaper success instrumentation. Add laundry as a later stress test.

Make the benchmark answer the research claim

State

Current multi-view observation, compact history of attempted subtasks and outcomes, remaining-goal representation, elapsed budget, and optional VLA confidence. Ablate raw history vs learned memory.

Action

A constrained but compositional language schema: verb(object, relation, target) rendered to natural text. Retain canonical IDs so paraphrases do not create fake distinct actions.

Transition

Execute the VLA until learned termination, success predicate, timeout, or safety stop. Log duration and executor outcome; this is a semi-Markov transition, so duration-aware discounting matters.

Reward

Terminal success + delta in predicate progress − time − invalid action − repeated futile attempt − intervention. Never reward matching the demonstrator’s exact plan when multiple valid decompositions exist.

Offline data

Demonstrations provide positive behavior support; simulator rollouts add alternatives; target failures/interventions add recovery. Use conservative Q-learning or advantage-weighted behavior learning and constrain actions to executor support.

Evaluation

Full success, partial progress, language actions per success, invalid-subtask rate, recovery after injected fault, calibration of value/success heads, and generalization across objects, compositions, rules, scenes, and paraphrases.

Baselines that isolate the contribution

  1. Oracle task graph: shortest valid next skill from simulator state.
  2. Behavior cloning: supervised next-subtask prediction from annotated trajectories.
  3. Open-loop VLM: plan once, then execute without revising.
  4. Closed-loop VLM, no memory/value: replan from the current image only.
  5. Memory + success head: isolates state/history modeling.
  6. Offline/off-policy language Q: the proposed learning contribution.
  7. Low-level oracle: replace the VLA with scripted skills to separate planner failure from executor failure.

Cross-cutting gaps and cautions

Counterfactual scarcity

Segmented expert demonstrations say what happened, not what would have happened under another subtask. Without exploration or conservative objectives, Q-values for unseen language actions are extrapolation.

Executor support

A knowledgeable VLM can name a sensible action the VLA cannot execute. Track empirical skill competence and mask or penalize unsupported proposals; otherwise high-level “reasoning” fails for low-level reasons.

Boundary leakage

Per-frame subtask labels may make termination trivial during training. At test time the boundary must come from observation, executor state, or a learned detector. Evaluate with annotations hidden.

Recovery bias

RoboMIND/AgiBot failures arise under particular teleoperators and policies. Their frequency is not the target robot’s failure distribution; recalibration and deliberately injected target faults are necessary.

“Long horizon” inflation

Five concatenated atomic skills do not necessarily require reasoning. Include irreversible choices, alternative valid orders, hidden state, and perturbations that force the correct plan to change.

License composition

Code, robot trajectories, 3D assets, textures, and derived labels can carry different terms. “Open source” on a project page is not enough; audit the exact files used before model redistribution.

Useful but not primary

LIBERO has 130 language-conditioned tasks in four continual-learning suites,19↗ usually 50 demonstrations/task, but its principal question is lifelong learning and its tasks are mostly isolated; use it for compositional/OOD executor sanity checks, not the main language-RL claim. PerAct² adds 13 bimanual RLBench tasks / 23 variations,20↗ but is smaller and less diverse than RoboTwin. BiGym offers 40 tasks × 50 human demos with sparse rewards,21↗ yet its mobile humanoid emphasis conflicts with the requested fixed bimanual setup. All three are legitimate comparison points, not first choices.

Scope, methodology, and stopping rationale

Scope and assumptions. The survey interprets “VLM or critic” as the high-level language-option selector/outcome model, “VLA” as a separately trainable short-horizon executor, and “real-world task set” as a lab-reproducible fixed or YAM-like bimanual workcell. Mobile manipulation, humanoids, and dexterous hands are included only when their annotations or failures transfer to the high-level problem. Search date is 3 September 2026; maximum related-work depth is 1.

Search procedure. The project brief and the supplied large-dataset survey were treated as anchors. Semantic searches covered language-conditioned manipulation, long-horizon/compositional benchmarks, subgoal and frame-level instruction labels, failure/recovery datasets, offline RL, bimanual simulation, and sim-to-real. Candidate claims were checked against canonical paper pages, papers, official repositories, dataset cards, and official documentation. Related-work sections were traversed one hop to find competing/specialized benchmarks. Versions were deduplicated (for example RoboMIND 1.x vs 2.0 and AgiBot Alpha vs Beta).

Evidence types. RoboCasa365 is peer-reviewed at ICLR 2026; CALVIN is peer-reviewed and received the 2022 RA-L Best Paper Award; VLABench appeared at ICCV 2025. Several newest resources—ABC-130K, RoboTwin 2.0, RoboMIND 2.0—are preprints or organization releases; their scale/performance figures are author-reported. Product/project-page updates after paper publication are labeled by date and treated as release facts, not independent validation.

Exclusions. Pure navigation, grasp-only sets, action-free web video, demo-only company showcases, and generic multimodal instruction data were excluded. Older broad datasets (OXE, Bridge, DROID) were retained only as baselines because they do not expose sufficiently rich subtask/outcome structure. Humanoid/mobile benchmarks were demoted when embodiment complexity was likely to obscure the high-level RL contribution.

Coverage gaps. No multi-terabyte dataset was downloaded for sample-level auditing. Some living dataset cards changed counts after associated papers; source-native current figures are shown and discrepancies noted. Asset-level licenses, exact failure-label prevalence in RoboMIND 2.0, and precise public trajectory counts for VLABench require direct release audit before implementation. Search did not benchmark simulator installation stability or hardware cost.

Stopping rationale. Search stopped when new candidates repeated one of the covered roles without improving the decision frontier: RoboCasa for hierarchical simulation, ABC for YAM-scale bimanual data, RoboVQA/RoboMIND/AgiBot for reasoning and recovery supervision, RoboTwin for bimanual robustness, CALVIN for mature long-horizon ablation, and BEHAVIOR/VLABench for harder semantic stress tests. Additional catalogs would add breadth but not change the recommended stack.

Canonical sources

  1. RoboCasa project and 2026 release notes — tasks, scenes, hours, and per-frame subtask annotations.
  2. ABC-130K official dataset card — YAM embodiment, current episode/hour counts, schema, access, size, and Apache-2.0 license.
  3. RoboVQA project and paper — QA types, hours, instructions, embodiments, and evaluation caveats.
  4. RoboCasa365 paper — 365-task design, 50-task target splits, human/synthetic data, and long-horizon statistics.
  5. ABC project and technical report — task taxonomy, ABC Sim, matched sim/real evaluation, DAgger, box and bag tasks.
  6. RoboMIND 2.0 project and paper — 310k real / 20k twin trajectories, 739 tasks, six embodiments.
  7. RoboMIND paper and project — 107k successful plus 5k failure demonstrations and failure-cause labels.
  8. RoboTwin 2.0 official task documentation — 50 tasks, step lengths, embodiments and success rates.
  9. RoboTwin repository, paper, and LeRobot integration — MIT code, generation, randomization and evaluation data.
  10. CALVIN repository and paper — benchmark goals and sequence protocol.
  11. CALVIN dataset documentation — 24 hours, data modalities, actions, and temporal language annotation schema.
  12. BEHAVIOR-1K project and paper — 1,000 tasks, simulation features, and sim-to-real study.
  13. BEHAVIOR 2026 challenge dataset documentation — 20k demos, 100 tasks, duration, skill counts, and language annotations.
  14. VLABench project, ICCV 2025 paper, and repository.
  15. AgiBot World / GO-1 paper and official dataset organization — scale, hierarchy, recovery data, embodiment, and license.
  16. Language Table repository and Google Research article — component episode counts and language diversity.
  17. RoboBrain / ShareRobot CVPR 2025 paper and official model/data repository.
  18. BridgeData V2 project and paper — trajectory/task scale, language, generalization, and offline-RL evaluation.
  19. LIBERO documentation and repository.
  20. PerAct² project — 13 bimanual RLBench tasks / 23 variations and downloadable data.
  21. BiGym project and paper — 40 tasks, human demonstrations, sparse rewards.

Prepared as a selective technical survey, not a legal opinion. Verify the current license and access conditions of every dataset, asset, and derived annotation before redistribution. Back to top ↑