RoboCasa365 SimDataBenchmark | MuJoCo; diverse mobile single-arm/humanoid-style embodiments, not native YAM | 365 tasks (65 atomic, 300 composite), 2,500 scenes; 600+ h human + 1,600+ h synthetic demos. Target set: 50 tasks / 25k human demos; composites reach 15+ subtasks.1↗,4↗ | Composite target data now labels every timestep with subtask index, atomic-skill name, stage (pick/place/navigate), and natural-language instruction. | Semantic household goals, long compositions, stateful appliances, unseen composite split, controlled scene variation. | Best Runnable success predicates and exact segment boundaries make high-level transitions easy to construct; generate counterfactual/off-policy rollouts serially. | Code/data public; verify licenses for individual assets/textures before redistribution. | Kitchen-only; most data are successful demonstrations; embodiment and contact physics differ from a fixed bimanual YAM. |
ABC-130K / ABC Sim RealSimData | Two 6-DoF YAM arms, parallel grippers; matching MuJoCo station | 130,703 real episodes / 3,590.7 h; 195 tasks in seven primitive families. 42,980 annotated episodes. ABC Sim adds ~400 h / 20 tasks.2↗,5↗ | Episode-level task name; optional separate annotation.mcap with timestamped segment labels. Text is an executable subtask segment, not a VQA rationale. | Bimanual folding, handover, insertion, tool use, assembly; durations ~7–469 s; public examples include multi-shirt folding/stacking, box folding, and bag packing. | Strong data Excellent VLA initialization and embodiment match. High-level RL still needs rewards, negatives, and an environment wrapper. | Apache-2.0; gated click-through; data ~22.6 TB; code, MJCF and hardware released. | Only one-third of episodes annotated; primarily behavior-cloning data; no unified runnable real benchmark across all 195 tasks. |
RoboVQA Real videoBrain data | Robot, human, human-with-tool across three office buildings | 238 h; 829,502 video-text pairs; ~90k medium-horizon and ~5k long-horizon instructions.3↗ | Six QA families: next-step planning, success detection, discriminative/generative affordance, past description, future prediction; hindsight temporal segmentation. | Direct supervision for closed-loop next-subtask choice, memory, and critic-like success judgment. | Best VLM seed Train planner and outcome heads; pair with robot-domain states before using as a Q-function. | Public Google Cloud release; research paper/project page. Verify bucket terms. | Many evaluations use prerecorded teleoperation; actions are text answers, not low-level executable trajectories; intervention labels are acknowledged as imperfect. |
RoboMIND 2.0 RealTwinData | Six heterogeneous dual-arm/mobile embodiments; digital twins | 310k real dual-arm trajectories / 1,000+ h / 739 tasks; 20k simulation trajectories; tactile and mobile subsets.6↗ | Task language and frame-level fine-grained instruction updates; predecessor documents 5k failure demonstrations with detailed causes.7↗ | Failure causes, longer dual-arm episodes, tactile contact, and matched simulation are unusually relevant to recovery and sim-to-real. | Strong critic data Candidate for success/failure and recovery pretraining; audit exact V2 failure-label carryover before depending on it. | Project links downloads through ModelScope; release terms and per-subset licensing require inspection. | Embodiment heterogeneity and evolving versions complicate normalization; public metadata do not provide a single benchmark protocol as clean as RoboCasa/CALVIN. |
RoboTwin 2.0 SimDataBenchmark | SAPIEN; ALOHA-AgileX plus four other dual-arm configurations | 50 bimanual tasks; 50 clean demos/task in standard LeRobot protocol; a public unified mirror reports 100k+ trajectories overall.8↗,9↗ | Task-level natural-language descriptions with configurable paraphrase count; trajectories are not generally segmented into compositional subtasks. | True two-arm coordination, tool use, handovers, articulated objects; clutter, lighting, camera, table-height and language randomization. | Good transfer test Sparse terminal success supports offline/off-policy RL, but the high-level action vocabulary must be authored. | MIT code; downloadable dataset/mirror. Confirm asset-level terms. | Most tasks are short/fixed rather than branching; scripted experts bias trajectories; some upstream task bugs have been documented. |
CALVIN SimDataBenchmark | PyBullet tabletop; one Franka Panda | 34 atomic tasks; 24 h teleoperated play (6 h × 4 environments); evaluation uses 1,000 chains of five language goals; only ~1% / ~24k windows are language-annotated.10↗,11↗ | ~400 crowdsourced phrasings mapped to task IDs and temporal start/end indexes; annotations describe atomic skills such as opening a drawer or moving a block. | Continuous play, sequential evaluation, explicit unseen-environment split, cheap and mature. | Fast baseline Excellent smoke test for memory and hierarchical logging; can synthesize alternate subtask rollouts. | Open-source code and downloadable splits (166–656 GB); research use, inspect repository/submodule licenses. | Single scene grammar, single arm, five-task chains are often fixed and offer little semantic uncertainty; old Python/PyBullet stack. |
BEHAVIOR-1K / 2026 Challenge SimDataBenchmark | OmniGibson; mobile manipulators; rigid, articulated, deformable, fluid objects | 1,000 full activities / 50 scenes / 10k+ objects. Challenge data: 20k human teleop demos over 100 tasks, average 351.5 s and 27.1 labeled skills/episode.12↗,13↗ | Episode language plus 270,600 segments from 31 skill types (open/close, pour, chop, wipe, attach, etc.). Goals are first-order-logic BDDL predicates. | Best uncertainty/memory stress test: stateful household processes, long horizons, deformables and many valid orderings. | High scientific fit Skill labels and progress predicates enable high-level RL; computationally heavy and mobile. | Code/tasks public; asset dataset partly encrypted due to third-party licensing; 2026 demos ~3.27 TB. | Direct conflict with preferred tabletop/YAM embodiment; simulator complexity can dominate the research; sim-to-real is expensive. |
VLABench SimGenerated dataBenchmark | MuJoCo tabletop; single-arm manipulator | 100 categories: 60 primitive + 40 composite; 2,000+ objects; automated expert trajectory generation, but no stable canonical episode/hour total advertised.14↗ | Natural instructions include implicit intent, world knowledge, spatial/physical constraints; composite task language is task-level rather than a released universal subtask segmentation. | Excellent semantic/OOD evaluation: ordering books by publication year, commonsense object selection, sorting with implicit rules. | Planner benchmark Strong test of VLM decomposition; author subtask action schemas and reward checkpoints for RL. | Open code/data-generation repository; verify code and asset licenses separately. | Single arm; synthetic expert data; reasoning benchmark can conflate perception/world knowledge with control; scale not reported in one auditable unit. |
AgiBot World Beta RealData | Mobile dual-arm humanoid; grippers/dexterous hands/tactile; five domains | 1,001,552 trajectories / 2,976.4 h / 217 tasks / 87 skills / 106 scenes.15↗ | Task, scene, item, and subtask-segmented skill annotations; imperfect episodes include annotated error states and recovery. | Longer real tasks (typically 30–60 s), real operational scenes, explicit recovery, many atomic skills. | Strong semantics Valuable planner/critic corpus, but action-policy transfer to YAM is indirect. | Public Hugging Face release, CC BY-NC-SA 4.0. | Non-commercial/share-alike; mobile humanoid embodiment and high-dimensional hands diverge from the preferred platform. |
Language Table RealSimDataBenchmark | Planar x-y pushing of colored blocks, one low-cost arm | 442,226 real + 181,020 human-sim episodes; 206k unique language strings reported across the mix.16↗ | Every trajectory has a low-level natural-language command; relabeling yields huge paraphrase diversity and interactive commands can change online. | Exceptional language grounding and correction-by-new-command; clean test of paraphrase and relational generalization. | Language pretrain Useful for instruction encoder/low-level interface, weak for value learning. | Apache-2.0 repository and public TFDS/GCS data. | 2D pushing, no grasping or bimanual control, short skills, little plan branching—do not use as the main benchmark. |