Multi-Turn RL for Robotics Agents

Language-space reinforcement learning for hierarchical manipulation agents.

Updated

Offline Critic on Language Table Making Line Task

Setting

  • Base model: R3 mid-trained model
  • Training data: 16 seeds × 50 scenes per seed = 800 seed-scene instances, with 16 trials per instance (12,800 trajectories total). This includes the 64 seed-200k evaluation scenes (8 seeds × 8 scenes).
  • Evaluation: Language Table make_line_n_colors task; 16 trials per scene

Offline Critic Training

The offline critic estimates a scalar value for a multimodal state and language action. It is trained with TD(lambda): successful terminal actions receive reward 1, while failed trajectories receive reward 0. An EMA target critic supplies bootstrap values, and prioritized episode replay emphasizes episodes with higher TD error.

Training Configuration
  • Critic backbone: Qwen/Qwen3.5-2B
  • Objective: TD(lambda) with MSE loss, gamma = 0.95, and lambda = 0.8
  • Target critic: EMA with tau = 0.01, updated after every optimizer step
  • Adaptation: LoRA on all targets (rank 16, alpha 32, dropout 0.1) and a 256-wide MLP value head
  • Learning rates: 5e-5 for LoRA and 5e-4 for the value head; 100 warm-up steps
  • Replay: pure prioritized episode replay with alpha = 0.4 and full importance correction
  • Training: 16 transitions per GPU, gradient accumulation of 4, for 100 epochs

Results

The 64×16 critic checkpoint was trained for 1,500 steps. The 800×16 critic checkpoint was trained for 6,000 steps.

Seed 200k

Method pass@1 pass@2 pass@4 pass@8 pass@16 Status
standard 0.3398 0.5626 0.8035 0.9548 1.0000 complete
best of 8 (64×16 critic) 0.3281 0.5365 0.7645 0.9232 0.9844 complete
best of 8 (800×16 critic) 0.4639 0.6954 0.8844 0.9729 1.0000 complete

Seed 300k

Method pass@1 pass@2 pass@4 pass@8 pass@16 Status
standard 0.3359 0.5448 0.7766 0.9411 1.0000 complete
majority of 8 0.3701 0.5859 0.8079 0.9549 1.0000 complete
best of 8 (64×16 critic) 0.3145 0.5172 0.7468 0.9201 1.0000 complete
best of 8 (800×16 critic) 0.5176 0.7428 0.9109 0.9754 0.9844 complete

Survey on Open Source Datasets

This is a summary of some large open source datasets that might be helpful for robotic reasoning / language-based critic.

Comparison Table

Verified against official dataset cards, documentation, repositories, and papers on 2026-09-08. Counts are release-specific; “derived” values are arithmetic from reported statistics.

Dataset Tasks Complex tasks Episodes / demos Hours World Horizon Subtask annotation coverage
AgiBot Alpha 36 35/36 catalogued multi-step 92,214 trajectories* 595.31 h Real Typically 30 s; some >2 min All tasks; 185,722 action slices
AgiBot Beta 217 215/217 catalogued multi-step 1,003,672 trajectories* 2,976.4 h Real Typically 30 s; some >2 min All tasks; 1,000,041 action slices
RoboCasa365 365 300 composite 655,000 demos 2,290 h Simulation Mostly 10–60 s; some >3 min 32 tasks; 16,000 demos
ABC-130K 195 real; 20 sim Not reported 130,703 real episodes 3,590.7 real; ~400 sim Real + simulation 98.9 s mean; 469 s maximum 42,980 episodes; annotated task count unknown
BEHAVIOR-1K 1,000 Long-horizon; exact count unreported 20,000 demos across 100 tasks ~1,953 subset hours, derived Simulation Hundreds–thousands environment steps 100/1,000 tasks; 20,000 demos
Behavior-Skill 50 50 long-horizon 10,000 demos; 235,492 skills ~1,086 skill-hours, derived Simulation 23.55 skills; ~6.5 min/demo 50 tasks; 10,000 demos

Subtask-Annotation Notes

  • AgiBot Alpha/Beta: The official task catalog provides representative ordered action_text lists. Counting lists longer than one gives 35/36 Alpha and 215/217 Beta task types with multiple semantic steps; the remaining tasks have one step. All catalogued tasks have action-slice annotations, and the release schema says each episode’s task JSON stores timestamped atomic-skill/language slices. The catalog reports 185,722 Alpha and 1,000,041 Beta slices. Alpha is a Beta subset.
  • AgiBot count caveat: The current official repository reports 92,214 Alpha and 1,003,672 Beta “trajectories”; the Beta paper instead reports 1,001,552. The table uses the repository counts and preserves the publisher’s term because the public materials do not cleanly distinguish headline trajectories from episode directories/action slices. Alpha’s updated 595.31 h notice supersedes its stale 300 h headline.
  • RoboCasa365: It defines 65 atomic and 300 composite tasks, but only the 32 target composites have documented per-frame subtask index, atomic-skill, stage, and language labels. At 500 target demonstrations per task, this gives 16,000 annotated multi-step demos. No equivalent decomposition is documented for the other 268 composites. Most tasks have one or two stages; the catalog maximum is 16, and the target maximum is 15.
  • ABC-130K: Only 42,980 of 130,703 released episodes have annotation.mcap (32.9%). These contain timestamped subtask segments aligned with the parent task. Neither the card nor paper reports how many of the 195 task types are represented, how many annotated episodes contain multiple segments, or how many annotations are clean semantic decompositions. Sample-level filtering is therefore required.
  • BEHAVIOR-1K: “1K” means 1,000 activity definitions, not 1,000 demonstrated tasks. The full benchmark provides 1,000 BDDL specifications describing initial and goal predicates across 50 scenes, but these predicates are not ordered subtask decompositions. The 2026 trajectory release covers 100 of the 1,000 tasks: all 20,000 demos in that subset have ordered skill_annotation and finer primitive_annotation records, totaling 270,600 skills (27.06 per trajectory; 5.9 min average). Thus verified decomposition coverage is 100/1,000 tasks, not 1,000/1,000.
  • Behavior-Skill versus BEHAVIOR-1K: Behavior-Skill packages a separate 50-task/10,000-demo BEHAVIOR-1K subset as 235,492 richer natural-language skill intervals. All 10,000 demos are segmented; only 500 additionally receive restorable start states and local BDDL goals for independent skill evaluation. Full BEHAVIOR-1K offers 1,000 task definitions; its 2026 data provide broader demonstrated coverage (100 tasks), while Behavior-Skill provides stronger skill-level training/evaluation packaging (50 tasks). Their 31-versus-34 skill-category counts use different taxonomies.

Primary Sources

  1. AgiBot Alpha card, Beta card, official task catalog, repository, and paper
  2. RoboCasa365 project page and release notes, current dataset overview, current data format guide, and RoboCasa365 paper
  3. ABC-130K official dataset card, ABC project, and ABC paper
  4. BEHAVIOR-1K project, full-suite paper, 2026 demonstration documentation, Behavior-Skill dataset, and Behavior-Skill paper

Examples

AgiBot World Examples

task_327: 209 episode records; 209 with action_text; 209 multi-stage 7 distinct exact multi-stage label sequence(s): different stages across episodes episode 648649: Retrieve cucumber from the shelf. -> Place the held cucumber into the plastic bag in the shopping cart. -> Retrieve tomato from the shelf. -> Place the held tomato into the plastic bag in the shopping cart. -> Retrieve corn from the shelf. -> Place the held corn into the shopping cart’s plastic bag. episode 649755: Retrieve pear from the shelf. -> Place the held pear into the plastic bag in the shopping cart. -> Retrieve orange from the shelf. -> Place the held orange into the plastic bag in the shopping cart. -> Retrieve carambola from the shelf. -> Place the held carambola into the shopping cart’s plastic bag. episode 649765: Retrieve pear from the shelf. -> Place the held pear into the plastic bag in the shopping cart. -> Retrieve orange from the shelf. -> Place the held orange into the plastic bag in the shopping cart. -> Retrieve carambola from the shelf. -> Place the held carambola into the shopping cart’s plastic bag. episode 648756: Retrieve cucumber from the shelf. -> Place the held cucumber into the plastic bag in the shopping cart. -> Retrieve tomato from the shelf. -> Place the held tomato into the plastic bag in the shopping cart. -> Retrieve corn from the shelf. -> Place the held corn into the shopping cart’s plastic bag. episode 649743: Retrieve pear from the shelf. -> Place the held pear into the plastic bag in the shopping cart. -> Retrieve orange from the shelf. -> Place the held orange into the plastic bag in the shopping cart. -> Retrieve carambola from the shelf. -> Place the held carambola into the shopping cart’s plastic bag.

task_352: 1387 episode records; 1387 with action_text; 1387 multi-stage 12 distinct exact multi-stage label sequence(s): different stages across episodes episode 649841: Open the refrigerator door with the left arm. -> Pick up the dried strawberries from the fridge with the left arm. -> Place the dried strawberries held in the left arm on the table. -> Push the refrigerator door with the right arm. episode 650149: Open the refrigerator door with the left arm. -> Pick up the dried strawberries from the fridge with the left arm. -> Place the dried strawberries held in the left arm on the table. -> Push the refrigerator door with the right arm. episode 648643: Open the refrigerator door with the left arm. -> Pick up the cola from the fridge with the left arm. -> Place the cola held in the left arm on the table. -> Push the refrigerator door with the right arm. episode 648809: Open the refrigerator door with the left arm. -> Pick up the cola from the fridge with the left arm. -> Place the cola held in the left arm on the table. -> Push the refrigerator door with the right arm. episode 650491: Open the refrigerator door with the left arm. -> Pick up the dried strawberries from the fridge with the left arm. -> Place the dried strawberries held in the left arm on the table. -> Push the refrigerator door with the right arm.

task_354: 516 episode records; 516 with action_text; 516 multi-stage 33 distinct exact multi-stage label sequence(s): different stages across episodes episode 650263: Retrieve from the shelf. -> Place the held cookie biscuit into the shopping cart. episode 650390: Retrieve from the shelf. -> Place the held cookie biscuit into the shopping cart. episode 652724: Retrieve from the shelf. -> Place the held boxed chocolate cookies into the shopping cart. episode 650445: Retrieve from the shelf. -> Place the held cup jelly into the shopping cart. episode 652682: Retrieve from the shelf. -> Place the held boxed chocolate cookies into the shopping cart.

RoboCasa Examples

Task WashLettuce Episode 000140: 0. turn on the water in the sink (frames 0-260)

  1. pick up the lettuce from the counter (frames 261-586)
  2. place the lettuce in the sink (frames 587-706)
  3. task complete (frames 707-737) Episode 000214:
  4. turn on the water in the sink (frames 0-285)
  5. pick up the lettuce from the counter (frames 286-593)
  6. place the lettuce in the sink (frames 594-645)
  7. task complete (frames 646-707) Episode 000230:
  8. turn on the water in the sink (frames 0-179)
  9. pick up the lettuce from the counter (frames 180-325)
  10. place the lettuce in the sink (frames 326-402)
  11. task complete (frames 403-444)

Task WashFruitColander Episode 000000: 0. pick up the colander from the counter (frames 0-121)

  1. place the colander in the sink (frames 122-241)
  2. pick up the tangerine from the counter (frames 242-475)
  3. place the tangerine in the colander (frames 476-594)
  4. pick up the peach from the counter (frames 595-751)
  5. place the peach in the colander (frames 752-911)
  6. turn the sink spout to the left (frames 912-1102)
  7. turn on the sink faucet (frames 1103-1232)
  8. task complete (frames 1233-1248) Episode 000081:
  9. pick up the colander from the counter (frames 0-114)
  10. place the colander in the sink (frames 115-269)
  11. pick up the orange from the counter (frames 270-381)
  12. place the orange in the colander (frames 382-469)
  13. pick up the banana from the counter (frames 470-561)
  14. place the banana in the colander (frames 562-673)
  15. pick up the apple from the counter (frames 674-809)
  16. place the apple in the colander (frames 810-943)
  17. turn on the sink faucet (frames 944-1108)
  18. task complete (frames 1109-1124) Episode 000098:
  19. pick up the colander from the counter (frames 0-140)
  20. place the colander in the sink (frames 141-252)
  21. pick up the banana from the counter (frames 253-398)
  22. place the banana in the colander (frames 399-489)
  23. turn on the sink faucet (frames 490-732)
  24. task complete (frames 733-748)
ABC130k Examples

clip_the_socks_to_the_hanger: annotations: annotation.mcap subtask @ +0.000s: Open the clamp at the hanger. subtask @ +3.319s: Pick up the sock from the bin. subtask @ +5.941s: Clamp the sock onto the hanger. subtask @ +14.290s: Open the clamp at the hanger. subtask @ +21.880s: Pick up the sock from the bin. subtask @ +24.767s: Clamp the sock onto the hanger. subtask @ +28.860s: move the distractor subtask @ +32.484s: Open the clamp at the hanger. subtask @ +35.613s: Pick up the sock from the bin. subtask @ +38.337s: Clamp the sock onto the hanger. subtask @ +41.965s: Open the clamp at the hanger. subtask @ +45.962s: Pick up the sock from the bin. subtask @ +48.211s: Clamp the sock onto the hanger. subtask @ +50.931s: Open the clamp at the hanger. subtask @ +54.796s: mistake subtask @ +56.878s: Pick up the sock from the bin. subtask @ +58.284s: Clamp the sock onto the hanger. subtask @ +62.908s: Open the clamp at the hanger. subtask @ +69.258s: Pick up the sock from the bin. subtask @ +72.112s: Clamp the sock onto the hanger. subtask @ +76.237s: reset both arms

annotations: annotation.mcap subtask @ +0.000s: Pick up the sock from the bin. subtask @ +8.061s: Open the clamp at the hanger. subtask @ +10.744s: Clamp the sock onto the hanger. subtask @ +14.752s: Pick up the sock from the bin. subtask @ +22.461s: Open the clamp at the hanger. subtask @ +26.034s: Clamp the sock onto the hanger. subtask @ +28.154s: Pick up the sock from the bin. subtask @ +35.464s: Open the clamp at the hanger. subtask @ +40.821s: Clamp the sock onto the hanger. subtask @ +43.341s: Pick up the sock from the bin. subtask @ +51.874s: Open the clamp at the hanger. subtask @ +57.896s: Clamp the sock onto the hanger. subtask @ +63.358s: Pick up the sock from the bin. subtask @ +70.495s: Open the clamp at the hanger. subtask @ +76.826s: Clamp the sock onto the hanger. subtask @ +83.085s: Pick up the sock from the bin. subtask @ +93.922s: Open the clamp at the hanger. subtask @ +96.477s: Clamp the sock onto the hanger. subtask @ +99.102s: reset both arms

Behavior Skill Examples

Episode: skill_annotations/task-0033/episode_00330110_step.json Task: wash dog toys Long-horizon instruction: In the utility room, take the two teddy toys, the tennis ball, and the softball out of the cabinet and wash them in the washer so that both teddies are free of dirt and dust, the tennis ball has no debris, and the softball has no dirt. Short subtasks: 19 00 [84, 378]: Move to the washer. 01 [378, 1605]: Open the door of the washer using the left gripper. 02 [1605, 2250]: Move to the bottom cabinet in the utility room. 03 [2250, 2971]: Open the door of the bottom cabinet no top using the left gripper. 04 [2971, 4164]: Open the door of the bottom cabinet no top using the right gripper. 05 [4164, 4434]: Pick up the softball from the bottom cabinet no top using the left gripper. 06 [4434, 4770]: Push the teddy bear to the edge of the bottom cabinet no top using the right gripper. 07 [4770, 5001]: Pick up the teddy bear from the bottom cabinet no top using the right gripper. 08 [5001, 5438]: Move back to the washer. 09 [5438, 6067]: Place the teddy bear into the washer using the right gripper. 10 [6068, 6753]: Place the softball into the washer using the left gripper. 11 [6753, 7276]: Move back to the tennis ball. 12 [7276, 7616]: Pick up the tennis ball from the bottom cabinet no top using the left gripper. 13 [7616, 7866]: Pick up the other teddy bear from the bottom cabinet no top using the right gripper. 14 [7866, 8193]: Move back to the washer. 15 [8193, 8717]: Place the teddy bear into the washer using the right gripper. 16 [8717, 9300]: Place the tennis ball into the washer using the left gripper. 17 [9300, 9877]: Close the door of the washer using the left gripper. 18 [9877, 10230]: Turn on the switch of the washer using the right gripper.

Episode: skill_annotations/task-0033/episode_00330120_step.json Task: wash dog toys Long-horizon instruction: In the utility room, take the two teddy toys, the tennis ball, and the softball out of the cabinet and wash them in the washer so that both teddies are free of dirt and dust, the tennis ball has no debris, and the softball has no dirt. Short subtasks: 21 00 [0, 461]: Move to the washer in the utility room. 01 [461, 1557]: Open the door of the washer using the left gripper. 02 [1557, 2187]: Move to the bottom cabinet in the utility room. 03 [2187, 2789]: Open the door of the bottom cabinet no top using the left gripper. 04 [2789, 3827]: Open the door of the bottom cabinet no top using the right gripper. 05 [3827, 4195]: Push the softball to the edge of the bottom cabinet no top using the left gripper. 06 [4195, 4267]: Pick up the softball from the bottom cabinet no top using the left gripper. 07 [4267, 4417]: Push the teddy bear to the edge of the bottom cabinet no top using the right gripper. 08 [4417, 4665]: Pick up the teddy bear from the bottom cabinet no top using the right gripper. 09 [4665, 5039]: Move back to the washer. 10 [5039, 5822]: Place the teddy bear into the washer using the right gripper. 11 [5822, 6505]: Place the softball into the washer using the left gripper. 12 [6506, 6908]: Move back to the tennis ball. 13 [6909, 7410]: Pick up the tennis ball from the bottom cabinet no top using the right gripper. 14 [7410, 8181]: Push the teddy bear to the edge of the bottom cabinet no top using the left gripper. 15 [8181, 8382]: Pick up the teddy bear from the bottom cabinet no top using the left gripper. 16 [8383, 8804]: Move back to the washer. 17 [8804, 9301]: Place the tennis ball into the washer using the right gripper. 18 [9301, 10170]: Place the teddy bear into the washer using the left gripper. 19 [10170, 10692]: Close the door of the washer using the left gripper. 20 [10692, 11744]: Turn on the switch of the washer using the right gripper.

Episode: skill_annotations/task-0033/episode_00330140_step.json Task: wash dog toys Long-horizon instruction: In the utility room, take the two teddy toys, the tennis ball, and the softball out of the cabinet and wash them in the washer so that both teddies are free of dirt and dust, the tennis ball has no debris, and the softball has no dirt. Short subtasks: 18 00 [30, 312]: Move to the washer. 01 [312, 1426]: Open the door of the washer using the left gripper. 02 [1426, 1951]: Move to the bottom cabinet in the utility room. 03 [1951, 2731]: Open the door of the bottom cabinet no top using the left gripper. 04 [2731, 3516]: Open the door of the bottom cabinet no top using the right gripper. 05 [3516, 3939]: Pick up the softball from the bottom cabinet no top using the left gripper. 06 [3940, 4179]: Pick up the tennis ball from the bottom cabinet no top using the right gripper. 07 [4179, 4495]: Move back to the washer. 08 [4496, 5101]: Place the tennis ball into the washer using the right gripper. 09 [5102, 5490]: Place the softball into the washer using the left gripper. 10 [5490, 5882]: Move back to the teddy bear. 11 [5882, 6542]: Pick up the teddy bear from the bottom cabinet no top using the right gripper. 12 [6542, 7278]: Pick up the other teddy bear from the bottom cabinet no top using the left gripper. 13 [7278, 7674]: Move back to the washer. 14 [7674, 8362]: Place the teddy bear into the washer using the right gripper. 15 [8362, 8994]: Place the teddy bear into the washer using the left gripper. 16 [8994, 9585]: Close the door of the washer using the left gripper. 17 [9585, 10379]: Turn on the switch of the washer using the right gripper.

Language-Space RL: Project Framing and Surveys

Project Framing

The proposed agent separates semantic decision-making from motor control. A high-level vision-language model observes the task, scene, and recent execution history, then selects a concise language subtask. A low-level vision-language-action policy executes that subtask and returns feedback for the next decision. This hierarchy makes the planning interface interpretable while retaining closed-loop correction.

The central research question is whether reinforcement learning over language subtasks can improve long-horizon manipulation beyond an end-to-end policy. The intended settings should make reasoning consequential: plans must adapt to uncertain observations, failures must be recoverable, and success may depend on memory or on interpreting an underspecified goal. Evaluation should also test transfer across tasks that share strategy or reusable atomic skills rather than only measuring in-distribution execution.

The preferred embodiment is a bimanual YAM-like setup. The low-level policy needs reliable instruction following for short actions and reasonable transfer to unfamiliar scenes. It can be developed with a strong pretrained policy, broad demonstration datasets with aligned instruction labels, and smaller controlled domains where targeted data collection makes high-quality execution practical.

Learning Direction

Simulation is primarily a proof-of-concept and ablation environment; the main target is one or two real-world task families. Since physical interaction is expensive, the algorithm should favor offline or off-policy learning and reuse logged trajectories instead of depending on parallel, rollout-heavy training. The main value-learning object is therefore a state value or, more directly, an action value over language subtasks: Q(state, language instruction).

Survey Reports

The reports below provide the current evidence base and candidate experimental directions.