Survey 2 · language-space RL for robotics

Frontier real-world bimanual manipulation tasks

A critical map of what leading labs and companies have actually shown, what remains marketing, and which task families best isolate high-level language-subtask reasoning over a capable low-level VLA.

Search date: 3 September 2026Related-work depth: 1Real robot onlyPrimary sources prioritizedProject-specific ranking
Executive answer

Start with mixed-garment fold–stack and recoverable box assembly; add semantic table bussing as the cleanest generalization test.

These tasks fit a fixed YAM-like bimanual station, contain meaningful high-level choices, naturally generate recoverable failures, and admit cheap state resets and dense progress labels. Espresso is the strongest ambitious follow-up because it introduces waiting, hidden appliance state, cleanup and order memory. Whole-kitchen dishwasher demos are impressive but confound manipulation with mobility and expensive hardware.

Core finding 1

Frontier ≠ good benchmark

Figure's four-minute dishwasher run and Gemini Robotics 2 whole-body demonstrations extend the capability frontier, but their locomotion, five-finger sensing and proprietary platforms obscure whether gains come from the high-level language policy.

Core finding 2

Recovery is the research signal

The most useful evidence is not a clean success clip: it is Figure returning an accidental extra towel, Physical Intelligence re-folding a malformed box, Generalist regrasping a displaced washer, or Sunday's robot retrieving dropped clothes.

Core finding 3

The architecture is already plausible

π0.5 explicitly predicts semantic subtasks before low-level actions, while Figure Helix 02 describes a slow semantic system conditioning faster control. These are close conceptual precedents, though neither validates language-space off-policy RL as proposed here.

Core finding 4

Claims vary by orders of evidence

Sunday reports 785 autonomous folds with a fixed rubric; Figure reports continuous autonomy but mostly company-hosted evaluations; DYNA reports live deployments and throughput without public protocols; many glossy clips disclose neither trial count nor edit policy.

Scope, assumptions, and method

Inclusion required a physical robot performing a task that is bimanual or meaningfully benefits from two arms, with at least one frontier dimension: long horizon, state uncertainty, memory/replanning, dexterous contact, open-world variation, or deployment reliability.

Project interpretation

  • High level: a VLM observes current scene plus history and chooses a short language subtask.
  • Low level: a VLA executes the subtask on a fixed bimanual YAM-like platform.
  • Learning: off-policy/offline RL over language actions, likely with V(s) or Q(s, language-subtask).
  • Goal: reasoning, correction, exploration and memory—not merely longer motor trajectories.

Search protocol

  • Anchors: Physical Intelligence, Figure, AgiBot, DYNA, Sunday, Generalist.
  • Semantic expansion: laundry, food, dishes, cleanup, assembly, packing, tool use, ALOHA, real-robot RL.
  • Depth 1: anchor pages/papers plus their directly linked reports, project pages and canonical predecessors.
  • Relevance before popularity; real-robot evidence before simulation; primary sources before commentary.

Exclusions

  • Simulation-only tasks; locomotion-only feats; single-arm tasks without transferable bimanual structure.
  • Pure teleoperation footage presented only as data collection.
  • Fixed industrial automation with no meaningful variation or decision-making.
  • Claims with no task detail beyond a generic capability label.

What “duration” means

Reported duration is used only when the source documents it. A company compilation, sped-up montage, “hours of operation,” per-item cycle time, and one continuous episode are not treated as equivalent. “Unreported” is preferable to inference from video.

Evidence discipline

E1

Protocol evidence

Paper or disclosed multi-trial evaluation, metrics and conditions; strongest here, still often author-evaluated.

E2

Continuous autonomy

Uncut/continuous autonomous run or deployment claim with useful operating statistics; protocol may remain private.

E3

Company demo

Official clip or blog says autonomous, but trial count, selection, edit policy or failure distribution is absent.

E4

Marketing/preview

Capability named or visually shown without enough disclosure to establish autonomy or repeatability.

Default skepticism. “Fully autonomous” establishes absence of live control in the shown run only. Unless a source provides a protocol, this survey does not infer typical success, zero resets outside the frame, independence of test objects, or lack of cherry-picking. Company superlatives (“first,” “state of the art,” “commercial-grade”) remain attributed claims.

Frontier landscape at a glance

Task / organizationConcrete sequence & horizonWhy it is frontierBimanual needPlatformAutonomy / evidenceReproducibility
Mixed laundry
Physical Intelligence π0 / π*0.6
Retrieve crumpled item → flatten → identify fold strategy → fold → stack; π0 eval capped at ≈5 min/item. π*0.6 shows 50 novel items in a new home over hours.Deformable state, item-dependent strategy, cascading alignment error, re-fold/re-stack recovery.Intrinsic hold/tension, paired corner manipulation.Fixed and mobile dual-arm variants.E1/E2 paper scoring + company-hosted long runs; later RL includes autonomous trials and interventions.π0 weights/code open, but task data/hardware setup and π*0.6 weights are not fully reproducible.
In-home laundry
Sunday ACT-2
Retrieve from pile/basket/floor → reorient → fold → stack; median 2:13 successful fold including retries.785 attempts, 9 garment classes, unseen homes, broad size/lighting/position variation and disturbances.IntrinsicMemo wheeled dual-arm home robot.E1 99.1% company-graded with fixed rubric; all eval videos linked, but no independent audit.Closed model, robot and adaptation pipeline.
Towel folding
Figure Helix
Pick from mixed pile → separate multi-pick → unravel/trace edge → fold.Multi-finger manipulation, tangled towels, learned multi-pick recovery.IntrinsicFigure 02 humanoid, dexterous hands.E3 official “fully autonomous”; no trials, timing or failures.Closed.
Napkin/towel production
DYNA
Retrieve → orient → multi-step fold → stack/bin; production claims span shifts and months.Deployment reliability and throughput, not maximal task diversity.IntrinsicProprietary dual-arm workcell.E2 company says 200k garments/3 months at one laundromat and 99% acceptance; no public audit/protocol.Closed; customers and real site add credibility, but claims are first-party.
Dishwasher round trip
Figure Helix 02
Walk/open → unload → carry/stack/store → return → load → start; 4 min, 61 actions.Whole-body control, tactile/in-hand sensing, task memory, many object transfers.Intrinsic + whole bodyFigure 03 humanoid.E2 claimed continuous, no reset or intervention; single highlighted run, no success rate.Closed and hardware-confounded.
Clear table / load dishwasher
Sunday Memo
Collect mixed dishes/utensils/napkins → dump scraps → trash waste → rack items → run machine.Semantic sorting, fragile objects, contamination, appliance state and long horizon.StrongMemo mobile dual-arm.E4 product claim/demo; Sunday says reliability/generalization still improving.Closed; beta-stage.
Espresso drinks
Physical Intelligence π*0.6 / π0.7
Prepare portafilter → grind/tamp → mount → place vessel → actuate/wait → pour → wipe/clean; >90% claimed after RECAP and day-long run.Long horizon, liquid/contact, appliance latency, recipe memory and cleanup.Strong stabilize/operate, parallel preparation.Fixed dual-arm workcell.E1/E2 preprint metrics + 5:30–23:30 company video; edits/interruptions described as absent, independent audit absent.Model family partly open, target data and evaluation workcell closed.
Cook shrimp / wash pan
Mobile ALOHA
Oil → shrimp → hold pan + flip with spatula → turn/pour to bowl; 75 s demonstrations. Wash pan: hold pan, operate faucet, rinse.Heat, liquids, tool use, role specialization, mobile transitions.IntrinsicLow-cost mobile ALOHA.E1 research paper, autonomous policy evals trained from 50 demos/task; success varies by task.Code/hardware design/datasets available; closest open predecessor, though mobile.
Origami / Ziploc / salad
Google DeepMind Gemini Robotics
Multi-step folds; stabilize and insert snack then seal; manipulate ingredients/tools.Fine contact, deformables, precise sequencing, replanning after disturbances.IntrinsicALOHA 2 / bi-Franka; later humanoids.E3 official demos and model pages; limited task-specific protocol.Private model; hardware platforms are accessible.
Box assembly
PI / Generalist
Pick blank → open/fold body → hold geometry → tuck flaps → center/label; ≈34 s prior systems, GEN-1 claims 12.1 s.Deformable/semi-rigid contact, error propagation, stuck blanks, refolding.IntrinsicFixed industrial dual arms.E1 PI task scoring and RL study; E2 GEN-1 200-run claim at 99%, company evaluated.Partial: π0 code/weights; datasets and exact cells closed. GEN-1 closed.
Automotive kitting
Generalist GEN-1
Select parts → orient → place/insert in kit; hour-long run; recover displaced washer via regrasp/extrinsic contact.OOD recovery, fine insertion, semantic part matching, production cadence.StrongFixed dual-arm.E2 official claims fully autonomous and one hour uninterrupted; no independent protocol.Closed; ≈1 h task-specific robot data claimed.
Package reorientation
Figure Helix
Find barcode → regrasp/flip/flatten mailer → orient for scan; ≈4.05–4.31 s/package reported.Deformable packages, temporal visual memory, force feedback, concurrent arm use.Opportunistic one arm sufficient, two improve throughput/regrasp.Figure 02 humanoid upper body.E1/E2 60 h training-data ablation, ≈95% barcode orientation; company-run.Closed.
Dual-arm pot handling
AgiBot World Challenge
Grasp two handles → lift/transport/place under real-robot final conditions.Load sharing, coupled motion, collision/tilt and language-to-action transfer.IntrinsicAgiBot G2 dual-arm humanoid.E1 standardized competition; dataset has hundreds of trajectories/task, but public results are less granular.High relative to companies: dataset/baselines released; exact robot costly.

1. Laundry and clothing

The best-developed frontier class and the strongest immediate fit. Cloth makes the state partially observed: identical RGB contours can hide layers, inside-outness, entanglement and grasp topology.

Why the frontier moved

SpeedFolding established a transparent academic reference: 93% success and under 120 seconds on average from random garment configurations after 4,300 annotated/self-supervised actions. π0 expanded the sequence from a flat shirt to retrieval from a crumpled bin, flattening, folding and stacking. Figure highlighted end-to-end multi-finger towel recovery. Sunday then reported a much larger in-home evaluation: 785 attempts, 778 successes, nine classes and median 2:13 per successful item.

Comparability warning: garment sets, fold definitions, platforms and initial states differ substantially.

Where reasoning actually enters

  • Is this one garment or two? Return an accidental multi-pick before proceeding.
  • Is it inside-out, occluded or sufficiently flattened? Choose inspect, shake, spread or fold.
  • Which semantic class and orientation determines the fold program?
  • Did a sleeve escape or corner misalign? Re-open locally instead of restarting.
  • Where can the next folded item be stacked without destabilizing the pile?

Within-class comparison

VariantInitial uncertaintyContact / dexterityMemory & replanningOOD evidenceBest research use
Prepared towel foldLow–mediumCorner pinch, smoothingLow unless disturbedUsually unclearLow-level VLA calibration, not main high-level RL benchmark.
Crumpled single garmentHigh topology/orientationBimanual spread, tension, regraspTrack inspected orientation and failed graspsSpeedFolding unseen color/shape/stiffness; PI held-out items.Best Tier-A value-learning task.
Mixed pile → folded stackVery high; multi-pick and occlusionSeparation plus class-specific foldsItem count, stack state, recovery historySunday unseen homes/garments; PI novel items, both first-party.Best long-horizon extension.
Zipper / inside-out garmentHidden slider, cloth layersPrecision pinch, tension, constrained motionNeed local progress and jam detectionSunday preview only; Gemini unzipping demo.Excellent failure-recovery module after basic cloth competence.

2. Cooking and food preparation

Cooking adds delayed effects and irreversible errors. It is a more faithful test of task memory than folding, but heat, liquids, hygiene and reset burden raise real-world RL cost.

Espresso: unusually good hierarchy

  1. Read order; choose cup and recipe.
  2. Load/grind dose; wait for completion.
  3. Tamp while stabilizing portafilter.
  4. Mount, place cup, start extraction.
  5. Monitor/wait; optionally prepare milk in parallel.
  6. Serve; knock puck, purge and wipe station.

Physical Intelligence's RECAP report explicitly says its espresso task uses a high-level language policy, and that RECAP improves a VLA from autonomous experience plus interventions. That is the closest evidence to this project's intended loop. The catch is costly resets, fluids and appliance wear.

Cook shrimp: impressive, hazardous

Mobile ALOHA's 75-second demonstrations comprise oil pouring, raw shrimp transfer, one arm tilting the pan while the other manipulates a spatula, then transferring the cooked food. The sequence has genuine bimanual role asymmetry and physical state change. Yet doneness is hard to score visually, hot oil makes exploration unsafe, and cleaning between trials dominates throughput.

Snack bag / lunch packing

Gemini Robotics' Ziploc and lunch-packing demos combine deformable containers, semantic item selection, occluded insertion and seal verification. A safe lab version using dry packaged goods preserves nearly all of the reasoning value and is much easier to reset than cooking.

Frontier but poor first benchmark

Salad preparation and stir-fry are visually compelling, but continuous material state, cutting hazards, contamination and ambiguous outcome metrics make them poor early real-world RL tasks. Use pre-portioned ingredients and blunt tools if included later.

3. Household cleanup and dishes

Cleanup is where semantic reasoning dominates: the goal specifies an end state, not a fixed sequence. Every scene induces a different plan.

Figure Helix 02: frontier ceiling

The official report describes a continuous four-minute, 61-action dishwasher episode: unload, carry, stack/store, return, reload and start, with no intervention or reset. It also claims implicit error recovery and bimanual transfers. This is strong qualitative evidence of horizon, but one highlighted company run cannot establish typical reliability. It furthermore depends on locomotion, balance, tactile fingertips and palm cameras absent from a YAM-like station.

Fixed-table distillation

Distill the same reasoning into table bussing: classify trash, utensils and dishes; clear occlusions; select a safe order; stabilize a bowl while removing nested cutlery; scrape a plate; place items in distinct racks; wipe only after objects are removed. PI's table-bussing task uses 12 objects and includes a chopstick on trash—an excellent partial-observability seed.

Best generalization axis: hold out category–destination rules rather than only colors or poses. Train “glass → rack A,” “napkin → bin,” and “fork → caddy”; test novel objects whose destination can be inferred from semantics, plus combinations of known skills under new clutter.

4. Assembly, fastening, and tool use

Semi-rigid assembly offers clearer rewards than cloth while preserving contact richness, long-range error propagation and intrinsically bimanual stabilization.

Recoverable box assembly

PI scores pickup, opening/folding, each flap closure and final centering; later RECAP work describes stuck-together blanks and malformed folds that must be undone. Generalist reports 200 consecutive autonomous box folds, 99% success and 12.1-second cycle time after roughly one hour of robot data. These are self-reported and not apples-to-apples with PI's complete box task, but together establish both frontier speed and rich recoveries.

Kitting and insertion

Generalist's automotive-kitting demo is especially relevant because the disclosed failure responses are discrete: set a displaced washer down and regrasp, exploit a slot for extrinsic reorientation, or recruit the second hand. A language policy can select among these strategies while a VLA executes each.

One-shot LEGO copying

Generalist presents end-to-end LEGO assembly from a demonstrated target structure. The task is fancy and tests goal interpretation, but exact part pose and high insertion precision can make success mostly a low-level bottleneck. Use larger magnetic or snap-fit pieces first, and reserve a held-out target structure for compositional evaluation.

PaperBot: an outer-loop frontier

PaperBot uses a real two-arm system to fold and throw paper airplanes, and to cut paper grippers, optimizing designs where simulation is inaccurate. It is compelling real-world learning, but its action space is design parameters and scripted skills rather than language-subtask planning. It is a useful later test of high-level exploration, not the first benchmark.

5. Logistics, packing, and commercial repetition

Logistics supplies scale and clear metrics, but many tasks are short-horizon. The research value comes from memory during inspection, OOD packaging and recovery—not repeated pick-place.

Package reorientation

Figure reports ≈95% barcode orientation and roughly 4.05–4.31 seconds per package after scaling curated demonstrations from 10 to 60 hours. Temporal vision memory remembers inspected sides; force feedback helps contacts; the policy smooths wrinkled mailers and uses both hands opportunistically. This is credible company-run quantitative evidence, but the unit task is too short for the proposed high-level RL unless expanded to multi-package sorting with limited bins and exception handling.

AgiBot's reproducibility advantage

AgiBot World reports more than one million real trajectories across 217 tasks, with step/segment-level language in its 2026 release. Its challenge includes logistics sorting, shelf straightening, popcorn scooping, desk clearing and dual-arm pot handling, each with hundreds of trajectories and a real-robot final. Public data and baselines make it more reproducible than startup demos, although the exact G2 hardware remains costly and task-level autonomous results are incompletely reported.

DYNA: deployment over spectacle

DYNA's strongest evidence is operational: a customer testimonial claims 200,000 garments in three months with 99% quality acceptance, while a later technical post reports deployment logs and 2.7× napkin throughput improvement. Its homepage also claims dryer-to-folded/stacked/bagged and restaurant bus-tub-to-rack workflows. The public evidence is first-party, protocols and autonomy interventions are unclear, and product pages may change; treat figures as product claims, not peer-reviewed findings.

Long-horizon upgrade

Create an order-fulfillment episode: read a multi-line order, inspect candidate items, pack fragile/heavy/food items in a safe order, handle an unavailable or damaged item, close and label the box, and verify completion. This forces memory and semantic constraints while staying at one fixed station.

Ranked shortlist for this project

Scores weight reasoning necessity (30%), fixed-bimanual fit (20%), reset/RL feasibility (20%), evaluation clarity (15%), and generalization design (15%). “Frontier glamour” is intentionally not a scoring term.

RankTask setScoreWhy nowHigh-level language actionsFeasibility
1Mixed garment → fold → stack, with perturbations91/100Strongest combination of bimanual necessity, partial observability, recoveries and frontier relevance.separate extra item; inspect orientation; turn inside-out; spread; align corners; fold left/right/body; repair sleeve; re-stackTier A begin towels/T-shirts; expand classes.
2Recoverable box assembly + pack + label88/100Cheap reset, explicit milestones, contact-rich and demonstrated real-world RL precedent.separate blanks; open body; square base; close flap A/B; refold; insert items; close lid; orient/apply labelTier A
3Semantic table bussing and wipe-down86/100Best test of VLM semantics and plan ordering; many combinatorial OOD splits.inspect occluded item; remove cutlery; sort trash/dish; scrape; rack; consolidate; wipe spill; verify clearTier A
4Snack/lunch order packing in deformable bags82/100Safe analogue of cooking; semantic constraints, bag opening, insertion order and sealing.read order; open/stabilize bag; choose item; orient; insert; protect fragile item; seal; verify countTier A/B
5Espresso workstation + cleanup78/100Best ambitious memory/waiting task and direct language-policy + RL precedent.select recipe; grind; tamp; mount; brew; wait/check; pour; discard puck; purge; wipe; recover jamTier B use dry/mock machine first.
6Heterogeneous kitting with insertion recovery76/100Clear rewards and strategy choice; strong industrial relevance.inspect part; choose slot; reorient on table; stabilize; insert; retry alternate grasp; mark exceptionTier B
7Dish rack loading with fragile/cluttered items69/100Rich ordering and geometry, but breakage and rack access complicate RL.empty scraps; separate nested item; open rack; place by class; rearrange collision; close/verifyTier B/C plasticware first.
8Cook-and-serve with pan/tool55/100True bimanual roles and irreversible state, but safety and reset cost swamp early algorithm work.pour; hold pan; stir/flip; check state; adjust heat; transfer; rinseTier C

Benchmark-ready decompositions

1

Laundry-LSRL

Episode: three-item mixed pile to stable stack. Randomize item classes, inside-out state, entanglement, table texture and stack location.

Reward: item separation → coverage/orientation → keypoint alignment → fold quality → stable stack; penalize dropped/off-table items and unnecessary retries.

Recovery events: multi-pick, failed corner pinch, escaped sleeve, fold occludes orientation cue, dropped garment, stack collapse.

Why language Q-values matter: after a bad fold, “fold body” may look locally plausible but “reopen and realign left sleeve” has higher long-term value.

2

Box-Repair

Episode: separate one blank, form box, insert order-specific items, close and label.

Reward: single blank, squared geometry, valid flap order, item checklist, lid closure, label pose.

Recovery events: double blank, inverted box, flap trapped under wall, partial collapse, wrong item, label upside down.

Key ablation: fixed scripted subtask order versus VLM-selected language actions with memory and a learned value.

3

Bus-and-Wipe

Episode: clear 8–12 mixed items into three destinations, then wipe a spill without contaminating clean objects.

Reward: correct destination, recovered occluded/nested items, spill coverage, no fragile collision, final semantic inventory.

Recovery events: utensil hidden under napkin, cup contains trash, two objects grasped, rack slot blocked, cloth dropped, spill expands.

Memory test: briefly visible instruction card is removed after the first action; agent must retain sorting rules.

4

Pack-by-Constraint

Episode: satisfy a textual order and packing constraints in a deformable zipper bag or box.

Reward: correct SKU/count, heavy-before-fragile order, compartment use, sealed closure, barcode visible.

Recovery events: unavailable SKU, damaged package, bag mouth collapses, item inserted in wrong compartment, zipper jams.

Generalization: test novel items whose attributes (“fragile,” “keep upright”) can be inferred visually or linguistically.

Suggested action vocabulary

Keep actions short, visually verifiable and compositional: inspect(X), separate(X,Y), reorient(X,goal), stabilize(X), open(container), insert(X,slot), spread/fold(X,region), undo(last), wait_until(event), verify(predicate). Natural-language surface forms can vary while mapping to a canonical semantic action for credit assignment.

Evaluation splits and measurements

SplitTrainTestWhat it isolatesAnti-shortcut rule
CompositionalAtomic skills and some two-skill chainsNew valid sequences of known skillsLanguage planning and value compositionMatch visual backgrounds across train/test.
Semantic OODSubset of garment/item categoriesNovel categories with inferable propertiesVLM knowledge transferBalance color/shape so class cannot be guessed from a nuisance cue.
State OODCommon initial statesInside-out, entangled, double-pick, blocked slotReplanning under partial observabilityDo not include identical perturbations in demonstrations.
Goal OODKnown goals/rulesNew fold style, sorting rule or packing constraintInstruction understandingUse the same objects under different goals.
Recovery holdoutSuccesses + subset of failuresNovel failures generated during executionValue-driven correctionReport first-attempt and eventual success separately.
MemoryInstruction always visibleInstruction removed; delayed appliance/event stateHistory use versus reactive policyRandomize delays and insert distractor actions.
EnvironmentTwo tables/cameras/lighting regimesHeld-out surface, lighting, camera perturbationVisual robustnessKeep task semantics constant.

Primary metrics

  • End-to-end strict success and graded task progress.
  • Successful completions/hour, including reset and recovery time.
  • Autonomous recovery rate conditioned on a failure event.
  • Language-action efficiency: useful subtasks / all issued subtasks.
  • Intervention rate, unsafe-stop rate and object damage.

Architecture-specific metrics

  • Value calibration by subtask and horizon.
  • Regret versus oracle subtask choice at decision points.
  • State aliasing errors with and without history.
  • Generalization gap by split, not only pooled average.
  • Ablations: scripted planner, greedy VLM, VLM+history, language-space RL.

Evidence limits, blind spots, and stopping rationale

What remains unknown

  • Most companies do not release per-trial logs, intervention policies, reset rules or selected-vs-total video counts.
  • “Autonomous” clips rarely disclose whether a human chose the initial state or accepted/rejected preceding attempts.
  • Cross-company success rates are not comparable because task definitions and strictness differ.
  • Fine-grained platform differences—gripper friction, tactile sensors, compliance and workspace—may dominate task success.
  • Public evidence from Chinese startups is often dataset- or event-centric rather than granular autonomous evaluation.

Coverage gaps

  • No exhaustive crawl of social video platforms; short clips without durable technical descriptions were deliberately downweighted.
  • Some embedded videos could not be inspected frame-by-frame through text search; claims use surrounding official descriptions.
  • Depth 1 misses distant citation descendants and non-English reporting.
  • DYNA and Sunday pages dated after early 2026 are unusually current and may evolve; archived versions were not available in this pass.

Anchor-bias check

The six requested companies were complemented with Google DeepMind, Generalist, Mobile ALOHA/ACT, SpeedFolding and PaperBot. This adds academic/open baselines, not just humanoid-product narratives. Repeated laundry examples reflect genuine field concentration, but also reveal selection bias toward visually legible demos.

Stopping rationale

Search stopped when each major task class had: (1) at least one primary frontier example, (2) a documented evidence caveat, and (3) a translation to a fixed-bimanual language-action benchmark; and additional results were repeating task labels without stronger protocols. Maximum related-work depth remained 1 as specified.

Canonical linked bibliography

Source labels describe the evidence type, not an endorsement. All links accessed for this survey on 3 September 2026.

  1. Physical Intelligence, “π0: Our First Generalist Policy” (2024) and technical report. Preprint + company project page Task definitions, platforms, evaluation windows and videos.
  2. Physical Intelligence, “π0.5: a VLA with Open-World Generalization” (2025) and paper. Preprint + company project page Unseen-home cleanup and explicit high-level subtask prediction.
  3. Physical Intelligence, “π*0.6: a VLA That Learns from Experience” (2025) and arXiv record. Preprint + company evaluation RECAP, espresso/laundry/box RL and long runs.
  4. Physical Intelligence, “π0.7: a Steerable Model with Emergent Capabilities” (2026). Company research release Cross-task and zero-shot claims.
  5. Figure, “Helix Learns to Fold Laundry” (2025). Company demo Mixed towel pile and multi-pick recovery.
  6. Figure, “Scaling Helix: a New State of the Art in Humanoid Logistics” (2025). Company quantitative report Package timing, barcode rate, memory/force modules.
  7. Figure, “Introducing Helix 02: Full-Body Autonomy” (2026). Company demo/report Four-minute dishwasher run and tactile dexterity.
  8. Sunday Robotics, “ACT-2 Preview: Generalizing Reliability” (2026). Company protocol report 785-fold evaluation, scope, speed, quality and recovery videos.
  9. Sunday Robotics, Memo product page (2026). Product claim Dishes, coffee and laundry; explicitly beta-stage.
  10. DYNA, Monster Laundry customer case study (2025) and “Not Just a Model, But a Product” (2026). First-party deployment claims
  11. Generalist, “GEN-1: Scaling Embodied Foundation Models to Mastery” (2026). Company quantitative report Kitting, box folding, phone packing, autonomous streaks and recovery.
  12. Generalist, “The Robots Build Now, Too” (2025). Company demo One-shot LEGO copying.
  13. Google DeepMind, “Gemini Robotics brings AI into the physical world” (2025) and model page. Company research release Origami, Ziploc, salad and disturbance replanning.
  14. Google DeepMind, “Gemini Robotics 2 brings whole body intelligence to robots” (2026). Company research release Knot tying, bag sealing, tight packing and whole-body tasks.
  15. Fu et al., “Mobile ALOHA” project page and paper (2024). Research paper + open project Cook shrimp, wash pan, cabinet and elevator tasks.
  16. Zhao et al., “Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware” (RSS 2023). Peer-reviewed + open project ACT, battery insertion and condiment-cup opening.
  17. Avigal et al., “SpeedFolding” (IROS 2022) and code. Peer-reviewed + open code
  18. AgiBot World Colosseo repository, paper, and 2026 challenge task page. Open dataset/model + competition
  19. Liu et al., “PaperBot: Learning to Design Real-World Tools Using Paper” (2024). Research project + code/paper

Prepared as Survey 2 for the Language-Space RL for Robotics project. Claims labeled as inference are the survey author's synthesis; company claims remain attributed. No external JavaScript is required.