Stateconversation/history + environment observation
Actionwhole utterance; tokens form inner MDP
Rewarddelayed environment return
Policy/dataoff-policy critic; replay; online + offline forms
Mechanism. The high-level critic learns utterance Q and V with target-network Bellman bootstrapping. The actor’s token log-probabilities are weighted by the utterance advantage. An offline variant uses IQL-style expectile V estimation plus advantage-weighted regression (AWR), explicitly preventing unconstrained out-of-distribution language actions. The authors report about 100× sample efficiency relative to existing methods across Twenty Questions, Guess My City, a detective game, and WebShop; their PPO implementation required at least 1,024 new trajectories per update.1↗
Burden/evidence. Research code supports replay buffers and GPT-2/RoBERTa-scale actor/critic models, but the repository warns that even Twenty Questions can take a week on one GPU.3↗ Evidence is multi-turn but text-only; 7B was the largest reported model and no physical deployment was attempted.
Exact transfer. Make each completed VLA subtask one utterance-level transition. Train double Q(s,a) and expectile V(s) from all robot logs; update a VLM subtask generator/ranker by AWR. Preserve a frozen or slowly updated VLA. This is the report’s primary blueprint.
Statetextual interaction state/history
Actionopen-ended language action
Critictext description/critique of future, not scalar
Policy/dataoff-policy; prioritized replay; distillation
Mechanism. A language successor model predicts future consequences; a language Bellman target merges the observed one-step successor with predicted downstream futures. A language evaluator turns this into an action critique, and a refinement policy proposes a better action. The policy distills refinements rather than using a policy-gradient ratio. Priority is tied to critic loss so replay emphasizes likely-suboptimal steps. On 20 Questions and τ-retail, the paper reports roughly 30% improvement over standard RL fine-tuning on multi-step tasks.4↗
Limitations. Its convergence claims rely on strong ordering and representation assumptions; one model plays several prompted roles; evaluations remain language/tool tasks. A fluent critique can be causally wrong. No code link was identifiable on the canonical paper page at search time.
Exact transfer. Pair scalar Q with a generative “future consequence” head: “if the VLA attempts a, likely object-state change / failure / recovery is …”. Use it to relabel or refine bad logged subtasks while scalar Q remains the decision/safety anchor. Especially attractive when random free-form instruction exploration is unsafe.
Statetext observation / anchor state
ActionLLM reasoning-action turn
Creditepisode macro + same-state micro advantage
Policy/dataon-policy grouped rollouts; critic-free
Mechanism. Complete trajectories from the same task give a macro relative advantage. Repeated observations across those trajectories form anchor groups; actions from the same state are compared to produce micro advantages without an auxiliary critic or extra rollouts. Reported gains exceed 12 points on ALFWorld and 9 points on WebShop over GRPO at the same rollout and GPU-memory budget.5↗
Limitations. “No extra rollouts” means no extra relative to grouped GRPO—not few rollouts. Exact repeated states are common in deterministic text simulators and scarce in camera-based robotics. It cannot naturally exploit an arbitrary historical behavior mixture.
Exact transfer. In simulation, reset to saved states and branch 4–8 language subtasks to create a clean action-ranking set. In reality, use naturally recurring canonical states or learned state clusters cautiously; do not make grouped branching the core real-data pipeline. Code.
Statestate + explicit thinking
Actionenvironment-facing action turn
Rewardtrajectory result; shaping studied
Policy/dataon-policy trajectory optimization
Mechanism/evidence. StarPO structures state–thinking–action–reward trajectories. StarPO-S adds trajectory filtering, a critic, and decoupled clipping after observing an “Echo Trap”: loss of reward variance accompanied by gradient spikes. Three stylized environments support useful findings—diverse initial states, medium action granularity, frequent resampling, and fine-grained reasoning-aware reward matter.6↗
Limitations. Primarily a stability/diagnostic contribution on small synthetic environments; still rollout-centric, and its reward-shaping conclusions may be environment-specific.
Exact transfer. Instrument reward variance, gradient norms, action/reason repetition and subtask diversity. Reject degenerate trajectories and vary initial object configurations. Treat it as an experimental hygiene checklist, not the core algorithm. Code/environments.
Statemulti-turn tool-use history
Actionone LLM turn
Creditturn-level MDP advantages
Policy/dataon-policy; GRPO/PPO-compatible
Mechanism/evidence. Recasts a bandit-like trajectory objective as an MDP and estimates finer turn advantages. On search/tool-use tasks, the paper reports 100% tool execution and 50% exact match versus baselines with no tool invocation and 20–30% exact match.7↗
Limitations. Strong local results but limited environments; does not solve replay or distribution shift. Credit quality depends on trustworthy intermediate reward or a learned value estimator.
Exact transfer. Use as a controlled on-policy baseline showing whether temporal advantages themselves help. If it beats terminal-only GRPO but loses badly in real interactions per episode, that supports the off-policy thesis.
Statededuplicated observation node
Actionedge / state transition
Creditgraph-aggregated V + normalized TD edge error
Policy/datagroup-rollout graph
Mechanism/evidence. Merges identical states across trajectories into a global transition graph, estimates state values from grouped visits, and uses edge-centric TD errors. The preprint reports gains up to 22.2% over GRPO on WebShop, ALFWorld, and AppWorld.8↗
Limitations. Very recent preprint; identical observations need not be Markov-equivalent, especially under hidden robot state. Graph construction presumes enough revisitation.
Exact transfer. Build an offline semantic transition graph over object-centric scene states, but merge only when task, scene state, held objects, subtask outcome, and relevant history agree. Use graph TD error as a replay priority, not as the sole truth.
Stateagent execution context
Actionindividual LLM call
Credithierarchical decomposition module
Policy/datadecoupled runner/trainer; algorithm-pluggable
Mechanism. Treats arbitrary agent execution as an MDP, records calls as structured spans, decomposes trajectories into per-call training transitions, and decouples execution from training. Microsoft reports experiments on text-to-SQL, RAG and tool-use; the open-source system supports external harnesses.9↗
Limitations. Much of the contribution is infrastructure. Decomposition alone is not a solution to long-horizon credit; the actual reward propagation/optimizer choice remains decisive. Therefore it is included as a useful implementation layer, not ranked as a scientific blueprint.
Exact transfer. Adopt its trace schema idea: log VLM prompt, image/memory digest, proposed subtask, safety gate, VLA result, duration, intervention, and terminal outcome as separable spans. Feed those records to ArCHer/IQL rather than assuming the middleware’s default learner solves the problem. Code.
Statedialogue history / latent context
Actiongenerated response
Rewardgoal/task outcome
Policy/datafully offline static human conversations
Mechanism/evidence. CHAI combines a pretrained language model with offline RL to pursue task goals from static human dialogue. Earlier human-centric offline dialogue RL adds KL-control to a language prior and pessimism under uncertainty, explicitly targeting overestimated out-of-distribution actions.10↗, 11↗
Limitations. Dialogue dynamics, reward quality, and action consequences differ substantially from physical manipulation. These methods predate modern VLMs and evaluate relatively constrained language domains.
Exact transfer. Pretrain only within the support of demonstrated/replayed language subtasks; add behavior-model KL or conservative Q penalties; require human/VLM paraphrase augmentation so “support” reflects semantic equivalence rather than exact wording. CHAI code.
Stateweb page + trajectory
Actionweb action/reasoning step
CreditMCTS backed outcomes + self-critique preferences
Policy/dataoff-policy DPO on searched interactions
Mechanism/evidence. Guided MCTS branches from agent states, a self-critic supplies process guidance, outcomes are backed up, and an off-policy DPO variant learns preferences from both successes and failures. The paper reports large gains in WebShop and booking tasks after iterative collection.12↗
Limitations. Search multiplies interactions; reported booking numbers involve web environments and online search, not irreversible physical actions. DPO preferences do not necessarily produce a calibrated Q.
Exact transfer. Run MCTS only in simulation or a learned world model; distill pairwise “subtask a beats b at state s” comparisons into the VLM, then validate the best branch once in reality. Do not branch physical rollouts merely to produce preferences.
Statetextual environment state
Actionagent action
Rewardembedding sim-to-real-gap + success
Trainingself-supervised next-state alignment
Mechanism/evidence. Trains action-conditioned textual world models to align imagined and realized successor states in representation space. Combined with success reward, it reports +6.9 points on ALFWorld and +5.7 on τ²-Bench over direct success-reward RL, while matching expert-data training.13↗
Limitations. Text-state semantic distance is much easier than predicting contact-rich visual dynamics. A plausible verbal successor can conceal an impossible grasp or irreversible physical failure.
Exact transfer. Predict symbolic/object-centric deltas and option success—not pixels. Train from every transition. Use uncertainty to veto model-planned candidates; periodically recalibrate on real transitions.
Stategame/web text + history
ActionLM-generated candidate command
CriticRL re-ranker / environment score
Trainingcandidate LM + RL or IL→RL
Mechanism/evidence. CALM learns an action proposal distribution from human play and lets an RL agent re-rank the compact candidate set; it reported a 69% relative score improvement on unseen Jericho games.14↗ WebShop supplies a stateful web MDP with 1.18M products and 12,087 crowd instructions, plus IL, RL, and IL+RL baselines and public code/data.15↗
Limitations. Candidate bottlenecks cap performance when the right action is never proposed. Text games/websites have cheap resets and exact rewards.
Exact transfer. Separate proposal from evaluation: a pretrained VLM produces 4–16 short, executable subtask candidates; Q ranks them after feasibility filtering. Always include recovery, observe, and terminate candidates.
Statecurrent robot/environment state
Actionlanguage-described pretrained skill
Valueskill success/affordance probability
Traininglow-level skill RL; frozen LLM planner
Mechanism/evidence. Candidate skill score is the product of LLM usefulness and a learned affordance/value estimate; selection repeats after execution. It completed abstract long-horizon instructions on a real mobile manipulator.2↗
Why only adjacent. It does not learn a high-level policy from long-term task return, and the original system receives environment grounding through current-step skill values; its project page explicitly notes that failures or environmental changes may not be available to planning.
Exact transfer. Preserve SayCan’s feasibility factor p(success of VLA | s,a), but multiply/add it with a learned high-level Q for downstream task completion. Do not confuse option feasibility with task value: “can pick cup” is not “should pick cup now.”
Statetext embodied observation/history
Actioninteractive plan/action turn
Rewardsparse completion
Trainingpure online grouped RL
Mechanism/evidence. Interactive Policy Optimization trains from grouped in-environment exploration and sparse completion. Reported completion is 97.78% on ALFWorld and 79.92% on ScienceWorld, with a 3.66-point drop on unseen environments.16↗
Limitations. It explicitly relies on parallel exploration—the opposite of this project’s real-world constraint—and uses text simulations with resettable state. Despite “embodied” branding, it is not physical VLM control.
Exact transfer. Include only as an upper-bound simulation comparator and a reminder that pure sparse success can work when rollouts are abundant. It should not determine the real-robot algorithm.
Stateinteraction tree node
ActionLM-proposed environment action
ValueLM score + self-consistency; return backup
Trainingnone; inference-time search
Mechanism/evidence. MCTS-like selection/expansion/simulation uses actual environmental feedback, self-reflection and outcome backup. It reported strong results including 75.9 WebShop average score; code and trajectories are public.17↗
Limitations. It is not RL training and consumes multiple environment branches. LM self-values can be miscalibrated.
Exact transfer. Its node/value interface is useful for simulator evaluation and for planning inside a learned symbolic world model. Distill resulting branches to Q/policy; real execution remains single-path.