Survey 3 · language-space control

Agentic RL for LLM/VLM Agents

A selective map of reinforcement-learning methods that treat interaction as a temporal decision problem—not a long transcript masquerading as a one-step completion—and a concrete offline-to-online design for a VLM choosing language subtasks for a low-level VLA.

Search date: 3 Sep 2026Citation depth: 1Primary-source weightedReal-rollout budget firstSelf-contained HTML
Executive answer

Learn Q over semantic decisions; do not RL-train the whole transcript.

Q(s, language subtask)

is the right primary object.

It can rank several feasible instructions at the same visual/history state and admits one-step Bellman updates from a replay buffer. A pure V(s) only scores states; action selection still needs candidate sampling, a model, or Q-like lookahead.

is the closest existing algorithmic template.

ArCHer explicitly learns utterance-level Q/V off-policy, then sends the utterance advantage into token-level policy learning. Its reported ~100× sample-efficiency advantage over on-policy baselines is unusually aligned with costly robot rollouts, though it was tested on text agents, not robots.1↗

Factorize

usefulness × feasibility × long-term return.

SayCan already demonstrates language candidate scoring with learned robotic affordances, but its high-level planner is not learned by RL and its original feedback is myopic.2↗ The project should retain the safety/feasibility gate and add a high-level return critic.

Key negative conclusion. GiGPO, RAGEN/StarPO, turn-level GRPO and similar methods are useful controls and simulation baselines, but remain fundamentally on-policy/group-rollout methods. They improve multi-turn credit assignment, not real-world data reuse. For the intended project they should not be the main algorithm.

Question & boundary

What counted as agentic RL?

Included

  • An agent repeatedly observes environment feedback and chooses language/tool/semantic actions.
  • The training method explicitly addresses temporal credit, turn-level values, hierarchy, replay/offline data, counterfactual branches, a learned world model, or multi-agent non-stationarity.
  • Embodied work where language is an actual decision interface, plus adjacent robotics systems that expose the exact value factorization needed here.
  • Peer-reviewed work and recent preprints through the search date; source type is labeled.

Excluded or demoted

  • Single-response RLVR (math/code) without environment transitions.
  • “Multi-turn” GRPO that concatenates a trace, broadcasts one terminal reward, and otherwise changes no learning design.
  • Harness-only work that improves rollout plumbing but contributes no temporal optimization.
  • Prompt-only ReAct/reflection agents. Search-only LATS is retained only as an adjacent model-based design, not claimed as RL training.
  • Low-level continuous robot RL with no language-action decision layer.

Project assumption. At high-level time t, the VLM observes recent images, task goal, execution history, low-level success/failure signals, and optionally a compact memory; it emits a short language instruction executed for a variable-duration option by a mostly fixed VLA. A real rollout therefore yields a small number of semantically meaningful transitions, not thousands of motor actions.

Problem formulation

The literature separates into six useful families

FamilyRepresentative workWhere temporal structure entersData regimeFit to expensive robotics
Hierarchical actor–criticArCHerUtterance MDP with Bellman Q/V; token MDP receives utterance advantageOff-policy replay; online and offline variantsBest fit
Language-valued criticNLACLanguage Bellman backup predicts and critiques future consequencesOff-policy prioritized replay; no policy gradientsHigh-upside
Conservative offline language RLCHAI; human-centric dialogue RLSequence/utterance return under dataset support constraintsFully offline; KL/pessimism or CQL-style controlStrong prior
Grouped turn creditGiGPO; turn-level GRPO; G2POCompare actions from repeated/identical states; graph TD edgesMostly on-policy grouped trajectoriesSim ablation
Search / model-based improvementAgent Q; LATS; RWML; WorldCoderBranch, predict next state, critique or back up terminal outcomesOften environment- or inference-heavy; can distill offlineUse in sim / model
Embodied language agentsCALM, WebShop, Embodied Planner-R1; SayCan (adjacent)Language action candidates, closed-loop feedback, skill affordancesMixed; many are rollout-heavyEvidence bridge
Major works

Methods, evidence, limits, and exact transfer lessons

ArCHer — Actor-Critic with a Hierarchical Structure ICLR 2024

Stateconversation/history + environment observation
Actionwhole utterance; tokens form inner MDP
Rewarddelayed environment return
Policy/dataoff-policy critic; replay; online + offline forms

Mechanism. The high-level critic learns utterance Q and V with target-network Bellman bootstrapping. The actor’s token log-probabilities are weighted by the utterance advantage. An offline variant uses IQL-style expectile V estimation plus advantage-weighted regression (AWR), explicitly preventing unconstrained out-of-distribution language actions. The authors report about 100× sample efficiency relative to existing methods across Twenty Questions, Guess My City, a detective game, and WebShop; their PPO implementation required at least 1,024 new trajectories per update.1↗

Burden/evidence. Research code supports replay buffers and GPT-2/RoBERTa-scale actor/critic models, but the repository warns that even Twenty Questions can take a week on one GPU.3↗ Evidence is multi-turn but text-only; 7B was the largest reported model and no physical deployment was attempted.

Exact transfer. Make each completed VLA subtask one utterance-level transition. Train double Q(s,a) and expectile V(s) from all robot logs; update a VLM subtask generator/ranker by AWR. Preserve a frozen or slowly updated VLA. This is the report’s primary blueprint.

NLAC — Natural Language Actor-Critic 2025 preprint

Statetextual interaction state/history
Actionopen-ended language action
Critictext description/critique of future, not scalar
Policy/dataoff-policy; prioritized replay; distillation

Mechanism. A language successor model predicts future consequences; a language Bellman target merges the observed one-step successor with predicted downstream futures. A language evaluator turns this into an action critique, and a refinement policy proposes a better action. The policy distills refinements rather than using a policy-gradient ratio. Priority is tied to critic loss so replay emphasizes likely-suboptimal steps. On 20 Questions and τ-retail, the paper reports roughly 30% improvement over standard RL fine-tuning on multi-step tasks.4↗

Limitations. Its convergence claims rely on strong ordering and representation assumptions; one model plays several prompted roles; evaluations remain language/tool tasks. A fluent critique can be causally wrong. No code link was identifiable on the canonical paper page at search time.

Exact transfer. Pair scalar Q with a generative “future consequence” head: “if the VLA attempts a, likely object-state change / failure / recovery is …”. Use it to relabel or refine bad logged subtasks while scalar Q remains the decision/safety anchor. Especially attractive when random free-form instruction exploration is unsafe.

GiGPO — Group-in-Group Policy Optimization NeurIPS 2025

Statetext observation / anchor state
ActionLLM reasoning-action turn
Creditepisode macro + same-state micro advantage
Policy/dataon-policy grouped rollouts; critic-free

Mechanism. Complete trajectories from the same task give a macro relative advantage. Repeated observations across those trajectories form anchor groups; actions from the same state are compared to produce micro advantages without an auxiliary critic or extra rollouts. Reported gains exceed 12 points on ALFWorld and 9 points on WebShop over GRPO at the same rollout and GPU-memory budget.5↗

Limitations. “No extra rollouts” means no extra relative to grouped GRPO—not few rollouts. Exact repeated states are common in deterministic text simulators and scarce in camera-based robotics. It cannot naturally exploit an arbitrary historical behavior mixture.

Exact transfer. In simulation, reset to saved states and branch 4–8 language subtasks to create a clean action-ranking set. In reality, use naturally recurring canonical states or learned state clusters cautiously; do not make grouped branching the core real-data pipeline. Code.

RAGEN / StarPO 2025 preprint

Statestate + explicit thinking
Actionenvironment-facing action turn
Rewardtrajectory result; shaping studied
Policy/dataon-policy trajectory optimization

Mechanism/evidence. StarPO structures state–thinking–action–reward trajectories. StarPO-S adds trajectory filtering, a critic, and decoupled clipping after observing an “Echo Trap”: loss of reward variance accompanied by gradient spikes. Three stylized environments support useful findings—diverse initial states, medium action granularity, frequent resampling, and fine-grained reasoning-aware reward matter.6↗

Limitations. Primarily a stability/diagnostic contribution on small synthetic environments; still rollout-centric, and its reward-shaping conclusions may be environment-specific.

Exact transfer. Instrument reward variance, gradient norms, action/reason repetition and subtask diversity. Reject degenerate trajectories and vary initial object configurations. Treat it as an experimental hygiene checklist, not the core algorithm. Code/environments.

Turn-level credit assignment (MT-GRPO / MT-PPO) 2025 preprint

Statemulti-turn tool-use history
Actionone LLM turn
Creditturn-level MDP advantages
Policy/dataon-policy; GRPO/PPO-compatible

Mechanism/evidence. Recasts a bandit-like trajectory objective as an MDP and estimates finer turn advantages. On search/tool-use tasks, the paper reports 100% tool execution and 50% exact match versus baselines with no tool invocation and 20–30% exact match.7↗

Limitations. Strong local results but limited environments; does not solve replay or distribution shift. Credit quality depends on trustworthy intermediate reward or a learned value estimator.

Exact transfer. Use as a controlled on-policy baseline showing whether temporal advantages themselves help. If it beats terminal-only GRPO but loses badly in real interactions per episode, that supports the off-policy thesis.

G2PO — Group-Graph Policy Optimization 2026 preprint

Statededuplicated observation node
Actionedge / state transition
Creditgraph-aggregated V + normalized TD edge error
Policy/datagroup-rollout graph

Mechanism/evidence. Merges identical states across trajectories into a global transition graph, estimates state values from grouped visits, and uses edge-centric TD errors. The preprint reports gains up to 22.2% over GRPO on WebShop, ALFWorld, and AppWorld.8↗

Limitations. Very recent preprint; identical observations need not be Markov-equivalent, especially under hidden robot state. Graph construction presumes enough revisitation.

Exact transfer. Build an offline semantic transition graph over object-centric scene states, but merge only when task, scene state, held objects, subtask outcome, and relevant history agree. Use graph TD error as a replay priority, not as the sole truth.

Agent Lightning / LightningRL 2025 preprint + system

Stateagent execution context
Actionindividual LLM call
Credithierarchical decomposition module
Policy/datadecoupled runner/trainer; algorithm-pluggable

Mechanism. Treats arbitrary agent execution as an MDP, records calls as structured spans, decomposes trajectories into per-call training transitions, and decouples execution from training. Microsoft reports experiments on text-to-SQL, RAG and tool-use; the open-source system supports external harnesses.9↗

Limitations. Much of the contribution is infrastructure. Decomposition alone is not a solution to long-horizon credit; the actual reward propagation/optimizer choice remains decisive. Therefore it is included as a useful implementation layer, not ranked as a scientific blueprint.

Exact transfer. Adopt its trace schema idea: log VLM prompt, image/memory digest, proposed subtask, safety gate, VLA result, duration, intervention, and terminal outcome as separable spans. Feed those records to ArCHer/IQL rather than assuming the middleware’s default learner solves the problem. Code.

CHAI and conservative offline dialogue RL NAACL 2022

Statedialogue history / latent context
Actiongenerated response
Rewardgoal/task outcome
Policy/datafully offline static human conversations

Mechanism/evidence. CHAI combines a pretrained language model with offline RL to pursue task goals from static human dialogue. Earlier human-centric offline dialogue RL adds KL-control to a language prior and pessimism under uncertainty, explicitly targeting overestimated out-of-distribution actions.10↗, 11↗

Limitations. Dialogue dynamics, reward quality, and action consequences differ substantially from physical manipulation. These methods predate modern VLMs and evaluate relatively constrained language domains.

Exact transfer. Pretrain only within the support of demonstrated/replayed language subtasks; add behavior-model KL or conservative Q penalties; require human/VLM paraphrase augmentation so “support” reflects semantic equivalence rather than exact wording. CHAI code.

Agent Q 2024 preprint

Stateweb page + trajectory
Actionweb action/reasoning step
CreditMCTS backed outcomes + self-critique preferences
Policy/dataoff-policy DPO on searched interactions

Mechanism/evidence. Guided MCTS branches from agent states, a self-critic supplies process guidance, outcomes are backed up, and an off-policy DPO variant learns preferences from both successes and failures. The paper reports large gains in WebShop and booking tasks after iterative collection.12↗

Limitations. Search multiplies interactions; reported booking numbers involve web environments and online search, not irreversible physical actions. DPO preferences do not necessarily produce a calibrated Q.

Exact transfer. Run MCTS only in simulation or a learned world model; distill pairwise “subtask a beats b at state s” comparisons into the VLM, then validate the best branch once in reality. Do not branch physical rollouts merely to produce preferences.

RWML — Reinforcement World Model Learning 2026 preprint

Statetextual environment state
Actionagent action
Rewardembedding sim-to-real-gap + success
Trainingself-supervised next-state alignment

Mechanism/evidence. Trains action-conditioned textual world models to align imagined and realized successor states in representation space. Combined with success reward, it reports +6.9 points on ALFWorld and +5.7 on τ²-Bench over direct success-reward RL, while matching expert-data training.13↗

Limitations. Text-state semantic distance is much easier than predicting contact-rich visual dynamics. A plausible verbal successor can conceal an impossible grasp or irreversible physical failure.

Exact transfer. Predict symbolic/object-centric deltas and option success—not pixels. Train from every transition. Use uncertainty to veto model-planned candidates; periodically recalibrate on real transitions.

CALM and WebShop — language actions before modern LLM RL EMNLP 2020 / NeurIPS 2022

Stategame/web text + history
ActionLM-generated candidate command
CriticRL re-ranker / environment score
Trainingcandidate LM + RL or IL→RL

Mechanism/evidence. CALM learns an action proposal distribution from human play and lets an RL agent re-rank the compact candidate set; it reported a 69% relative score improvement on unseen Jericho games.14↗ WebShop supplies a stateful web MDP with 1.18M products and 12,087 crowd instructions, plus IL, RL, and IL+RL baselines and public code/data.15↗

Limitations. Candidate bottlenecks cap performance when the right action is never proposed. Text games/websites have cheap resets and exact rewards.

Exact transfer. Separate proposal from evaluation: a pretrained VLM produces 4–16 short, executable subtask candidates; Q ranks them after feasibility filtering. Always include recovery, observe, and terminate candidates.

SayCan — the closest robotics factorization (adjacent, not agent-RL training) CoRL 2022

Statecurrent robot/environment state
Actionlanguage-described pretrained skill
Valueskill success/affordance probability
Traininglow-level skill RL; frozen LLM planner

Mechanism/evidence. Candidate skill score is the product of LLM usefulness and a learned affordance/value estimate; selection repeats after execution. It completed abstract long-horizon instructions on a real mobile manipulator.2↗

Why only adjacent. It does not learn a high-level policy from long-term task return, and the original system receives environment grounding through current-step skill values; its project page explicitly notes that failures or environmental changes may not be available to planning.

Exact transfer. Preserve SayCan’s feasibility factor p(success of VLA | s,a), but multiply/add it with a learned high-level Q for downstream task completion. Do not confuse option feasibility with task value: “can pick cup” is not “should pick cup now.”

Embodied Planner-R1 / IPO 2025 preprint

Statetext embodied observation/history
Actioninteractive plan/action turn
Rewardsparse completion
Trainingpure online grouped RL

Mechanism/evidence. Interactive Policy Optimization trains from grouped in-environment exploration and sparse completion. Reported completion is 97.78% on ALFWorld and 79.92% on ScienceWorld, with a 3.66-point drop on unseen environments.16↗

Limitations. It explicitly relies on parallel exploration—the opposite of this project’s real-world constraint—and uses text simulations with resettable state. Despite “embodied” branding, it is not physical VLM control.

Exact transfer. Include only as an upper-bound simulation comparator and a reminder that pure sparse success can work when rollouts are abundant. It should not determine the real-robot algorithm.

LATS / tree search (adjacent model-based agent) ICML 2024

Stateinteraction tree node
ActionLM-proposed environment action
ValueLM score + self-consistency; return backup
Trainingnone; inference-time search

Mechanism/evidence. MCTS-like selection/expansion/simulation uses actual environmental feedback, self-reflection and outcome backup. It reported strong results including 75.9 WebShop average score; code and trajectories are public.17↗

Limitations. It is not RL training and consumes multiple environment branches. LM self-values can be miscalibrated.

Exact transfer. Its node/value interface is useful for simulator evaluation and for planning inside a learned symbolic world model. Distill resulting branches to Q/policy; real execution remains single-path.

Decision

Ranked algorithmic blueprints for expensive real rollouts

01

Offline ArCHer-IQL with semantic candidate ranking

Recommended core

One replay transition per executed VLA option; double Q over (visual/history state, language subtask); expectile V; AWR or weighted SFT actor. Highest direct evidence for off-policy multi-turn language learning and most compatible with heterogeneous logged robot behavior.

02

Conservative Q + SayCan feasibility gate

Safety layer

Factor downstream task return from immediate executability. CQL/KL support penalty prevents attractive but unsupported instructions. The low-level VLA success model can be pretrained independently and updated from every attempted option.

03

Hybrid scalar Q + NLAC consequence critique

Research upside

Scalar critic makes ranking and calibration testable; language successor/critique generates repairs and supplies interpretable failure hypotheses. More speculative than ArCHer, but unusually well matched to open-ended instruction actions.

04

Sim/model branching → preference distillation

Auxiliary

Agent Q or GiGPO-style branches from checkpointed simulator states generate action comparisons; distill them by DPO/AWR, then run only the chosen path on hardware. This converts cheap branching into offline data without demanding physical resets.

05

Semantic transition graph + TD replay priority

Experimental

G2PO-inspired merging may pool repeated scene configurations and identify pivotal subtask edges. Requires robust object-centric state identity and tests against perceptual aliasing.

06

On-policy turn-level GRPO

Baseline only

Valid multi-turn method and useful scientific control, but group rollouts are costly and old data expires. Use in simulation to measure the value of turn credit; do not center real-world learning on it.

Concrete proposal

A semi-MDP with Q over language instructions

Why Q, not only V?

  • The scientific question is which language subtask to choose now; Q exposes this comparison directly.
  • Q supports off-policy one-step updates from old trajectories.
  • Q can rank a bounded candidate set without sampling full future trajectories.
  • V remains valuable as the IQL expectile baseline and for early stopping, recovery, and state-progress diagnostics.

Action representation

  • Phase 1: canonical schema: verb, object(s), target/relation, constraints, stop condition.
  • Phase 2: allow free-form VLM proposals, map them to embeddings plus parsed slots.
  • Train Q on both structured slots and text/image embeddings; paraphrase consistency is an explicit regularizer.
  • Always offer observe closer, retry with correction, undo/recover, ask human, and terminate.

Important distinction. The low-level option-success model answers “will this VLA execute the instruction from here?” Q answers “does choosing it now increase eventual task return?” They must be separately supervised and jointly used.

Training recipe

Offline first, then budgeted online improvement

1 · Build replay

Segment demonstrations and autonomous runs at semantic option boundaries. Log failures, interventions, duration, and terminal outcomes—not only successes.

→

2 · Conservative offline RL

Behavior-clone the proposal model; train feasibility, double Q, expectile V, and AWR actor. Use simulator branches and paraphrase augmentation.

→

3 · Small online loop

Deploy conservatively, prioritize uncertain/high-value states, add transitions to replay, update critic frequently and actor slowly behind safety gates.

  1. Transition construction. Boundary = VLA stop detector, success detector, timeout, or high-level replanning. Store raw video pointers plus frozen visual embeddings and symbolic object-state deltas.
  2. Reward labeling. Prefer state-based task predicates. Add dense progress only if it is potential-based or separately ablated; never let an LLM judge be the only reward on hardware.
  3. Offline pretraining. BC/SFT on demonstrated subtasks; train psuccess; train Q/V with n-step or option-duration targets; apply conservative penalty and action paraphrase consistency.
  4. Counterfactual augmentation. In simulation, restore recorded high-level states and execute alternative language candidates. In real logs, use hindsight goal/subtask relabeling only when a state predicate verifies the relabeled outcome. AgentHER is suggestive but, as a March 2026 preprint, its language relabeling evidence is WebArena/ToolBench rather than robotics.18↗
  5. Policy extraction. Candidate generator proposes K actions; reject low feasibility/unsafe candidates; Q ranks survivors. AWR fine-tunes the generator only on high-advantage supported actions.
  6. Online allocation. Spend real episodes on state-action pairs with high epistemic uncertainty and plausible upside, not broad random exploration. Keep a fixed fraction of evaluation episodes completely untouched by training.
  7. Stability. Target networks, clipped double Q, calibrated ensembles, replay balance by task/failure mode, and a rollback checkpoint. Monitor RAGEN’s reward-variance and repetition-collapse signals.
Failure analysis

Likely ways this project fails

Language aliasing

Paraphrases fragment Q data; superficially similar instructions differ in grounding. Canonicalize slots and regularize paraphrase Q consistency while retaining raw text.

State aliasing / non-Markov history

Same image can hide grasp force, prior failed attempts, or object damage. Include belief, option result, recent actions, and uncertainty; test recurrent memory against snapshot state.

Extrapolation error

Q overvalues unseen, eloquent instructions. Use conservative Q, behavior likelihood, ensemble disagreement, and low-level feasibility; force an ask/observe fallback.

Feasibility–utility confusion

Easy skills dominate even when irrelevant. Train feasibility only on local completion and Q only on downstream task return; report both calibrations.

VLA drift

If the low-level policy changes, replay transition dynamics become stale. Version every VLA checkpoint; condition critics on version or freeze VLA during high-level RL phases.

Reward hacking

Progress detectors reward moving objects without satisfying relations. Use terminal predicates, negative tests, intervention cost, and video audits of top-Q failures.

Duration bias

Fixed-step discount treats 2-second and 30-second options equally. Use semi-MDP γk and report wall-clock/robot-time efficiency.

Simulator exploitation

Branch-generated policy learns visual or dynamics artifacts. Domain randomize, keep real-only validation, and down-weight synthetic transitions using uncertainty.

Critic self-confirmation

A VLM actor and VLM critic share blind spots. Use separate seeds/architectures, grounded predicates, double critics, and human review of disputed high-impact actions.

Decisive evaluation

Experiments that can falsify the design

ExperimentComparisonPrimary metricDecision it resolves
E1 · Q vs V vs no criticQ candidate ranker; V + sampled lookahead; SayCan-style feasibility×LM; VLM policy aloneSuccess per 100 real episodes; calibration/Brier scoreWhether action-conditioned long-term value is necessary.
E2 · data reuseOffline IQL/AWR; online ArCHer; turn-level GRPO; terminal GRPOLearning curve by environment transitions, not optimizer stepsWhether off-policy reuse creates the claimed real-rollout advantage.
E3 · critic factorizationTask Q only; feasibility only; product/additive combination; joint monolithic scoreInvalid-action rate, task success, recovery successWhether separating “can” and “should” matters.
E4 · offline supportBC; IQL; IQL+conservative penalty; unconstrained Q actorOOD action rate and real successWhether conservative control prevents language extrapolation.
E5 · state/historysingle frame; frame stack; object belief; recurrent episodic memorySuccess on tasks with hidden state/repeated mistakesWhat must be in the high-level state.
E6 · sim branchingno synthetic; random alternatives; Q-uncertain alternatives; Agent-Q/MCTS alternativesreal transfer per synthetic transitionWhether branch data helps or induces sim bias.
E7 · language action formfixed skill IDs; structured language slots; free-form instructionheld-out compositional tasks + VLA execution validityWhether language adds generalization beyond a discrete option library.
E8 · consequence critiquescalar Q; NLAC-like critique; scalar+critiquerecovery quality and human-rated causal accuracyWhether language-valued criticism adds actionable signal.
E9 · frozen vs changing VLAfrozen; periodically fine-tuned; joint alternating updatescritic calibration under policy versionsHow damaging non-stationary low-level dynamics are.
E10 · generalizationtrain task compositions A/B/C; test unseen D/E/F sharing atomic skillszero/few-shot success and subtask-sequence edit distanceWhether RL learns strategy rather than memorized scripts.

Go/no-go criterion. If offline Q does not outperform behavior cloning under a fixed real-transition budget and its calibration does not predict failures, do not scale real-world RL. Improve state representation/reward/support before collecting more episodes.

Synthesis & gaps

What the literature still does not establish

Evidence gap

No included agentic-LLM RL paper demonstrates off-policy training of a visual high-level language planner over a real manipulation VLA under a deliberately small rollout budget. Most “embodied” evidence is textual ALFWorld/ScienceWorld; the closest physical system, SayCan, does not RL-train its high-level LM.

Value gap

Scalar critics are easy to rank and calibrate but hard in open language spaces; language critics are expressive but hard to verify. A hybrid is promising, not settled.

Identity gap

GiGPO/G2PO assume repeated or mergeable states. Visual manipulation needs a defensible equivalence relation over observations, latent object state, and history.

Benchmark gap

Leaderboards usually report success or score, not success per real transition, robot-minutes, irreversible failures, intervention load, or calibration—the metrics that should govern this project.

Cross-cutting disagreement. Recent group-policy work argues that a learned critic is avoidable and potentially unstable; ArCHer/NLAC argue that explicit temporal value is precisely what yields sample efficiency. Both can be true in cheap simulators. Under expensive physical data, the decisive axis is not GPU memory but how many independent environment transitions an update can reuse; this strongly favors off-policy critics unless their extrapolation error dominates.

Survey method

Coverage, decisions, and stopping rationale

Search process

Search date: 3 September 2026. Default related-work depth: 1. Queries combined “multi-turn/agentic/interactive/embodied” with “reinforcement learning, actor critic, Q learning, offline, off-policy, credit assignment, replay, tree search, world model, language action.” Anchors were resolved to arXiv/OpenReview/PMLR/ACL, official project pages, and official GitHub repositories. Backward/forward neighbors were semantically filtered rather than included by citation proximity.

Evidence labels

venue peer-reviewed venue or proceedings. preprint not treated as established. system implementation evidence. Numeric claims in this report are author-reported unless explicitly stated; no cross-paper leaderboard comparison is implied because models and environments differ.

Explicit inclusion decisions

  • Included: ArCHer, GiGPO, turn-credit work, RAGEN, G2PO, NLAC because each changes the temporal learning problem.
  • Included as foundations: CHAI/CALM/WebShop because they expose offline language RL, candidate action, and environment-MDP lessons.
  • Adjacent: SayCan, LATS, Agent Q, RWML because their factorization/search/world-model lessons transfer directly.
  • Demoted: Agent Lightning—useful system, credit algorithm insufficiently specific for this question.

Exclusions

Ordinary RLHF/RLVR, single-turn reasoning, prompt-only agents, pure harness papers, and low-level robot control were not enumerated. WebGPT is a historical interactive-RL anchor but uses largely trajectory/answer-level human preference optimization and lacks the multi-turn temporal design required here, so it is bibliographic context rather than a recommended method.19↗

Blind spots

Several 2026 works are extremely recent and only available as preprints. Full reproducibility was not independently rerun. Citation counts were deliberately not used because age dominates. Some commercial agent-training details are not public. Robotics transfer statements are labeled as this survey’s inference, not authors’ claims.

Stopping rationale

Search stopped after all requested families had at least one primary representative; repeated queries yielded variants of grouped policy optimization or generic RLVR rather than new off-policy high-level mechanisms; and the main design alternatives—Bellman critic, conservative offline policy, group-relative turn credit, search/model, and affordance factorization—were saturated. The remaining gap is empirical, not another citation branch.

Canonical linked bibliography

Primary sources

Zhou et al. “ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.” ICLR 2024. Paper · Code.
Ichter et al. “Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.” CoRL 2022 / PMLR 2023. Proceedings · Project.
ArCHer official repository. Implementation and experiment configuration. GitHub.
Hong et al. “Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space.” 2025 preprint. Paper.
Feng et al. “Group-in-Group Policy Optimization for LLM Agent Training.” NeurIPS 2025. Paper · Code.
Wang et al. “RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.” 2025 preprint. Paper · Code.
Zeng et al. “Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Credit Assignment.” 2025 preprint. Paper.
Wang et al. “Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning.” 2026 preprint. Paper.
Luo et al. “Agent Lightning: Train ANY AI Agents with Reinforcement Learning.” 2025 preprint. Paper · Code · Official overview.
Verma et al. “CHAI: A Chatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning.” NAACL 2022. Paper · Code.
Jaques et al. “Human-centric Dialog Training via Offline Reinforcement Learning.” 2020. Paper.
Putta et al. “Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.” 2024 preprint. Paper.
Yu et al. “Reinforcement World Model Learning for LLM-based Agents.” 2026 preprint. Paper.
Yao et al. “Keep CALM and Explore: Language Models for Action Generation in Text-based Games.” EMNLP 2020. Paper · Code.
Yao et al. “WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.” NeurIPS 2022. Paper · Code/data · Project.
Fei et al. “Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning.” 2025 preprint. Paper.
Zhou et al. “Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models.” ICML 2024. Paper · Code.
Ding. “AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling.” 2026 preprint. Paper.
Nakano et al. “WebGPT: Browser-assisted Question-answering with Human Feedback.” 2021. Paper · Official project.
Ji et al. “Tree Search for LLM Agent Reinforcement Learning.” 2025 preprint. Paper.
Wang & Ammanabrolu. “A Practitioner’s Guide to Multi-turn Agentic Reinforcement Learning.” 2025 preprint. Paper · Code.
Tang et al. “WorldCoder, a Model-Based LLM Agent.” 2024 preprint. Paper.
Pang et al. “Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts.” 2024 preprint. Paper.