Chat with Aviral (09/07): From Internalized Knowledge to Robot Reasoning
1. Aviral’s LLM Experiment
Modern LLM agents solve tasks by searching repositories and reading files through tool calls. This raises a question: can a model internalize such external knowledge—not merely memorize it, but understand it well enough to act more effectively at test time?
Aviral explored this in two stages:
-
Internalize repository knowledge. The model explored a codebase, generated questions about it, found the answers, and was fine-tuned on those question–answer pairs. Although this should have taught the model about the repository, agentic performance did not improve. Knowledge acquisition alone was therefore insufficient.
-
Learn to use the knowledge. The agent first collected multi-turn interaction rollouts in the environment. Another model then compressed parts of those interactions into reasoning traces framed as recalling relevant internal knowledge before acting. Fine-tuning on both the repository knowledge and these action-oriented reasoning traces improved performance.
The reasoning data required neither human demonstrations nor a stronger expert model: it was generated automatically by summarizing the agent’s own interactions.
2. Main Takeaway
Possessing useful knowledge and knowing how to use it are distinct capabilities. Auxiliary training may inject knowledge without improving the target task unless the model is also trained to retrieve and apply that knowledge when choosing actions.
The central challenge is therefore not only what knowledge to internalize, but also how to create training data that teaches the model to use it.
3. Robotics Analogy and Open Questions
The experiment suggests a possible direction for robotics: use robot trajectories themselves to generate reasoning-oriented training data, without relying on additional human or expert annotations.
However, robotics introduces several harder problems:
- Reasoning and control are both underdeveloped. A robot may lack both the ability to reason about an action and the ability to execute the correct action after reasoning. The mapping must therefore be learned in both directions.
- Long context remains difficult. Robot policies still struggle to use long observation histories and long reasoning traces. Could trajectories be compressed into useful memory or reasoning supervision?
- Language and action are not naturally aligned. LLM reasoning and actions often share the same token space. Robot actions are continuous and non-linguistic, so the model must also learn how linguistic reasoning corresponds to low-level control.
4. This Leads to the Broader Research Question
Can we automatically summarize (part of) a robot’s trajectories into supervision that somehow helps it learn important reasoning-related abilities, such as having memory and using reasoning to derive actions?
I call it the “free lunch” in current robotic datasets, similar to “self-supervised learning,” but for reasoning purposes.
I want the main gain to come from understanding the data themselves, not from distilling large expert models—for example, asking GPT how to understand something. We can use VLMs, but mainly for data generation, not for gaining knowledge.
From now on, these are LH’s own thoughts.
Key Research Question: What Is the “Free Lunch” in Robotic Trajectories?
What Robot Generalist Policies Need
- Ultimate goal: generalizable to unseen tasks (zero-shot); robust to variations of seen tasks; capable of fast few-shot adaptation; …
- Abilities:
- History conditioning
- In-context learning: learn from in-context demos, metadata conditioning, etc.
- Reasoning for high-level semantic understanding and planning
- Reasoning for low-level action
Note: Here, “reasoning” refers to the general human thinking process. It may take the form of chain-of-thought, as in LLMs, but may not be limited to that.
Why They Are Difficult for Robotics
-
One-sentence reason: A robotic policy involves multiple modalities: vision, language, action, and reward/value. In contrast, an LLM works only with language, which makes history, context, and reasoning feasible. As a result, there are no robotic foundation models trained on all of them.
-
Existing pre-training paradigms (disentangled):
- Vision → language: VLM training
- Vision → future action: normal robotic policy training, regardless of the base model (both a VLA and a WAM can be viewed as just a policy)
- Past vision (+ action) → future vision: world model training
- Past vision + future vision → action: inverse dynamics model, which is usually a co-training objective with a world model
- Vision (+ action) → reward/value: usually a co-training objective for world models
Given so many modalities and so many abilities to learn, the above are not enough. Even worse, in methods involving language, language may serve only as a label without involving in-depth reasoning. It is therefore hard to say that those models “understand language,” and “reasoning for action” remains missing in robotics.
- Existing efforts on reasoning for robotics:
- Check the related work in this paper: https://arxiv.org/abs/2608.26053
- Basically, they all require either (a) reasoning data generated by an expert or (b) structured annotations, which are not actually language-based free-form reasoning, but merely auxiliary objectives/representations.
Your Job – Survey and Summarize
- Works related to how people curate data and train models to align action and language spaces
- LLM works that use agentic trajectories to generate data; be sure to mention what ability this is for (apart from Aviral’s approach above, what else?)
- Works related to the idea of “summarizing (part of) a robot’s trajectories into supervision”
Post-Survey Thoughts
A Practical Project: Reasoning in Language + Action Token Space
- What abilities do we want?
- Reasoning before taking action
- Using actions naturally in reasoning, e.g., tree search over actions
- Imagining the outcome of an action (world modeling)
- Comparing two action sequences, or scoring an action (critic)
- In-context learning (how?)