Research Thoughts

Updated

Roadmap for Harmonic Reasoning

Stage 1 Stage 2 Stage 3
Focus Alignment Sync reason Streaming reason / ICL
Train WM: ot, at → ℓ(ot+1)
IDM: ot, ot+1 → at
ot, ot+1 → ℓ(at), at
ot−k, at−k, …, ot → reason, at —
Eval WM, IDM
Multiple choice, etc.
Task success
In-distribution, zero-shot
—
Comparison Action-as-text VLA w/o reason —

The handwritten function symbol is transcribed as ℓ(·); its exact notation is slightly unclear.

Discussion Notes (09/11)

Data Structure for Streaming Sensing, Reasoning, and Action

Original version:

Obs 1 -> Reason 1 -> (Interrupt) Action 1 -> Reason 1 cont'd -> (Interrupt) Obs 2 -> Reason 2 ...

Simplified version (assuming the action and new observation arrive simultaneously):

Obs 1 -> Reason 1 -> (Interrupt) Action 1 -> Obs 2 -> Reason 2 ...

Here, action tokens may serve the following purposes:

  • In the output
  • In the input (action history)
  • (Optional) In in-context examples
  • (Optional) In reasoning, e.g., “If I take , …”

Main Challenge – How Does a VLM Understand and Use Action Tokens, Not Just as Output?

The desired approach to data annotation is the “free lunch” in existing large-scale robotic datasets—we use future observations as oracle information to generate supervision for action-language alignment.

Here are some objectives:

  • Imitation learning: ot → lt, at
  • World modeling (WM): ot, at → l(ot+1)
  • Inverse dynamics model (IDM): ot, ot+1 → lt, at
  • Action captioning: proprio + at → textual description for at
  • Critic / preference model: ot, at1, at2 → reasoning about which action is better
  • Action matching – multiple-choice version of the objectives above: choose which action corresponds to the desired outcome, with other actions in the dataset as negatives

Notes:

  • We simplify the action as a one-step action, but it actually represents an action chunk. The observation can also mean multiple images.
  • If we use an observation as output, we use a rich language description of it; if we use an action as output, we optionally add reasoning annotations before action prediction.
  • For critic training, we probably need counterfactual data, so we need simulation.

A Plan for the Pipeline

  • Stage 1: pre-train the VLM to understand and use action tokens, combining some objectives above.
  • Stage 2: mid-train the VLM on the streaming data structure

The following are post-discussion thoughts.

Evaluation (for Stage 2)

  • Zero-shot performance on unseen tasks
  • (If possible) in-context learning or continual learning

Potential Failure Modes (by GPT Astra)

  1. Future-informed annotations become hindsight rationalizations. For example, a teacher sees a successful future grasp and writes: “The object is securely grasped, so lift it.” At decision time, the student may only see ambiguous contact.
  2. The model ignores action inputs while still achieving good training loss.
  3. Natural language loses the action distinctions that matter for control.
  4. Embodiment and coordinate conventions make action meaning ambiguous.
  5. Continuation preserves an obsolete plan instead of incorporating new evidence. Suppose the ongoing reasoning assumes a successful grasp. A new image shows the object slipping, but the model continues toward “lift and transport.”
  6. Offline demonstrations support factual description better than counterfactual ranking.
  7. The project improves through auxiliary supervision, while streaming reasoning contributes little.

Solutions:

  • For (1) and (5), we need to generate some suboptimal/wrong reasoning in a decision-time way and correct it in later steps.
  • For (2) and (7), the minimal effect of our project is that all those actions in the input/reasoning don’t work, and we only have actions as the final output. But even if only this works, it is already interesting.
  • For (3), this is why a language interface isn’t enough, but this is not a bottleneck for us because we have the action-token prediction objective.
  • For (4), we want a unified action space in this project. The dataset/evaluation tasks can have different embodiments, but the action spaces need to be aligned.
  • For (6), we want the dataset to contain/use simulation to generate counterfactual data.

Thoughts for Embodied Reasoning Project

– Combining two topics: harmonic reasoning + “free lunch” in robotic datasets

Motivation

Frontier LLMs (GPT Astra, Claude Fable) can control robots by directly controlling end effectors or calling skills/primitives.

But the limitations are:

  • Too slow
  • Cannot really understand complex embodiments, such as bimanual systems or dexterous hands, without in-context demos. Even if they can, their performance should be suboptimal unless they are trained on action prediction.

Solutions:

  • Harmonic reasoning
  • Tokenized actions, aligning actions with language by training on mixed tokens

From First Principles, What Do We Require for Harmonic Reasoning?

Assumption: We need to output an action every dt seconds. The initial thinking should not take more than T seconds—for example, dt = 1.0, T = 10.0. We receive a new observation (short video or image) every dt seconds.

  • The ability to continue/reuse previous reasoning rather than starting from scratch
  • The ability to output an action based on truncated reasoning
  • The ability to take many observation images/videos as context in a streaming way

Design

KV Cache

  • Key question: How do we keep a sliding-window context with recent media while also having a long-term text-based summary?

  • Solution: Cache observations and evict them in FIFO order when the context length is exceeded (see StreamingLLM).

    [Recent turns with observations: cached] [Textual Memory M_t: fresh] [Action query] → response

  • Update rules:

    Persistent base: [... Oₜ₋₁ Oₜ]
    Temporary use:  [... Oₜ₋₁ Oₜ] [Mₜ] [query] [response]
    Next base:      [... Oₜ₋₁ Oₜ Oₜ₊₁]
    Next use:       [... Oₜ₋₁ Oₜ Oₜ₊₁] [Mₜ₊₁] [query] ...
  • If we want, we can even remove the explicit textual memory because historical information is implicitly contained in cached turns.

Unified or Hierarchical

  • Unified: Simple and closer to our “ultimate version” of an embodied foundation model, but less flexible and efficient for action prediction.
  • Hierarchical: Less elegant, but more practical in terms of low-level policy choice. We can even use a WAM as the low-level policy. However, the high-level model never really learns low-level actions, and the interface design is a problem.

Data Annotation

Open-Source Models for Data Annotation

Requirement: Can be served on 8 × H100 GPUs

  • Qwen3.8-Flash-Next
  • GLM-5.3-Flash

How to Curate Data from Open-Source Datasets

  • The annotator is given future videos (the outcome of the action)
  • The reasoning tries to imagine the future outcome and explain why that action should be taken
  • The reasoning describes low-level actions in rich language; during training, we replace that language with action tokens

Chat with Aviral (09/07): From Internalized Knowledge to Robot Reasoning

1. Aviral’s LLM Experiment

Modern LLM agents solve tasks by searching repositories and reading files through tool calls. This raises a question: can a model internalize such external knowledge—not merely memorize it, but understand it well enough to act more effectively at test time?

Aviral explored this in two stages:

  1. Internalize repository knowledge. The model explored a codebase, generated questions about it, found the answers, and was fine-tuned on those question–answer pairs. Although this should have taught the model about the repository, agentic performance did not improve. Knowledge acquisition alone was therefore insufficient.

  2. Learn to use the knowledge. The agent first collected multi-turn interaction rollouts in the environment. Another model then compressed parts of those interactions into reasoning traces framed as recalling relevant internal knowledge before acting. Fine-tuning on both the repository knowledge and these action-oriented reasoning traces improved performance.

The reasoning data required neither human demonstrations nor a stronger expert model: it was generated automatically by summarizing the agent’s own interactions.

2. Main Takeaway

Possessing useful knowledge and knowing how to use it are distinct capabilities. Auxiliary training may inject knowledge without improving the target task unless the model is also trained to retrieve and apply that knowledge when choosing actions.

The central challenge is therefore not only what knowledge to internalize, but also how to create training data that teaches the model to use it.

3. Robotics Analogy and Open Questions

The experiment suggests a possible direction for robotics: use robot trajectories themselves to generate reasoning-oriented training data, without relying on additional human or expert annotations.

However, robotics introduces several harder problems:

  • Reasoning and control are both underdeveloped. A robot may lack both the ability to reason about an action and the ability to execute the correct action after reasoning. The mapping must therefore be learned in both directions.
  • Long context remains difficult. Robot policies still struggle to use long observation histories and long reasoning traces. Could trajectories be compressed into useful memory or reasoning supervision?
  • Language and action are not naturally aligned. LLM reasoning and actions often share the same token space. Robot actions are continuous and non-linguistic, so the model must also learn how linguistic reasoning corresponds to low-level control.

4. This Leads to the Broader Research Question

Can we automatically summarize (part of) a robot’s trajectories into supervision that somehow helps it learn important reasoning-related abilities, such as having memory and using reasoning to derive actions?

I call it the “free lunch” in current robotic datasets, similar to “self-supervised learning,” but for reasoning purposes.

I want the main gain to come from understanding the data themselves, not from distilling large expert models—for example, asking GPT how to understand something. We can use VLMs, but mainly for data generation, not for gaining knowledge.

From now on, these are LH’s own thoughts.

Key Research Question: What Is the “Free Lunch” in Robotic Trajectories?

What Robot Generalist Policies Need

  • Ultimate goal: generalizable to unseen tasks (zero-shot); robust to variations of seen tasks; capable of fast few-shot adaptation; …
  • Abilities:
    • History conditioning
    • In-context learning: learn from in-context demos, metadata conditioning, etc.
    • Reasoning for high-level semantic understanding and planning
    • Reasoning for low-level action

Note: Here, “reasoning” refers to the general human thinking process. It may take the form of chain-of-thought, as in LLMs, but may not be limited to that.

Why They Are Difficult for Robotics

  • One-sentence reason: A robotic policy involves multiple modalities: vision, language, action, and reward/value. In contrast, an LLM works only with language, which makes history, context, and reasoning feasible. As a result, there are no robotic foundation models trained on all of them.

  • Existing pre-training paradigms (disentangled):

    • Vision → language: VLM training
    • Vision → future action: normal robotic policy training, regardless of the base model (both a VLA and a WAM can be viewed as just a policy)
    • Past vision (+ action) → future vision: world model training
    • Past vision + future vision → action: inverse dynamics model, which is usually a co-training objective with a world model
    • Vision (+ action) → reward/value: usually a co-training objective for world models

Given so many modalities and so many abilities to learn, the above are not enough. Even worse, in methods involving language, language may serve only as a label without involving in-depth reasoning. It is therefore hard to say that those models “understand language,” and “reasoning for action” remains missing in robotics.

  • Existing efforts on reasoning for robotics:
    • Check the related work in this paper: https://arxiv.org/abs/2608.26053
    • Basically, they all require either (a) reasoning data generated by an expert or (b) structured annotations, which are not actually language-based free-form reasoning, but merely auxiliary objectives/representations.

Your Job – Survey and Summarize

  • Works related to how people curate data and train models to align action and language spaces
  • LLM works that use agentic trajectories to generate data; be sure to mention what ability this is for (apart from Aviral’s approach above, what else?)
  • Works related to the idea of “summarizing (part of) a robot’s trajectories into supervision”

Post-Survey Thoughts

A Practical Project: Reasoning in Language + Action Token Space

  • What abilities do we want?
    • Reasoning before taking action
    • Using actions naturally in reasoning, e.g., tree search over actions
    • Imagining the outcome of an action (world modeling)
    • Comparing two action sequences, or scoring an action (critic)
    • In-context learning (how?)