Action-Language Alignment

Updated

Oct 2 Progress

Data Overview

Source Data Annotation SFT Data
ABC 130k IDM:
Ot, Ot+1 →VLM Lt(Ot, Ot+1)
IDM: Ot, Ot+1 → Lt, At
WM: Ot, At → Lt
L2A: Ot, Lt → At
WM-QA: Ot, At, Q(Lt) → A(Lt)
L2A-MC / IDM-MC: multiple choice
WM-Correct: Ot, At, Perturb(Lt) → Lt

SFT Recipe

Stage 1

  • Naive SFT (IDM + WM)
  • Issue: severely overfit to training questions

Stage 2 (Tuesday)

  • Improvement:
    • SFT tricks: freeze backbone warmup (implemented, but no clear evidence that it works)
    • More data: (10% -> 50%, annotated, not used yet)
    • General VLM training data as regularizer (works)
  • Issue: still overfit. Hard to understand error

Stage 3 (Today)

  • Improvement
    • Better human understanding -> change to OAT v3 tokenizer with tool-frame Cartesian action
    • More question classes: add L2A & WM-QA
    • Compositional question templates within each question class
  • Issue: code accuracy too low for OAT universal V3 tokenizer.
    • Hypothesis: vocab size too large (4096)
    • Now running tokenizer with vocab size of 64

Tokenizer Evaluation

[placeholder]

Eval Notes — September 29, 2026

Checkpoints

  1. Full SFT for 4 epoch (ckpt abbr: Full-4)
  2. Freeze backbone for 1 epoch, then full SFT for 1 epoch. (ckpt abbr: FB-1, FB-1+Full-1)

Analysis

Seen question classes: wm, idm

On seen tasks, all seems fine

On OOD tasks (Full-4 ckpt):

  • Bad example: wm example 1
  • Bad example: wm example 2 (the task is OOD even for human and the expert model)
  • Good example: wm example 3

Unseen question class: idm multiple choice

  • The accuracy is close to 20% (5 choices), so it is just random guessing

Unseen question class: reason for action

  1. Full-4 ckpt cannot do any question classes beyond the training ones
  2. Prompt trick: For the “reason for action” task, saying “[action tokens]” doesn’t work (i.e., make the model output valid action tokens), so I changed prompt to “Action: <left_eef>…</left_eef><right_eef>…</right_eef>” Answers with correct format
  • Full-4 ckpt: 1/10 (and the text overfits to the wm/idm task so it only describe the action)
  • FB-1 ckpt: 1/10 (and the text follows the instructed structure)
  • FB-1+Full-1 ckpt: 9/10 (but the text overfits to the wm/idm task so it only describe the action)

Other discussion – ambiguity, data annotations

  1. Direction Ambiguity
  • There is less ambiguity in “downward”, “upward”, “inward”, but more ambiguity in “left”, “right”. because they can mean the different things for top view and wrist view
  • My current plan is to leave it as it is now, because they are annotated by the same model, so ideally the wording should align.
  1. Blurry camera

  2. Some task are ood for the expert model, so annotation could be wrong

  3. Ambiguity in gripper state

  • If a gripper is half-open, It could be grasping something, or just being open.
  1. Do we need priprio state as input (currently no)

  2. What’s next

To-do

  • Change to Zheyuan’s v3 tokenizer – end-effector-frame delta end-effector pose actions
  • Train on 5x data (annotation finished)

Auxiliary SFT Data and Evaluation Monitoring

To reduce overfitting to robot data and action tokens, I include general-purpose language and vision-language data alongside ABC130K:

  • OpenHermes-2.5: 5,000 training samples and 500 validation samples for general language instruction following.
  • LLaVA-OneVision-Data: 100 training samples and 10 validation samples from each of 89 subsets, totaling 8,900 training and 890 validation samples for general vision-language tasks.

Evaluation runs before training and every 50 optimizer steps. W&B logs validation loss separately as eval/abc130k/loss, eval/openhermes/loss, and eval/onevision/loss. The auxiliary losses provide a signal of forgetting; they do not directly measure downstream task accuracy.

Broader Discussion

  • What’s next – new tasks, new training data
  • Is reusing existing VLM tokens a good design choice

How the Data Is Curated

Data Curation

  • Sample data: Sample ABC130K episodes and divide them into 1-second action chunks.
  • Annotate data: Use synchronized top, left-wrist, and right-wrist videos to generate fine-grained robot action descriptions.
  • Create training data: Convert each annotated chunk into one IDM or WM SFT item.
  • IDM objective: Predict the action description and action tokens from observations before and after the action.
  • WM objective: Predict the action description from the current observation and the provided action tokens.
Data Processing and Sampling Configuration
  • Dataset: ABC130K
  • Episode sampling:
    • 10% of train episodes, rounded down globally
    • 100% of val_custom episodes
    • Other splits excluded
    • Random seed: 42
  • Chunk sampling:
    • Train: up to 8 chunks per episode
    • Validation: up to 4 chunks per episode
    • Chunk length: 1.0 s
    • Start times: integer seconds
    • Sampling: uniform among feasible, non-overlapping chunks
    • Short episodes: use all available chunks
  • Annotation inputs:
    • 3 synchronized views: top, left wrist, right wrist
    • Composite video resolution: 544 × 448
    • Video sampling rate: 5 FPS
    • Task goal inserted into the prompt
  • Annotation model:
    • Qwen/Qwen3.8-Flash-Next-FP8
    • Thinking enabled
    • Maximum output: 16,384 tokens
    • 2 vLLM endpoints × 192 concurrent episode workers
    • Chunks remain sequential within each episode
    • Up to 8 concurrent FFmpeg preprocessing jobs
  • Annotation output:
    • One JSON file per episode
    • One action description per sampled chunk
    • Checkpoint saved after every chunk
  • SFT conversion:
    • One training item per annotated chunk
    • IDM: 6 images—three views at t, three views at t + 1 s; predicts the action description and action tokens
    • WM: 3 images at t plus the action tokens; predicts the action description
    • Images and action tokens are resolved later by the training loader

Prompts

Annotation Prompt
The videos show synchronized top camera, left wrist camera, and right wrist camera views of the same robotic manipulation task.
The left and right wrist cameras are mounted near the gripper on two robot arms and move with that arms.

The task goal is: {goal}

Your job is to describe the robot's physical action in detail.
- Describe fine-grained details like end-effector movement, direction.
- Focus on the robot itself rather than objects and enviroments.
- Do not describe static scene. Do not say something like "the first frame ..." or "the video ends with ..."
- Do not guess history, future, or the intent of the robot.
- Describe the robot's actions directly, without mentioning timestamps, frames, camera views, or phrases such as "the video" or "the clip".
IDM Training Prompt
The images show top camera, left wrist camera, and right wrist camera views of the same robotic manipulation task. The three views within each observation are synchronized.
Each wrist camera is mounted near its arm's gripper and moves with that arm.

The task goal is: {goal}

You are given observations immediately before and after one robot action chunk lasting exactly 1 second.

<before_observation>
Top camera: <image>
Left wrist camera: <image>
Right wrist camera: <image>
</before_observation>

<after_observation>
Top camera: <image>
Left wrist camera: <image>
Right wrist camera: <image>
</after_observation>

Your job is to infer the robot's physical action.
First describe the robot action in language, then output the action.

<output_format>
[action description]
Action: [action tokens]
</output_format>
WM Training Prompt
The images show top camera, left wrist camera, and right wrist camera views of the same robotic manipulation task. The three views within each observation are synchronized.
Each wrist camera is mounted near its arm's gripper and moves with that arm.

The task goal is: {goal}

You are given the current observation and the action chunk that will be executed. The action chunk lasts exactly 1 second.

Top camera: <image>
Left wrist camera: <image>
Right wrist camera: <image>

Action: <action_chunk_1>

Your job is to describe the robot's physical action in language.