AR-VLA on Language Table

Updated

Debugging AR-VLA Training on Language Table

Training Codebase steps x batch_size Block-to-Block Block-to-Absolute-Location Block-to-Block-Relative-Location Block-to-Relative-Location Separate Overall
OpenPI (Pi0.5, exe10) 35k x 128 90% 80% 68% 74% 98% 82%
Loong 7k x 384 80% 90% 78% 82% 100% 86%
Loong 14k x 384 74% 72% 74% 74% 100% 78.8%
Llamafactory 11k x 256 98% 86% 88% 90% 100% 92.4%
Llamafactory 19k x 256 98% 86% 82% 84% 100% 90%

The Loong checkpoint has proprioceptive information (end-effector state) in its input, while Llamafactory doesn’t, which probably makes the difference.

For both the Llamafactory model and Loong model, the lowest-validation checkpoint is better than the overfitting checkpoint.

Language Table AR-VLA

Action Tokenizer

  • Converts continuous Language Table actions into discrete tokens.
  • Input action: 2D planar motion.
  • Action chunk: 20 steps at 10 Hz, covering 2 seconds.
  • Represents each action chunk using 8 discrete tokens.
  • Uses Finite Scalar Quantization (FSQ) with levels [8, 8, 4, 4].
  • Shared action vocabulary: 1,024 codes.
  • Normalizes both action dimensions to the fixed range [-0.03, +0.03].
  • Uses a Transformer encoder-decoder with causal registers and nested dropout.
  • Trains with continuous-action regression plus a DCT-domain reconstruction loss.
  • Uses the EMA weights from checkpoint step 20,000.

AR-VLA

  • Base model: Qwen3.5-0.8B.
  • Input: one image and a language instruction.
  • Output: autoregressive action tokens.
  • Each predicted sequence contains 8 action tokens representing a 20-step continuous action chunk.
  • The action tokenizer converts continuous training actions into token targets.
  • At inference, the action tokenizer decodes predicted tokens back into continuous actions.
  • Uses the 1,024 action codes as an added action vocabulary.
  • Uses the non-thinking Qwen chat template.
  • Fine-tunes the full VLM:
    • Language model
    • Vision tower
    • Multimodal projector
  • No backbone-freezing or embeddings-only phase.

Training Sampling

  • Samples 4 different action chunks from each episode.
  • Consecutive epochs use different sampled chunks.
  • Every 4 epochs form one meta-epoch.
  • Because the sampled chunks change between epochs, loss and token accuracy may change abruptly at epoch boundaries.

Task-Level Evaluation: Single-Instruction Execution

Execution Settings

The three VLAs use different action-chunk and execution horizons:

Model Action-Chunk Horizon Execution Horizon
Pre-Trained VLA by Google 1 1
Fine-Tuned Pi0.5 10 5
AR-VLA (Fine-Tuned Qwen3.5-0.8B) 20 10

Each model was evaluated on 50 episodes per task class. Neither model produced invalid predictions.

Model Training Point Block-to-Block Block-to-Absolute-Location Block-to-Block-Relative-Location Block-to-Relative-Location Separate
qwen_oat_0917_2152 Epoch 4/4 12% 14% 32% 26% 56%
qwen_oat_0918_0908 Epoch 10/32 10% 28% 20% 44% 82%

Offline Train/Validation Evaluation

Checkpoint 7040 from epoch 10/32 was evaluated on 100 samples from each checkpoint-verified training and validation split. The action metrics use valid free-generation outputs only.

Teacher-forced code accuracy is reported for each of the eight action-token positions. The summary calls its generated-action MSE normalized_mse.

Split Pos. 1 Pos. 2 Pos. 3 Pos. 4 Pos. 5 Pos. 6 Pos. 7 Pos. 8 Normalized MSE MSE / Zero Normalized DCT L1
Train 23% 64% 68% 74% 75% 73% 74% 75% 0.1221 0.8249 0.2145
Validation 0% 2% 3% 6% 3% 4% 3% 8% 0.1800 1.1571 0.2683

Comments and Summary

  1. The tokenizer’s current reconstruction is not perfect, but it is acceptable.
  2. There is a severe overfitting issue, as shown by the train-versus-validation token accuracy.
  3. Sometimes, completely different token sequences lead to similar trajectories. This could be good, or it could indicate redundancy that makes it difficult for the model to learn the relationship between semantically similar tokens.
  4. The hardest token to learn is at position 1.
  5. Performance increases on some instructions while dropping on others. Can we draw any conclusions from this?