Action Tokenizer
- Converts continuous Language Table actions into discrete tokens.
- Input action:
2D planar motion.
- Action chunk:
20 steps at 10 Hz, covering 2 seconds.
- Represents each action chunk using
8 discrete tokens.
- Uses Finite Scalar Quantization (FSQ) with levels
[8, 8, 4, 4].
- Shared action vocabulary:
1,024 codes.
- Normalizes both action dimensions to the fixed range
[-0.03, +0.03].
- Uses a Transformer encoder-decoder with causal registers and nested dropout.
- Trains with continuous-action regression plus a DCT-domain reconstruction loss.
- Uses the EMA weights from checkpoint step
20,000.
AR-VLA
- Base model:
Qwen3.5-0.8B.
- Input: one image and a language instruction.
- Output: autoregressive action tokens.
- Each predicted sequence contains
8 action tokens representing a 20-step continuous action chunk.
- The action tokenizer converts continuous training actions into token targets.
- At inference, the action tokenizer decodes predicted tokens back into continuous actions.
- Uses the
1,024 action codes as an added action vocabulary.
- Uses the non-thinking Qwen chat template.
- Fine-tunes the full VLM:
- Language model
- Vision tower
- Multimodal projector
- No backbone-freezing or embeddings-only phase.
Training Sampling
- Samples
4 different action chunks from each episode.
- Consecutive epochs use different sampled chunks.
- Every
4 epochs form one meta-epoch.
- Because the sampled chunks change between epochs, loss and token accuracy may change abruptly at epoch boundaries.
Task-Level Evaluation: Single-Instruction Execution
Execution Settings
The three VLAs use different action-chunk and execution horizons:
| Model |
Action-Chunk Horizon |
Execution Horizon |
| Pre-Trained VLA by Google |
1 |
1 |
| Fine-Tuned Pi0.5 |
10 |
5 |
| AR-VLA (Fine-Tuned Qwen3.5-0.8B) |
20 |
10 |
Each model was evaluated on 50 episodes per task class. Neither model produced invalid predictions.
| Model |
Training Point |
Block-to-Block |
Block-to-Absolute-Location |
Block-to-Block-Relative-Location |
Block-to-Relative-Location |
Separate |
qwen_oat_0917_2152 |
Epoch 4/4 |
12% |
14% |
32% |
26% |
56% |
qwen_oat_0918_0908 |
Epoch 10/32 |
10% |
28% |
20% |
44% |
82% |
Offline Train/Validation Evaluation
Checkpoint 7040 from epoch 10/32 was evaluated on 100 samples from each checkpoint-verified training and validation split. The action metrics use valid free-generation outputs only.
Teacher-forced code accuracy is reported for each of the eight action-token positions. The summary calls its generated-action MSE normalized_mse.
| Split |
Pos. 1 |
Pos. 2 |
Pos. 3 |
Pos. 4 |
Pos. 5 |
Pos. 6 |
Pos. 7 |
Pos. 8 |
Normalized MSE |
MSE / Zero |
Normalized DCT L1 |
| Train |
23% |
64% |
68% |
74% |
75% |
73% |
74% |
75% |
0.1221 |
0.8249 |
0.2145 |
| Validation |
0% |
2% |
3% |
6% |
3% |
4% |
3% |
8% |
0.1800 |
1.1571 |
0.2683 |
- The tokenizer’s current reconstruction is not perfect, but it is acceptable.
- There is a severe overfitting issue, as shown by the train-versus-validation token accuracy.
- Sometimes, completely different token sequences lead to similar trajectories. This could be good, or it could indicate redundancy that makes it difficult for the model to learn the relationship between semantically similar tokens.
- The hardest token to learn is at position 1.
- Performance increases on some instructions while dropping on others. Can we draw any conclusions from this?