Eval Notes — September 29, 2026
Checkpoints
- Full SFT for 4 epoch (ckpt abbr: Full-4)
- Freeze backbone for 1 epoch, then full SFT for 1 epoch. (ckpt abbr: FB-1, FB-1+Full-1)
Analysis
Seen question classes: wm, idm
On seen tasks, all seems fine
On OOD tasks (Full-4 ckpt):
- Bad example: wm example 1
- Bad example: wm example 2 (the task is OOD even for human and the expert model)
- Good example: wm example 3
Unseen question class: idm multiple choice
- The accuracy is close to 20% (5 choices), so it is just random guessing
Unseen question class: reason for action
- Full-4 ckpt cannot do any question classes beyond the training ones
- Prompt trick: For the “reason for action” task, saying “[action tokens]” doesn’t work (i.e., make the model output valid action tokens), so I changed prompt to “Action: <left_eef>…</left_eef><right_eef>…</right_eef>”
Answers with correct format
- Full-4 ckpt: 1/10 (and the text overfits to the wm/idm task so it only describe the action)
- FB-1 ckpt: 1/10 (and the text follows the instructed structure)
- FB-1+Full-1 ckpt: 9/10 (but the text overfits to the wm/idm task so it only describe the action)
Other discussion – ambiguity, data annotations
- Direction Ambiguity
- There is less ambiguity in “downward”, “upward”, “inward”, but more ambiguity in “left”, “right”. because they can mean the different things for top view and wrist view
- My current plan is to leave it as it is now, because they are annotated by the same model, so ideally the wording should align.
-
Blurry camera
-
Some task are ood for the expert model, so annotation could be wrong
-
Ambiguity in gripper state
- If a gripper is half-open, It could be grasping something, or just being open.
-
Do we need priprio state as input (currently no)
-
What’s next
To-do
- Change to Zheyuan’s v3 tokenizer – end-effector-frame delta end-effector pose actions
- Train on 5x data (annotation finished)
Auxiliary SFT Data and Evaluation Monitoring
To reduce overfitting to robot data and action tokens, I include general-purpose language and vision-language data alongside ABC130K:
- OpenHermes-2.5: 5,000 training samples and 500 validation samples for general language instruction following.
- LLaVA-OneVision-Data: 100 training samples and 10 validation samples from each of 89 subsets, totaling 8,900 training and 890 validation samples for general vision-language tasks.
Evaluation runs before training and every 50 optimizer steps. W&B logs validation loss separately as eval/abc130k/loss, eval/openhermes/loss, and eval/onevision/loss. The auxiliary losses provide a signal of forgetting; they do not directly measure downstream task accuracy.
Broader Discussion
- What’s next – new tasks, new training data
- Is reusing existing VLM tokens a good design choice