Single-Instruction Evaluation of Fine-Tuned Pi0.5
Evaluation Results
Eight models were evaluated across five single-instruction Language Table tasks, with 50 episodes per task (250 episodes per model).
| Model |
Abs. loc. |
Block-to-block |
Rel. loc. |
Block rel. loc. |
Separate |
Overall |
| bc_resnet_sim_checkpoint_955000 |
78% |
90% |
58% |
56% |
92% |
74.8% |
| language_table_full_pytorch_bs128_25000_exe1 |
80% |
72% |
74% |
80% |
100% |
81.2% |
| language_table_full_pytorch_bs128_30000_exe1 |
82% |
90% |
82% |
74% |
100% |
85.6% |
| language_table_full_pytorch_bs128_35000_exe1 |
84% |
90% |
78% |
88% |
100% |
88.0% |
| language_table_full_pytorch_bs128_35000_exe10 |
80% |
90% |
74% |
68% |
98% |
82.0% |
| language_table_full_pytorch_bs128_35000_exe5 |
78% |
84% |
86% |
90% |
98% |
87.2% |
| language_table_full_pytorch_bs128_40000_exe1 |
78% |
80% |
66% |
72% |
100% |
79.2% |
| language_table_full_pytorch_bs128_50000_exe1 |
82% |
88% |
66% |
76% |
100% |
82.4% |
Best overall: language_table_full_pytorch_bs128_35000_exe1 with 220/250 successes (88.0%).
Analysis
For execution-length-1 checkpoint variants, the 35k checkpoint performs best.
At the same 35k checkpoint, execution lengths 1 and 5 are comparable. Execution length 10 drops significantly, primarily on blocktoblockrelativelocation.