Impact of Physical-State Inputs on Hybrid Language Model for Manipulation
Research explores how physical-state inputs influence a hybrid language model's performance.
This study investigates the effects of physical-state inputs on a hybrid language model designed for manipulation tasks. The model, which has 0.8 billion parameters, was trained under six conditions across three LIBERO-Spatial tasks, evaluated over three seeds and 540 rollouts.
Key Findings
- Conditioning recurrent decay gates on geometric increments achieved a success rate of 28.9%.
- Shuffling increments during training improved success to 36.7%, while no explicit geometry led to 24.4% success.
- A token adapter using the same increments scored 27.8%, with variations across seeds.
- Token-clock conditioning yielded a low success rate of 11.1%.
- Robustness tests showed a state-only relative-coordinate policy maintained 7/10 success under frame relabeling.
- Visual policies struggled, achieving a maximum of 3/20 success after a 5 cm object displacement.
The results indicate no consistent advantage from geometric alignment during training and highlight the challenges in achieving coordinate invariance and physical-layout generalization.