Skip to content
SIG//21742026.09.10 · 20:36 UTCARXIV

Impact of Physical-State Inputs on Hybrid Language Model for Manipulation

Research explores how physical-state inputs influence a hybrid language model's performance.

This study investigates the effects of physical-state inputs on a hybrid language model designed for manipulation tasks. The model, which has 0.8 billion parameters, was trained under six conditions across three LIBERO-Spatial tasks, evaluated over three seeds and 540 rollouts.

Key Findings

  • Conditioning recurrent decay gates on geometric increments achieved a success rate of 28.9%.
  • Shuffling increments during training improved success to 36.7%, while no explicit geometry led to 24.4% success.
  • A token adapter using the same increments scored 27.8%, with variations across seeds.
  • Token-clock conditioning yielded a low success rate of 11.1%.
  • Robustness tests showed a state-only relative-coordinate policy maintained 7/10 success under frame relabeling.
  • Visual policies struggled, achieving a maximum of 3/20 success after a 5 cm object displacement.

The results indicate no consistent advantage from geometric alignment during training and highlight the challenges in achieving coordinate invariance and physical-layout generalization.

More from Research & Papers

Impact of Physical-State Inputs on Hybrid Language Model for Manipulation — RoboSignal