Skip to content
SIG//08912026.07.24 · 20:03 UTCARXIV

A vision-language-action model that runs on-device

Inference moves off the datacentre and onto an embedded module drawing under 30 watts.

The team distilled a 7B parameter policy down to something that fits in 8GB.

Trade-offs

Success rates drop roughly four points against the full model.

More from AI & Embodied Intelligence

A vision-language-action model that runs on-device — RoboSignal