Enhancing Failure Detection in Vision-Language-Action Policies for Robotics
A new framework improves failure detection in robotic manipulation tasks.
Vision-language-action (VLA) policies are promising for robotic manipulation but struggle with failure detection during long tasks. Traditional methods often detect failures post-action or rely on inadequate supervision, leading to inaccuracies.
This study introduces a data-efficient framework that utilizes unlabeled VLA action segments to create weak supervision signals, identifying abnormal behaviors. By employing active learning, the approach selectively annotates uncertain trajectories, enhancing both timestamp-level and trajectory-level failure detection.