Topology-Informed Visual Prompting For Vision Language Action Policies

Haoyang Wu, Abhinav Kumar, Dmitry Berenson
Download Paper

In submission to ICRA, 2027

TIVP Teaser

Vision-language-action policies often struggle to manipulate objects around complex obstacles. We use simulated planning and topological reasoning to generate additional training trajectories, then visually prompt the policy with predicted robot waypoints during deployment. Our method improves performance across simulated and real-world tasks, increasing hardware success by 40% over the strongest baseline.