Act to See: Structured Articulated Representations through Robot Interaction

Act2See reconstructs a structured model of an unknown articulated object while a robot is interacting with it. Instead of treating perception and manipulation as separate stages, the system closes the loop between observing the object, choosing an informative action, estimating its joint structure, and updating an explicit URDF representation online.

Approach

A vision-language model planner reasons over the current RGB-D observation, the maintained URDF state, the discovered joint list, and the history of previous attempts. It decides whether to probe the object, execute an action, or skip an unhelpful action. The selected interaction is grounded into a grasp and constrained motion, and the resulting end-effector trajectory provides evidence for estimating the object’s joint structure.

The URDF is both the reconstruction output and structured memory for the planner. Each interaction can therefore improve the representation used to plan the next action.

Team and recognition

The project is a collaboration with Yifeng Liu, Yuzhen Chen, Bingyang Wang, Jingyi Lu, Han Li, Kaichen Zhou, Mengyu Wang, and Fangneng Zhan. It was accepted to the RSS 2026 FM4RoboPlan Workshop.

Chengqi (William) Li
Chengqi (William) Li

I build on-device multimodal AI products and real-time perception systems.