<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Vision-Language Models | Chengqi Li</title><link>http://iamwilliamli.github.io/tag/vision-language-models/</link><atom:link href="http://iamwilliamli.github.io/tag/vision-language-models/index.xml" rel="self" type="application/rss+xml"/><description>Vision-Language Models</description><generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Wed, 01 Jul 2026 00:00:00 +0000</lastBuildDate><image><url>http://iamwilliamli.github.io/media/icon_hu57a9add8d6f814cad03555fc46bfb714_1053646_512x512_fill_lanczos_center_3.png</url><title>Vision-Language Models</title><link>http://iamwilliamli.github.io/tag/vision-language-models/</link></image><item><title>Act to See: Structured Articulated Representations through Robot Interaction</title><link>http://iamwilliamli.github.io/projects/act2see/</link><pubDate>Wed, 01 Jul 2026 00:00:00 +0000</pubDate><guid>http://iamwilliamli.github.io/projects/act2see/</guid><description>&lt;p>Act2See reconstructs a structured model of an unknown articulated object while a robot is interacting with it. Instead of treating perception and manipulation as separate stages, the system closes the loop between observing the object, choosing an informative action, estimating its joint structure, and updating an explicit URDF representation online.&lt;/p>
&lt;h2 id="approach">Approach&lt;/h2>
&lt;p>A vision-language model planner reasons over the current RGB-D observation, the maintained URDF state, the discovered joint list, and the history of previous attempts. It decides whether to probe the object, execute an action, or skip an unhelpful action. The selected interaction is grounded into a grasp and constrained motion, and the resulting end-effector trajectory provides evidence for estimating the object&amp;rsquo;s joint structure.&lt;/p>
&lt;p>The URDF is both the reconstruction output and structured memory for the planner. Each interaction can therefore improve the representation used to plan the next action.&lt;/p>
&lt;h2 id="team-and-recognition">Team and recognition&lt;/h2>
&lt;p>The project is a collaboration with Yifeng Liu, Yuzhen Chen, Bingyang Wang, Jingyi Lu, Han Li, Kaichen Zhou, Mengyu Wang, and Fangneng Zhan. It was accepted to the RSS 2026 FM4RoboPlan Workshop.&lt;/p></description></item></channel></rss>