
Act to See
A closed-loop robotic system that interacts with an unknown articulated object and reconstructs an explicit URDF online, using the representation as both output and planner memory.
Read the projectSelected work
All projects
A closed-loop robotic system that interacts with an unknown articulated object and reconstructs an explicit URDF online, using the representation as both output and planner memory.
Read the project
An Open-Vocabulary 3D Affordance Grounding system that combines 3D segmentation, language-guided proposals, and VLM verification to locate functional parts in 3D scenes.
Read the project
Tested the accuracy–latency trade-off of image-pretrained vision transformers, reaching 78.1% AUC-ROC at 200 FPS on DoTA.
Read the studyBuilt LiDAR–camera visualization and annotation systems used across perception teams, then trained a point-cloud sequence model to support dense pedestrian tracking workflows.
Project detailsHow I work
I turn ambiguous needs into coherent native workflows, then carry them through privacy, reliability, distribution, and support.
I integrate speech, language, and vision models under real constraints: latency, memory, battery, offline use, and replaceable runtimes.
I build tools and models that connect sensor data to action — from production LiDAR pipelines to interactive 3D scene understanding.
Recent recognition
Get in touch