Lightweight scene proxies let creators control scene layout and motion while leaving room for imagination in appearance, lighting, and visual effects. However, a suitable proxy is not uniquely defined, making paired proxy-video data difficult to construct automatically at scale. We present Proxy2World, a controllable w...
Hong-Li Xu, Wei-Long Yan, An-Bang Wang et al.· 0 citations
Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model (VLM) pass. We present D$^...
Zi-Jian Ye, Chen Wei, Wei Huang et al.· 0 citations
ArtiMo, a novel agent-driven framework for text-guided articulated mesh animation, develops an agentic pipeline powered by Large Language and Vision-Language Models (LLMs/VLMs) to orchestrate motion generation without requiring model fine-tuning.
Chun-Yu Zou, Peng Dai, Yi-Hua Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.