Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric tok...
Xiang-Qi Li, Li-Bo Huang, Jia-Rui Zhao et al.· 0 citations
Large-scale video diffusion models (V-DMs) have achieved remarkable text-to-video generation quality, yet their massive computational complexity makes deployment costly. Post-Training Quantization (PTQ) offers an appealing route to accelerate inference without retraining, but existing diffusion PTQ methods remain fragi...
Wei-Lun Feng, Chuan-Guang Yang, Haotong Qin et al.· IEEE Transactions on Pattern...· 3 citations
P predictive Prompting (PrePrompt) is proposed, a novel CIL framework that circumvents correlation-based limitations by leveraging the inherent classification ability of pre-trained models to predict task-specific prompts and decomposes CIL into a two-stage prediction process: task-specific prompt prediction followed b...
Libo Huang, Xiang-Qi Li, Jia-Rui Zhao et al.· Proceedings of the 32nd ACM...· 4 citations
3D scene generation has rapidly evolved, significantly promoting the innovation of content creation. In this context, interaction techniques serve as a pivotal bridge connecting user intent with the generative models, thereby enabling precise control, real-time feedback and personalized customization of complex 3D scen...
Yuqi Li, Si-Wei Meng, Chuan-Guang Yang et al.· Proceedings of the Thirty-Fi...· 30 citations· ⚡1
DA-Nav is proposed, a Direction-Aware vision-language Navigation framework that reformulates navigation as a discrete spatial grounding problem on the egocentric 2D image plane, outperforming existing State-of-The-Art (SoTA) methods while maintaining a substantially stronger recovery capability.
Ye Yuan, Kehan Chen, Xinqiang Yu et al.· arXiv.org· 0 citations
Prompt-based learning has emerged as a promising paradigm for Class Incremental Learning (CIL), enabling pre-trained models to adapt efficiently to open-world scenarios. Existing methods often employ correlation-based strategies, where an image's feature serves as a query to retrieve the most relevant key prompts, with...
Libo Huang, Xiangqi Li, Jiarui Zhao et al.· Proceedings of the 32nd ACM...· 0 citations
RSIAT significantly outperforms state-of-the-art methods in both performance and parameter efficiency, achieving superior stability–plasticity trade-offs with minimal trainable parameters.
Jiarui Zhao, Libo Huang, Xiangqi Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.