Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring substantial costs and risking catastrophic forgetting of previously learned tasks. To address...
Zai-Jing Li, Rui Shao, Bing Hu et al.· 0 citations
While radiance field representations have achieved remarkable success in photo-realistic novel view synthesis with densely captured images, their performance sharply degrades under sparse input conditions, resulting in floater artifacts and missing regions caused by its inherent shape-radiance ambiguities and lack of i...
Hao-Yu Zhang, Shuai-Feng Zhi, Zhen-Hua Du et al.· IEEE Transactions on Visuali...· 0 citations
This work introduces Coherent4D, a large-scale egocentric dataset for continuous 4D interaction forecasting, comprising approximately 233K samples across three domains, and proposes HIGFlow, a Hand Interaction Guided Residual Flow framework that models forecasting as a cascaded where-to-how process.
Qiao-Hui Chu, Hao-Yu Zhang, Meng Liu et al.· 0 citations
Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for preventing harmful content. Existing adversarial probes evaluate these methods by testing whether erased concepts can still be recovered. However, existing erasure and probe methods remain largely t...
Qian-Long Xiang, Miao Zhang, Kun Wang et al.· 0 citations
UAV-MAS is proposed, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error...
Hao-Yu Zhang, Shuoxun Zhang, Peng Ye et al.· 0 citations
By systematically revealing the strengths and limitations of existing models in aerial-ground collaborative reasoning, AeroGround provides a foundation for developing more capable aerial-ground collaborative embodied intelligence systems.
Shenghong Yi, Lin Zhang, Muzian Li et al.· 0 citations
This work proposes a novel information overloading method that is equipped with both extensive text and multi-dimensional image attacks, underscoring the need for stronger defenses against complex multimodal jailbreak inputs.
Haoyu Zhang, Yangyang Guo, Mohan S. Kankanhalli· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.