Large Vision-Language Models (LVLMs) face significant computational inefficiencies caused by the large number of visual tokens. Existing visual token pruning methods mainly focus on either retaining individually important tokens or selecting mutually diverse ones. In this work, we revisit visual token pruning from a co...
Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions ma...
Meng-Yang Zhao, Zhuo-Lin He, Hai-Yang Yu et al.· 0 citations
SAGE is an evidence-grounded multi-agent framework that reformulates Chinese ancient document understanding as evidence-grounded inference rather than direct answer generation, highlighting the importance of structured, evidence-grounded inference beyond model scaling.
Yu-Chuan Wu, Xuan Luo, Yinglian Zhu et al.· 0 citations
This paper proposes IterCAD, an iterative framework that reformulates orthographic-view-to-CAD generation as a progressive program repair process that consistently improves code executability and geometric fidelity over strong one-shot baselines.
Yu-Chuan Wu, Ke Niu, Hai-Yang Yu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.