Skip to content

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

ACPruner: Visual Token Pruning as Biased Attention Coverage Maximization in LVLMs

Large Vision-Language Models (LVLMs) face significant computational inefficiencies caused by the large number of visual tokens. Existing visual token pruning methods mainly focus on either retaining individually important tokens or selecting mutually diverse ones. In this work, we revisit visual token pruning from a co...

Xu Li, Yuxuan Liang, Yi Zheng et al. · 0 citations
Preprint Sep 2026

Counterfactual Attention Policy Distillation for Temporal Video Grounding

Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of On-policy distilla...

Shao-Bo Ju, Hai-Yang Yu, Xue-Cheng Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection

Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions ma...

Meng-Yang Zhao, Zhuo-Lin He, Hai-Yang Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.