Skip to content

Author

Zhen-Bin Wang

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection

Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost. To reduce this overhead without retraining, existing methods select compact token subsets that prioritize query relevance, visual coverage, or a fixed trade-off between t...

Zhen-Bin Wang, Lei Zhang, Li-Tuan Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models

Multimodal large language models (MLLMs) can answer knowledge-intensive visual questions by combining visual evidence from images with facts retrieved from external sources. However, MLLMs may overlook relevant evidence in both modalities, attending weakly to the textual sentences or visual regions needed for the corre...

Zhen-Bin Wang, Lei Zhang, Li-Tuan Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.