Skip to content

Author

Zhi-Xiang Wei

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision

VIVAS is proposed, a framework built upon the unified token space paradigm, which introduces a dense-structural-semantic vision tokenizer, which expands the textual vocabulary into a unified vision-language vocabulary by incorporating a visual vocabulary.

Zhe-Han Kan, Yu-Bo Zhu, Xing-Hua Jiang et al. · 0 citations
Preprint Aug 2026

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

UVU effectively synergizes pixel-level visual perception with semantic-level visual understanding, internalizing visual reconstruction capabilities and unlocking the facilitative role of visual supervision in enhancing understanding in the pre-training stage.

Zhe-Han Kan, Xing-Hua Jiang, Yu-Bo Zhu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.