Skip to content

Author

Yi-Yi Zhou

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Visual Token Coding for Video Multimodal Large Language Models

In this paper, we propose a new token compression paradigm for video Multimodal Large Language Models (MLLMs), termed Visual Token Coding (VTC). Inspired by classical video coding principles, e.g., HEVC, VTC performs structured compression by predicting the I/P frames of a video and measuring their frame-wise residuals...

Chenxin Fang, Tao Chen, Junshuang You et al. · 0 citations
Preprint Sep 2026

Counterfactual Attention Policy Distillation for Temporal Video Grounding

Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of On-policy distilla...

Shao-Bo Ju, Hai-Yang Yu, Xue-Cheng Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.