Skip to content

Author

Zitong Yu

We have 3 of 24 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models

Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindrance to efficient reasoning. This paper introduces a training-free $\mathbf{P}$ersistence-Aware $\mathbf{C}$ompression and $\mathbf{A}$ggregation (PCA) method designed to preserve high-fidelity raw visual information before the encoding stage. PCA can be built on arbitrary VLLMs and consists of two modules: 1) A Dynamic Downsampling (DD) module that adaptively removes redundant frames by analyzing frame-wise similarity. 2) A Persistence-Aware Motion Enhancement (PAME) module that enriches each selected keyframe by aggregating the temporal context of its neighbors, ensuring that essential information is preserved even after aggressive frame reduction. Our approach substantially reduces the computation of long-context modeling, while enhancing the performance of the baseline model. Extensive experiments demonstrate that PCA consistently outperforms existing state-of-the-art approaches in both efficiency and accuracy, achieving a speedup of 1.8$\times$ to 2.5$\times$ compared to the baseline VLLM. The code is open-sourced at https://github.com/Heisenberg10110/PCA.

Zihan Song, Shuo Ye, Bo Zhao et al. · 0 citations
Jul 2026

GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs

This work proposes GMoT, a Gated Motion-Aware Tokenization module that explicitly distills sparse kinematic evidence into a compact sequence prior to temporal modeling, and introduces Body-Region Grounding (BRG) Recall as an anatomical-grounding proxy conditioned on correct predictions.

Taorui Wang, Wei Xia, Hui Ma et al. · 0 citations
Jul 2026

AC-VLA: Robust Out-of-Distribution Action Execution via Compositional Learning

AC-VLA is introduced, a plug-and-play Action Compositional learning framework comprising two architecture-agnostic components that achieves a ~28% absolute improvement on compositional OOD tasks while maintaining near-perfect in-distribution performance.

Xiaojiang Peng, Kai Peng, Jie Lu et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.