Skip to content

Author

Rui Tian

We have 3 of 16 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Representation Dynamics Reveal Semantic Saliency and Similarity for Visual Token Pruning in MLLMs

Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several recent approaches also exploit representation changes, but when and how these changes reflec...

Wei-Xuan Li, Zi-Kun Zhou, Xin-Yi Zhuang et al. · 0 citations
Preprint Sep 2026

MintAct: A Unified Visual Agent for Digital Environments

We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain speci...

Ming-Fei Gao, Rui Tian, Hai-Ming Gang et al. · 0 citations
Jul 2026

TinyMem: Condensing Multimodal Memory for Long-Form Video Action Detection.

This paper introduces TinyMem, a model built upon compact multimodal memory for long-form video action detection that outperforms a range of state-of-the-art models on AVA v2.2 while using 5 times fewer memory tokens than the baseline with dense visual memory embeddings.

Rui Tian, Qi Dai, Hang-Rui Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.