Skip to content

Large model-assisted video summarization via global entity unification and robust importance scoring

Jul 2026 · The Visual Computer · Vol 42 · 0 citations · 53 references
Computer Science

TL;DR

A summarization pipeline around Global Entity Unification and Robust Importance Scoring is built, but unlike earlier efforts, each object is traced across frames and attached consistent identifiers to it, and Fragmented, isolated descriptions become a single, object-aware text corpus that unifies the storyline.

View source

Similar papers

Aug 2026

PHA-Net: Prototype-based hierarchical alignment network for text-video retrieval

A new prototype-based hierarchical alignment network (PHA-Net) to align individual/local/global level representations across modalities and introduces multiple modality-shared prototypes as the bridge to efficiently optimize text and video representations for cross-modal alignment.

Xiaolun Jing, Kezhao Yin, Xinxing Yang et al. · 0 citations
Jul 2026

Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

This work proposes Self-SiMS, a self-similarity-based Moment Proposal and Scoring that exploits intrinsic relationships within videos, enabling robust span generation and scoring and introduces a query-aware MLLM-based reasoning stage to further sharpen alignment between text and video.

Jihyun Lee, Cheol-Ho Cho, Woojin Jun et al. · 0 citations
Preprint Aug 2026

Persistent Object Narratives for Token-Efficient Video Language Models

Experimental results establish persistent object narratives as a compact, structured, and temporally organized visual interface for Video-LLMs as a favorable trade-off between accuracy and visual-token count compared with prior compact Video-LLM interfaces.

Jun-Zhe Chen, Siyuan Meng, Xiaojie Guo · 0 citations
Preprint Sep 2026

MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval?

Text-video retrieval requires representations that can distinguish videos with similar scenes, actions, and temporal patterns. Recent multimodal large language models have been adapted as embedding models, but they often represent each input using a single token from the final layer. This can compress diverse video-text cues into a single vector and limit fine-grained retrieval. To address this limitation, we propose MARS, a multi-layer and multi-slot embedding framework for text-video retrieval. MARS constructs multiple adaptive representation slots by combining hidden states from different decoder layers, compares corresponding text and video slots, and aggregates their similarities for retrieval. To better handle confusing candidates, we further introduce a hard-negative-aware slot specialization objective that encourages the slots to capture discriminative matching cues. Experiments on four text-video retrieval benchmarks show that MARS achieves state-of-the-art results in both direct similarity-based retrieval and reranking settings. Ablation studies and analyses demonstrate that multi-layer fusion, multiple slots, and hard-negative-aware slot specialization provide complementary gains. Code is available at https://github.com/sejong-rcv/MARS.

Uicheol Jung, Juyoung Hong, Geuntaek Lim et al. · 0 citations
Preprint Aug 2026

Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering

Long-video question answering requires identifying sparse yet critical evidence from videos containing thousands of frames under a constrained visual-token budget. Existing methods either select query-aware frames in a single pass or rely on timestamped text solely as retrieval guidance, leading to two key limitations. First, selected frames tend to cluster around local relevance peaks, and once the budget is exhausted, omitted evidence cannot be recovered. Second, textual and visual evidence remain weakly aligned. We propose GCR, a training-free framework that casts fixed-budget frame selection as a joint evidence curation problem. Ground converts timestamped text into temporal events, selects query-relevant real frame anchors, and renders each event text onto its temporally aligned frame. Cover supplements grounded events with direct visual anchors for complementary visual evidence and applies global maximal marginal relevance to preserve diverse context. Refine revisits omitted temporal regions and replaces the weakest revisable context frame with a real-frame medoid---but only when the medoid offers greater evidence value. GCR maintains a fixed number of chronologically ordered frames and requires no VLM training or architectural modification. Experiments on LongVideoBench and Video-MME, across three 7B backbones and frame budgets of 8, 32, and 64, demonstrate consistent improvements in long-video QA. With the 7B LLaVA-OV backbone and 32 frames, GCR achieves 64.25% and 62.15% on the two benchmarks, outperforming the strongest reproduced baselines by 2.54 and 1.93 percentage points, respectively.

Fan Wei, Siru Zhong, Runmin Dong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.