Skip to content
Preprint

ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

ICM-Bench (Identity-Centric Memory Benchmark), which, to the best of the knowledge, is the first benchmark specifically designed to evaluate identity-centric reasoning over long video memories in multimodal agents, is introduced.

Abstract

Long-horizon multimodal agents should remember not only what happened but also who participated. This capability depends on linking recurring faces, voices, names, person-associated objects, events, and social relations to consistent identities over time. Existing long-video and multimodal-agent benchmarks measure broad memory question answering, but they do not isolate the ability to maintain recurring person identities and reason over their cross-time relations. We introduce ICM-Bench (Identity-Centric Memory Benchmark), which, to the best of our knowledge, is the first benchmark specifically designed to evaluate identity-centric reasoning over long video memories in multimodal agents. The benchmark contains 839 synthetic clips spanning 141 minutes and 1,217 open-ended questions about six recurring adults in a one-year life album. A theme-configurable pipeline generates the video collection and associates each question with its target identities and traceable supporting evidence. We compare direct caption-memory baselines, memory-augmented agents, and graph-retrieval systems. Gemini 3.1 Pro achieves the highest overall accuracy of 74.0%, yet its score falls to 60.3% on questions that require long-term identity profiles. The results show that current systems recover many event-level memories but remain less reliable when evidence must be accumulated around a stable person.

View source

Similar papers

Review Aug 2026

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

The Identity-conditioned Queries task is introduced, in which models are required to jointly associate and interpret an input video and a reference image of a person, and leverage this conditioning to address identity grounding, behavior understanding, and temporal reasoning, among other challenges.

Shibo Gao, Chongxiao Wang, Chenglong Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories

Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text...

Guang-Zhi Xiong, Xin-Yuan Zhang, Xiao Yang et al. · 0 citations
Preprint Sep 2026

What Does It Mean to Forget a Person? Individual-Level Unlearning in Vision-Language Models

Erasing individual identities from Vision-Language Models (VLMs) is uniquely challenging because personal data is entangled across modalities rather than stored as isolated attributes. However, existing multimodal unlearning benchmarks primarily evaluate attribute-centric forgetting, overlooking the more critical objec...

Xiong-Tao Sun, Hui Li, Tian-Tong Wu et al. · 0 citations
#natural language process... Preprint Sep 2026

To Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives

As AI systems evolve into personalized digital companions, a central capability is reasoning over a user's long-term personal history: not merely storing past events, but tracking longitudinal experiences and evolving preferences. Progress here is bottlenecked by evaluation, existing long-term memory benchmarks are lar...

Wen-Qi Zhou, Zhuo-Rui Yu, Kai-Ao Wen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations

Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying part...

Wen-Xu Jia, Xi-Ze Cheng, Zi-Han Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across...

Hui-Hui Ren, Lei Fan, Henry Pao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.