Skip to content
Preprint

R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

Aug 2026 · 0 citations · 42 references
Computer Science

TL;DR

R4DSG introduces a relative 4D scene graph memory for long egocentric video, built on recent RGB-only advances in promptable video segmentation, temporal propagation, and relative 3D lifting, and produces a retrieval-ready memory directly usable for long-horizon question answering.

Abstract

Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based memories rarely preserve persistent object identity or structured spatial change. Existing long-video QA methods mainly emphasize temporal grounding and clip retrieval, while prior 3D scene-graph methods typically assume stronger geometry than free-motion wearable RGB video provides, including point clouds, RGB-D input, posed views, sparse reconstruction, or reconstructed scenes. R4DSG introduces a relative 4D scene graph memory for long egocentric video. Instead of storing raw graph sequences, R4DSG converts video into compact queryable memory entries indexed by time, place, persistent objects, anchor-relative change, and local interaction context. The main idea is to separate stable anchors from dynamic objects, maintain persistent object identity across frames, and represent object state through anchor-relative transitions rather than a globally aligned world model. Built on recent RGB-only advances in promptable video segmentation, temporal propagation, and relative 3D lifting, the method produces a retrieval-ready memory directly usable for long-horizon question answering. Evaluation on a 255-question object-related subset from EgoLifeQA shows, under question-only retrieval, a 6.7-point overall gain over EgoRAG-Text and a 12.5-point gain on when questions, which highlights the value of temporally organized object memory. These results position relative 4D scene graphs as a practical memory substrate for wearable assistants, AR systems, and embodied multimedia agents. GitHub Page: https://dualtransparency.github.io/R4DSG/.

View source

Similar papers

Preprint Sep 2026

VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

Given a compact semantic scene graph, long-term indoor video relocalization estimates a map-frame trajectory after lighting and furniture changes. Visual methods rely on appearance and become unreliable under these changes; localizing one frame at a time from object classes and geometry instead leaves sparse, ambiguous...

Qian-Ru Li, Xu-Yang Chen, Xu-Qin Wang et al. · 0 citations
Preprint Sep 2026

Structured Spatio-Temporal Evidence Graphs for Open-Vocabulary Object Retrieval in Videos

Open-vocabulary object retrieval in videos requires answering free-form object queries under bounded query-time cost. Existing index-based systems typically store independent frame-level regions and retrieve them with vision-language similarity, which is effective for appearance queries but mismatched with predicates w...

Jing-Dan Wang, Chuanyou Li · 0 citations
Preprint Sep 2026

TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking

Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TRACKGRAPH, an online...

Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis et al. · 0 citations
#artificial intelligence Preprint Sep 2026

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested v...

Wang-Bo Yu, Kunhao Liu, Wen-Bo Hu et al. · 1 citation
Preprint Sep 2026

FRAME: Factored Retrieval via Attribute Readouts for Object-Centric Scene Memory

Language-guided robots need persistent scene memories to follow instructions, revisit objects, and resolve references to objects encountered over time. While much of language-guided scene-memory retrieval has emphasized spatial or relational references, many everyday object references specify objects by multiple persis...

Woosang Jeon, Sanghyeok Choi, Minwoo Kim et al. · 0 citations
Preprint Sep 2026

SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding

Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric tok...

Xiang-Qi Li, Li-Bo Huang, Jia-Rui Zhao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.