High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little. This ambiguity makes absolute revisit scores sensitive to rendering stability, repetitive content, and failed motion. We introduce \emph{R2M-Bench} (\textbf{R}elative \textbf{R}evisit \textbf{M}emory Benchmark), a benchmark of observable revisit-selective consistency. For every detected return, R2M-Bench compares the revisit pair with two controls from the same rollout: a gap-matched non-revisit pair that measures generic temporal stability and a short-range pair that estimates short-horizon consistency. These comparisons produce \emph{MemoryGain} (MG), the revisit advantage over the temporal baseline, and the \emph{Normalized Memory Ratio} (NMR), which normalizes this advantage by the short-to-baseline dynamic range. R2M-Bench combines 100 reference scenes with three leave-and-return trajectories to form 300 instances and evaluates appearance fidelity, scene and object identity, local geometry, and persistent state. Across seven action-conditioned video world models, Overall NMR correlates with human consistency judgments at Spearman's $\rho=0.547$ (95\% CI $[0.45,0.63]$). Its within-model correlation magnitude with generated motion is $0.072$, compared with $0.207$ for raw revisit similarity, indicating that relative calibration substantially reduces the slow-motion shortcut. DreamX-World-Memo achieves the highest Overall NMR among the evaluated video models. Together, these results support same-rollout relative calibration as a practical way to distinguish revisit-specific consistency from generic temporal stability.
Qiwen Gu, Bingjie Gao, Rui Chen et al.· 0 citations
As Multimodal Large Language Models (MLLMs) expand the scope of retrieval-augmented generation, recommendation, and multimedia search, audio retrieval is expected to become a dependable retrieval component. Yet existing systems struggle to reconcile lexical precision, non-verbal acoustic evidence, and efficient, transparent retrieval. Current approaches mainly follow two paradigms. Cascaded pipelines transcribe audio with automatic speech recognition (ASR) and retrieve over text. They inherit the strengths of lexical matching, but amplify recognition errors and systematically discard non-verbal cues such as acoustic events and speaking style. Dense retrievers bypass transcription, yet they compress each clip into a single embedding, making relevance difficult to inspect and incurring non-trivial indexing and query-time overhead as collections grow. To address these limitations, we propose LSAR (Learned Sparse Audio Retrieval), the first learned sparse retrieval framework for audio that maps audio directly into a sparse lexical space compatible with inverted-index search. LSAR employs cross-modal sparse alignment and multiple complementary branches to model spoken content, acoustic context, and logic associations, producing an indexable representation with explicit term activations. At inference, retrieval proceeds directly from audio without an ASR transcription stage. Experiments on diverse benchmarks covering speech, captioning, and audio question answering show that LSAR effectively maps continuous audio representations into a sparse lexical space. Acting as a robust first-stage retriever, it enables high-recall candidate retrieval with low latency to significantly narrow the search space, while offering keyword-level interpretability. Overall, LSAR introduces a new retrieval paradigm for audio, establishing an interpretable, index-friendly primitive for real-time multimodal RAG and beyond.
Haoyu Li, Yuzhe Bai, Li Niu· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.