Skip to content

AILMIR: Agent and incremental learning-based multimodal information retrieval for complex multi-hop multimedia queries

Aug 2026 · International Journal of Multimedia Information Retrieval · Vol 15 · 0 citations · 58 references
Computer Science

TL;DR

The Agent and Incremental Learning-based Multimodal Information Retrieval (AILMIR) framework is introduced, shifting the paradigm toward a dynamic Plan-Execute-Reflect cognitive loop, and a Non-Parametric Case-Based Memory is proposed that sediments successful reasoning trajectories, enabling efficient Domain Incremental Learning without destructive gradient updates.

View source

Similar papers

Jul 2026

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

A Cognitive-structured Multimodal Agent that externalizes visual information into an Episodic Visual Memory and selectively reactivates relevant episodes during reasoning is proposed, enabling reinforcement learning to optimize abstraction and retrieval policies.

Feng Wang, Canmiao Fu, Zhipeng Huang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

WeAgent-Harness, a multimodal agentic harness that supports native text-vision interaction and runtime recovery, and WeAgent-MMSearch, an integrated system spanning data construction, agentic post-training, and multimodal rollout that outperform similarly sized open-source models and rival models with roughly ten times its parameter count.

Zongkai Liu, Hui Zhang, Li-Qiang Niu et al. · 0 citations
Jul 2026

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

This work proposes ViSAGE, a multimodal agentic memory framework that constructs self-correcting, entity-centric memories and applies bidirectional memory refinement to propagate delayed identity evidence, retroactively unifying historical records and improving future reasoning.

Xinkui Zhao, Enbo Chen, Yifan Zhang et al. · 0 citations
#reinforcement learning Book Open access Aug 2026

NaviRAG: Learning to Navigate Knowledge Graphs for Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) has become a fundamental paradigm for enhancing Large Language Models (LLMs) with external knowledge. However, while recent structure-augmented approaches organize documents into graphs to improve information access, their retrieval strategies remain largely static, relying on similarity ranking or static probability diffusion. We identify that this paradigm suffers from two inherent limitations in complex reasoning: popularity bias, where retrieval paths are trapped by high-degree distractors, and signal decay, where relevance signals attenuate over long reasoning chains. To overcome these challenges, we propose NaviRAG, a novel framework that reformulates retrieval as a reinforcement learning-driven dynamic navigation problem on schema-less knowledge graphs (KGs). Unlike passive diffusion, NaviRAG employs an agent that actively traverses the graph to act as a search-space pruning engine, identifying logical multi-hop reasoning paths. Technically, we introduce three key components: (1) Structure-Aware Query Expansion, which bridges the modality gap between unstructured queries and structured graph seeds for precise initialization; (2) Target-Driven Reward Shaping, which provides dense supervision based on semantic progress toward gold documents, effectively mitigating the sparse reward problem in large-scale graph traversal; and (3) a Multi-View Hybrid Reranking strategy that operates on the highly-pruned candidate subgraph, integrating policy confidence, semantic relevance, and global structural importance to ensure robust candidate selection. Extensive experiments on three multi-hop QA datasets and two single-hop QA datasets demonstrate that NaviRAG significantly outperforms baselines, achieving state-of-the-art performance in multi-hop QA while maintaining robustness in single-hop QA. Our code and data are available at https://github.com/CkingEW/NaviRAG.

Jinghong Lei, Wang Kun, Zhigang Chen et al. · 0 citations
Preprint Aug 2026

TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory

TRAM (TRajectory-derived Auxiliary Memory), a training-free method that augments standard decoding with an auxiliary memory pathway derived from the model's own reasoning trajectory, shows that TRAM improves performance over vanilla decoding on mathematical, scientific, and general visual reasoning tasks without additional training.

Kang Liu, Zijing Wang, Yongkang Liu et al. · 0 citations
Jul 2026

Mixture of Cognitive Experts in Large Vision-Language Models

An evidence-driven multimodal reasoning framework that utilizes a Bloom-inspired taxonomy as a hierarchical reasoning protocol and quantitatively analyzes the trace to make evidence usage and reasoning progression explicit is proposed.

Robert Wijaya, Ngai-Man Cheung · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.