The Document Question Answering (DocQA) task necessitates the synergistic interpretation of visual and textual information embedded within documents. Although Retrieval-Augmented Generation (RAG) has enhanced the capabilities of Large Vision-Language Models (LVLMs), existing approaches still encounter significant bottlenecks when processing large-scale documents: the inability to capture long-range contextual dependencies within non-textual modalities, the difficulty in facilitating interaction and mutual complementation between different modalities, and the inefficient integration of heterogeneous modal information. To address these challenges, we introduce MMHRAG, a novel Multi-Modal Hierarchical Retrieval-Augmented Generation framework that achieves cross-modal interaction on DocQA for the first time. First, we construct a cross-modal hierarchical retrieval tree via a bottom-up recursive clustering and summarization mechanism. A key innovation of our approach is the structural injection of visual information, where image semantics are integrated as high-level abstract summaries of textual segments, thereby bridging the semantic gap between modalities. Furthermore, we design a Summarizing Agent to resolve logical conflicts, information redundancy, and granularity discrepancies among retrieved cross-modal evidence. Extensive experiments across multiple multi-modal long-document benchmarks demonstrate that MMHRAG significantly outperforms state-of-the-art baselines, achieving superior accuracy and consistency in complex reasoning tasks.
Jiayuan Wang, Jie Lian, Fu Zhao et al.· Proceedings of the 32nd ACM...· 0 citations
Unsupervised person re-identification (USL-ReID) typically relies on clustering to generate pseudo-labels, but significant cross-view appearance variations often cause images of the same identity to be split into different clusters. Training on such noisy pseudo-labels severely degrades the learned representations. Therefore, learning robust view-invariant features is paramount. Data augmentation provides a direct way to enhance invariance, yet its trade-offs in USL-ReID remain under-explored: weak augmentations usually preserve identity semantics but lack diversity, whereas strong augmentations provide richer appearance diversity at the cost of partially corrupting identity-consistent semantic cues. To address this challenge, we propose Invariant Representation learning with Progressive Prototype Refinement (IRPP), a unified framework that learns invariant and discriminative features from noisy pseudo-labels. IRPP consists of three synergistic components. First, an Augmented Dual-Contrastive Learning (ADCL) module performs dataset-level prototype-guided invariant learning by contrasting weakly and strongly augmented views against cluster-derived prototypes. Second, an Alignment and Uniformity Learning (AUL) module regularizes the mini-batch-level weak–strong feature geometry, leading to more stable feature distributions under data augmentation. Third, a Progressive Prototype Refinement (PPR) mechanism progressively optimizes cluster centroids into cleaner prototypes, thereby mitigating the influence of noisy pseudo-labels and further strengthening invariant representation learning. This closed-loop design enables prototype-guided contrastive learning, weak–strong regularization, and prototype refinement to mutually reinforce each other. Extensive experiments on standard USL-ReID benchmarks demonstrate that IRPP achieves state-of-the-art performance with a simple and efficient training pipeline. Code is available at https://github.com/Trangle12/IRPP
Xuan Tan, Qixian Zhang, Ding Qi et al.· IEEE Transactions on Image P...· 0 citations
This work introduces MMHRAG, a novel Multi-Modal Hierarchical Retrieval-Augmented Generation framework that achieves cross-modal interaction on DocQA for the first time, and designs a Summarizing Agent to resolve logical conflicts, information redundancy, and granularity discrepancies among retrieved cross-modal evidence.
Jiayuan Wang, Jie Lian, Fu Zhao et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.