Skip to content
Book

DocLayout-MM-RAG: A Layout-Aware Annotation Framework for Grounded Question Answering over Documents

Aug 2026 · Proceedings of the 2026 ACM Symposium on Document Engineering · 0 citations · 38 references

Abstract

Question answering over visually structured documents remains difficult when evidence is distributed across prose, tables, figures, captions, visual layout, and document structure. We present DOCLAYOUT-MM-RAG, a layout-aware annotation framework for grounded question answering over documents. The framework links each question-answer instance to supporting layout elements, preserving element-level provenance for annotation, retrieval, citation, generation, and evaluation. We instantiate the framework on annual reports and release an initial curated corpus of 650 accepted grounded question-answer instances across 30 documents. The corpus captures evidential complexity, with 36.6% of instances requiring cross-page support and 40.8% requiring multimodal support. Exploratory analyses compare flat-text, structure-aware, and layout-derived multimodal retrieval representations, and show how element-level provenance enables retrieval-to-generation and oracle-evidence analysis. DOCLAYOUT-MM-RAG provides a concrete basis for studying provenance-preserving retrieval-augmented generation over visually structured documents.

View source