DocLayout-MM-RAG: A Layout-Aware Annotation Framework for Grounded Question Answering over Documents
Question answering over visually structured documents remains difficult when evidence is distributed across prose, tables, figures, captions, visual layout, and document structure. We present DOCLAYOUT-MM-RAG, a layout-aware annotation framework for grounded question answering over documents. The framework links each question-answer instance to supporting layout elements, preserving element-level provenance for annotation, retrieval, citation, generation, and evaluation. We instantiate the framework on annual reports and release an initial curated corpus of 650 accepted grounded question-answer instances across 30 documents. The corpus captures evidential complexity, with 36.6% of instances requiring cross-page support and 40.8% requiring multimodal support. Exploratory analyses compare flat-text, structure-aware, and layout-derived multimodal retrieval representations, and show how element-level provenance enables retrieval-to-generation and oracle-evidence analysis. DOCLAYOUT-MM-RAG provides a concrete basis for studying provenance-preserving retrieval-augmented generation over visually structured documents.