Improving large-scale DLA datasets through semantic validation and relation-aware multimodal LLMs
A preliminary version of a framework that improves the most widespread DLA datasets quality by assessing and correcting layout coherence in scholarly documents and provides a more reliable ground truth with improved structural and semantic coherence for training and evaluating document segmentation and understanding models.