Skip to content

Improving large-scale DLA datasets through semantic validation and relation-aware multimodal LLMs

· 1 citation · 59 references

TL;DR

A preliminary version of a framework that improves the most widespread DLA datasets quality by assessing and correcting layout coherence in scholarly documents and provides a more reliable ground truth with improved structural and semantic coherence for training and evaluating document segmentation and understanding models.

View source

Similar papers

Open access Jul 2026

HRE-LSC: A Hyper-Relational Data Enhancement Framework for Long Tail Distribution and Structural Consistency

The results indicate that the proposed framework effectively enhances the quality and structural consistency of generated hyper-relational data while mitigating the effects of long-tail distributions and pseudo-negative samples without requiring additional manual annotations.

Xinzhi Du, Yan Chen, Siqi Xu et al. · 0 citations
Preprint Aug 2026

Ontology-Driven Structural Regularization for Document-Level Relation Extraction

This work introduces an ontology-driven framework to quantify and enforce structural consistency in DocRE datasets and reveals substantial structural noise in DocRED distant and demonstrates that such inconsistencies propagate to model predictions.

Laura Menotti, Stefano Marchesin, Gianmaria Silvello · 0 citations
2026

Augmenting Datasets for Fine-Tuning Large Language Models Using Semantic Variations

This study explores a semantic variation methodology to augment training data by generating question-answer pairs with explicit control over semantic similarity, and shows that semantically controlled augmentation improves domain-specific knowledge acquisition while preserving consistency.

Alexander Chen, Caroline Tang, Jennifer Sleeman · 0 citations
#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounded on a semantic reference label defined in the narrative, requiring models to resolve target entities from context before aggregating evidence across multiple charts. To enable systematic evaluation, we construct DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, covering 2,074 examples across six task categories. Experiments on a wide range of proprietary and open-source MLLMs show a substantial gap to human performance: annotators achieve over 90% accuracy, while the best model reaches only 62.83%. Reasoning-enhanced models consistently show improved results, but performance degrades as reasoning complexity increases. Overall, DocHop provides a controlled testbed for challenging multi-hop document reasoning.

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 0 citations
Preprint Jul 2026

KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quality from Text

A novel, interpretable metric for intrinsic KG quality assessment that measures how closely an automatically extracted graph approximates an"ideal"graph capturing the key noun phrases, predicate relations, and basic linguistic phenomena such as negation expressed in the source text is proposed.

Nipun Misra, Vikranth Udandarao, Aanchal Gupta et al. · 0 citations

Automatic Domain Classification of Tabular Datasets Using Large Language Models

It is asserted that the present contribution consists of an interpretable domain palette, a constructed benchmark of diverse tabular datasets, and reproducible code and data to enable further research on domain discovery and domain-aware tooling for tabular data.

Elizaveta Gamper, Irina Deeva · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.