Skip to content
Preprint

Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

Aug 2026 · 0 citations · 53 references
Computer Science

TL;DR

A schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard, and demonstrates generalisability of this framework across different generative AI models and transferability across different HTA organisations and languages.

Abstract

We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard. The schema, serving as an information model encoding domain knowledge, provides a unified, systematic, and consistent framework for extraction of hierarchical, nested information, with attributes of variable cardinality, and subsequent evaluation of the results. Information extraction from a document is performed in a single call to the model, in zero-shot mode. In the evaluation step, we introduce a path-based semantic matching algorithm to align the nested, variable-cardinality attributes in the extracted results with those in the gold standard. We use generative AI for semantic comparison of the extracted and gold standard values of an attribute, and introduce a rubric to classify the result of the comparison, according to domain-specific considerations, as an exact, semantic, useful, or non-match. We were able to extract 12 out of 14 attributes with an F1 score of $>$90\% from documents published by the health technology assessment organisation NICE, using the generative AI model Claude Opus 3. The time needed to extract the attributes from a document was $\sim$30 times lower than the time taken by a human domain expert. We further demonstrate generalisability of this framework across different generative AI models and transferability across different HTA organisations and languages.

View source

Similar papers

Jul 2026

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

LlamaExtract Agentic Plus ranks first on all three metrics, with accuracy comparable to coding agents at a fraction of the cost, and is the first to score value accuracy, record completeness at scale, grounding, and measured cost together.

Boyang Zhang, Adrian Lyjak, Elizabeth Stewart et al. · 1 citation
Conference Open access Aug 2026

AI-Driven Knowledge Externalisation: From Unstructured Documents to Structured Data Models

The findings suggest that AI-based structured extraction may redefine how organisations formalise expertise, shifting from document-centric storage toward schema-driven knowledge architectures.

Dilyan Georgiev, E. Gourova · 0 citations
Jul 2026

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

A production extraction layer that converts a live document stream into a validated knowledge graph aligned to a formal ontology, and improved search recall from roughly 70 to 95 percent with no false merges, and corrected seven classes of silent quality defect.

Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik · 0 citations
Open access Aug 2026

A Synergistic Knowledge Graph and LLM-Driven Framework for Intelligent Process Decision-Making Systems

The proposed knowledge graph construction method for the workpiece machining distortion domain is proposed, together with an intelligent decision-making framework driven by the collaboration of knowledge graphs and large language models, providing a feasible pathway for the structured organization, intelligent retrieval, and decision support of workpiece machining distortion knowledge.

Deguo Yao, Zhaoze Sun, Jie Gao et al. · 0 citations
Review Open access Jul 2026

LLM-driven materials knowledge extraction: multimodal parsing, ontology, and agentic systems

Extracting reliable knowledge from unstructured materials literature remains a central bottleneck for data-driven and AI-enabled materials discovery. Large language models (LLMs) are reshaping this task by integrating multimodal document parsing, ontology-guided semantic grounding, structured extraction, and agentic verification into increasingly unified workflows. This review analyzes these developments through a Perception–Cognition–Action lens. At the perception layer, we examine how scientific document parsers, multimodal LLMs, table and chart readers, and optical chemical-structure-recognition systems convert visually rich papers into computable evidence. At the cognition layer, we discuss how ontologies and knowledge graphs constrain LLM outputs, support entity alignment, and reduce semantic ambiguity. At the action layer, we compare schema-based extraction, schema-free discovery, and agentic extraction as a control–coverage–autonomy spectrum rather than a simple succession of tools. We further argue that reliability is the decisive criterion for large-scale deployment, and synthesize failure modes, layered defenses, and evaluation protocols that connect source grounding, ontology constraints, physical verification, and human-in-the-loop review. By distinguishing demonstrated extraction capabilities from more speculative AI-scientist and self-driving-laboratory visions, this review provides a comparative and risk-aware account of how LLM-driven systems can produce evidence-linked, physically meaningful, and reusable materials knowledge.

Shuai Yang, Yi-Meng Wang, Qiong Tu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.