This paper describes the ingestion and ontology-tagging layer that turns a validated extraction stream into a knowledge graph of 537,157 entities and 2,198,567 relationships drawn from 98,795 government documents, and describes a record-identity ladder that decides sameness from identifier columns, name columns, display names and type-scoped position rather than from name similarity.
Abstract
Extraction produces candidate entities and relationships; writing them into a graph is where identity is decided, and identity decisions are destructive in a way extraction errors are not. A wrong type can be corrected later, but two records merged under one identity cannot be separated once their properties have been combined, and the merge leaves no error behind. This paper describes the ingestion and ontology-tagging layer that turns a validated extraction stream into a knowledge graph of 537,157 entities and 2,198,567 relationships drawn from 98,795 government documents. We describe a record-identity ladder that decides sameness from identifier columns, name columns, display names and type-scoped position rather than from name similarity. The ladder governs de-duplication within parsed tables, while the graph write applies a coarser canonical-name key, so records sharing a canonical name merge automatically on exact equality. We argue rather than demonstrate that this is where the automation line belongs: no identity benchmark is reported, and the over-merges the key permits are undetectable by construction. That policy, under which entity resolution only ever flags candidates, followed an incident in which two surface forms of one name were merged, corrupting a correct record and deleting eight entities from an unrelated document. We then describe multi-class ontology tagging and an evidence asymmetry we did not anticipate: an entity name is an instance label rather than a type assertion, so matching name fragments against a class index invents classifications. Requiring anchored evidence cut role assignments on an enriched sample from 36 to 4, all confirmed correct. We quantify the graph's conformance debt, show secondary classifications compensating for a mis-parented primary class, and describe a curation queue grown to 48,403 pending proposals against 775 human decisions.
A production extraction layer that converts a live document stream into a validated knowledge graph aligned to a formal ontology, and improved search recall from roughly 70 to 95 percent with no false merges, and corrected seven classes of silent quality defect.
Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik· arXiv.org· 0 citations
: Entity Matching (EM) is a core challenge in data integration, requiring the identification of records that refer to the same real-world entity across heterogeneous sources. Practical experience shows that overall performance depends on end-to-end system design rather than isolated algorithms: candidate generation, threshold calibration, provenance tracking, consolidation policies, and iterative error analysis often determine effectiveness. Ontologies, knowledge graphs, and persistent identifiers provide semantic context and stable references, but introduce additional complexity in handling uncertainty and evolving representations. We present a Named Entity Management System (NEMS) that integrates entity lifecycle management with scalable matching in a unified workflow for knowledge graph creation. Instead of treating reconciliation as post-processing, NEMS embeds matching and validation during ingestion, combining attribute-level similarity with graph-structured and ontology-aware signals to guide merge decisions. By integrating canonical identifiers, provenance tracking, and configurable decision thresholds, NEMS enables conservative merging, incremental updates, and explainable outcomes. The architecture accommodates diverse matching paradigms while leveraging structural context, providing a robust foundation for scalable and consistent entity integration.
Andrea Leoni, Andrea Molinari, Simone Sandri· International Conference on...· 0 citations
This work targets a KG for Sophocles’ Antigone that supports two coupled uses: structured retrieval, through integrity and competency questions expressed in SPARQL over dramatic structure and interpretive annotations; and interactive exploration, through a lightweight read client that navigates lines across languages, shows scene context, and reports corpus statistics.
A two-stage LLM-assisted workflow for French maintenance regulations is presented: ontology engineering from a SEMLEG-based core ontology, followed by construction of an ontology-grounded French legal knowledge graph.
Génesis Montenegro, M. Billami, Catherine Faron et al.· arXiv.org· 0 citations
We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2.0 contains 38.4M triples over 1.6M canonical entities, together with 207.6K consolidated relations and 66K consolidated classes. Unlike prior LLM-derived knowledge bases that largely identify entities by surface strings, GPTKB 2.0 performs context-guided disambiguation during recursive KB construction, separating homonyms and merging synonymous mentions as facts are elicited. The demo makes this process inspectable: users can browse entities, follow links across the KB, and audit the provenance of individual facts, including surface forms, candidate matches, source triples, and disambiguation decisions. The interface further supports structured SPARQL queries, natural-language questions translated to SPARQL, and entity linking from user-provided text to canonical GPTKB 2.0 entries. GPTKB 2.0 is available at https://gptkb.org/, with the full KB downloadable for offline use.
Yujia Hu, Tuan-Phong Nguyen, S. Razniewski· 0 citations
A seven-stage graph-grounded pipeline that converts domain documents into a complete, auditable Web Ontology Language (OWL) Terminological Box (TBox) without any unconstrained generation step is presented, demonstrating that the pipeline produces stable, reusable domain representations from large document corpora.
Maruf Ahmed Mridul, A. Talukder, O. Seneviratne· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.