Skip to content
Preprint

BZKO: An Ontology for the Card Index of German Post-War Compensation Records

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

A two-layer ontology for historical archival data that separates ontologically grounded domain semantics from interoperability-oriented extension constructs and lays the groundwork for future knowledge graph generation, ontology validation, and the incorporation of additional historical entities and uncertain temporal and spatial information.

Abstract

The Central Federal Card Index (Bundeszentralkartei) of Germany is a key archival resource documenting compensation claims submitted by victims of National Socialist persecution and their relatives, within the German Wiedergutmachung process. To enable semantically enriched representation, integration, and reuse of this historically significant collection, we present the BZK Ontology (BZKO). We propose a two-layer ontology for historical archival data that separates ontologically grounded domain semantics from interoperability-oriented extension constructs. The approach combines BFO-based realism with archival standards (RiC-O, PROV-O, PiCo), enabling provenance-preserving semantic integration, while maintaining logical rigor, modularity, and reuse across digital humanities infrastructures. The proposed approach establishes a reusable semantic foundation for the integration of Wiedergutmachung archival materials into digital humanities infrastructures and lays the groundwork for future knowledge graph generation, ontology validation, and the incorporation of additional historical entities and uncertain temporal and spatial information. The ontology is available on https://github.com/ISE-FIZKarlsruhe/bzko.

View source

Similar papers

Review Aug 2026

Semantic Intelligence Against CSAM: The PreventCSA@EU Ontology Framework for Classification and Investigation

This work presents the PreventCSA@EU ontology, a semantically grounded framework designed to support the identification, classification, annotation, and analysis of online Child Sexual Abuse and Child Sexual Exploitation Material (CSAM/CSEM). The growing circulation and dissemination of CSAM/CSEM across digital environments, combined with inconsistencies in legal definitions and classification practices across jurisdictions, highlights the need for semantically interoperable frameworks capable of supporting cross-organizational cooperation and automated processing. The proposed ontology is developed through a systematic review and comparative analysis of existing CSA/CSE-related, metadata oriented, and investigative ontologies and taxonomies, with its primary design aimed at addressing the operational needs and domain-specific requirements of national LEA Directorates. It introduces a hierarchical semantic model built around core entities such as Media Object, Content, Person, Depiction, and Investigative Report, while enabling structured alignment with INHOPE UCS labels, Dublin Core-DMCI Metadata Terms, and Schema.org. The proposed framework emphasizes ontology-driven interoperability for structured annotation and analysis of CSA/CSE-related data, supporting consistent classification, child identification, and investigative processes for offender prosecution. The design aims extend existing classification approaches with additional conceptual structures for database conceptualization, process modeling, and ontology-driven data management. By integrating established classification standards with a novel hierarchical ontology, the proposed framework enhances cross-system compatibility, with particular relevance to emerging EU-level data infrastructures, including the envisaged EU Center database under the proposed Child Sexual Abuse Regulation (CSAR).

Elias Tzortzakakis, E. Kokolaki, Evangelia Daskalaki et al. · 0 citations
Book Open access Sep 2026

ReSB²: Retrieving Similar Brazilian State Bills

Legislative knowledge evolves as an intricate hypertext in which documents are interconnected through complex, often implicit relationships. In this paper, we introduce ReSB2, a framework for retrieving and linking similar legislative bills that supports human–machine collaboration and helps reduce redundancy in the lawmaking process. The framework fine-tunes two ModernBERT-based language models on authentic legislative data, incorporating domain-specific formatting and procedural constraints derived from real workflows in a Brazilian state-level legislative assembly. To ensure transparency, ReSB2 integrates an explainability module based on Integrated Gradients, enabling analysts to inspect which textual elements most influence model decisions. Evaluated on a large corpus of official bills, the framework outperforms both general-purpose and domain-specific baselines in identifying semantically similar documents, achieving recall values of approximately 0.9. Human-centric evaluation with domain experts further demonstrates that ReSB2 serves as an effective human-centered augmentation tool, supporting the consistency and governance of legislative knowledge.

Lucas G. L. Costa, Átila Souza, Elves Rodrigues et al. · 0 citations
Book Open access Sep 2026

A Graded Model of Semantic Commitment for Curated Cultural Heritage Data Integration

This paper presents a semantic hypermedia framework for documenting and interpreting cultural heritage artifacts, with particular attention to decontextualized and refunctionalized architectural components, through curated integration of heterogeneous sources. Built on Linked Open Data principles, it addresses two key challenges: mediating semantic complexity for domain experts, and ensuring data quality when integrating external resources. Knowledge patterns map CIDOC-CRM to a project-specific model aligned with scholarly practice, while external datasets—such as the Getty vocabularies, OpenStreetMap, Zotero, and Iconclass—are incorporated through a graded model of semantic commitment, distinguishing authority alignment, partial reuse, and full ontology adoption. External information is materialized in a knowledge graph to ensure reproducibility and long-term accessibility. By framing external linking and semantic mediation as a hypermedia design problem, the approach shows how controlled integration can support navigation, enrichment, and scholarly reuse of cultural heritage data while preserving curatorial control. We contribute a practical model for dataset curation, semantic mediation, and graph-based knowledge representation in cultural heritage hypermedia systems, developed through the design and implementation of a case study.

Polina Voronova, Alessandro Adamou · 0 citations
Jul 2026

LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph

A two-stage LLM-assisted workflow for French maintenance regulations is presented: ontology engineering from a SEMLEG-based core ontology, followed by construction of an ontology-grounded French legal knowledge graph.

Génesis Montenegro, M. Billami, Catherine Faron et al. · 0 citations
Open access Sep 2026

TEI, Taxonomies, and Semantic Data Governance for Documented LLM-Assisted Literary Analytics

This article examines how TEI-based digital archives and human-curated annotation layers can support documented uses of large language models in literary analytics. It argues that structured textual infrastructures strengthen the conditions for evaluating and interpreting LLM outputs by making scholarly decisions explicit, inspectable, auditable, and contestable. Here, semantic data governance refers to provenance, documentation, controlled vocabularies, interpretive criteria, and critical evaluation. The case study is the LdoD Archive and its Virtual Edition “Philosophical Intertextuality,” built on TEI-encoded sources for Fernando Pessoa’s Livro do Desassossego, a fragmentary and posthumously assembled work. The virtual edition adds a human-authored semantic layer through taxonomies and tagging tools that connect fragments with philosophical categories and named thinkers. Building on this architecture, the article proposes a layered workflow in which TEI, taxonomies, authority records, annotation tables, machine-readable metadata, and validation procedures interact. A curated dataset derived from the virtual edition can support qualitative comparison between human annotations and model outputs, exploratory evaluation of omissions, unsupported associations, and interpretive divergences, and prompt design based on concise scholarly examples. The annotations function as situated scholarly claims, allowing interpretive plurality to coexist with stronger provenance, inspection, and critical assessment in transparent and reproducible conditions for future LLM-assisted literary analysis across different analytical contexts.

Diego Emanuel Giménez Celano · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.