Skip to content
Open access

TEI, Taxonomies, and Semantic Data Governance for Documented LLM-Assisted Literary Analytics

Sep 2026 · Analytics · 0 citations · 17 references

Abstract

This article examines how TEI-based digital archives and human-curated annotation layers can support documented uses of large language models in literary analytics. It argues that structured textual infrastructures strengthen the conditions for evaluating and interpreting LLM outputs by making scholarly decisions explicit, inspectable, auditable, and contestable. Here, semantic data governance refers to provenance, documentation, controlled vocabularies, interpretive criteria, and critical evaluation. The case study is the LdoD Archive and its Virtual Edition “Philosophical Intertextuality,” built on TEI-encoded sources for Fernando Pessoa’s Livro do Desassossego, a fragmentary and posthumously assembled work. The virtual edition adds a human-authored semantic layer through taxonomies and tagging tools that connect fragments with philosophical categories and named thinkers. Building on this architecture, the article proposes a layered workflow in which TEI, taxonomies, authority records, annotation tables, machine-readable metadata, and validation procedures interact. A curated dataset derived from the virtual edition can support qualitative comparison between human annotations and model outputs, exploratory evaluation of omissions, unsupported associations, and interpretive divergences, and prompt design based on concise scholarly examples. The annotations function as situated scholarly claims, allowing interpretive plurality to coexist with stronger provenance, inspection, and critical assessment in transparent and reproducible conditions for future LLM-assisted literary analysis across different analytical contexts.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.