Skip to content
Preprint

Metadata Reconstruction from Values Alone: Recovering Column Semantics in Undocumented Warehouses

Aug 2026 · 1 citation · 32 references
Computer Science

TL;DR

Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean, and Rosetta addresses the problem they pose first: recovering what columns and values mean from the data itself.

Abstract

Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean. Production warehouses are the inverse: cryptic identifiers, partial or absent documentation. We address the problem they pose first: recovering what columns and values mean from the data itself. Rosetta places a language model inside a verification harness: a deterministic profiler extracts structural evidence (value fingerprints, a 26-pattern library, checksum verdicts), the model proposes semantics conditioned on that evidence, and every fact carries provenance and a confidence bounded by its evidence class. Against human documentation on 680 paired columns across eleven BIRD databases, identifiers destroyed, the harness delivers metadata that is 0.475 accurate on the 42% of columns it commits to, against 0.223 on 94% for the same model used directly. Restricted to the 283 columns where both arms speak, the harness writes no better prose than the model alone; the gain is selection: deterministic evidence governs whether the system speaks (coverage +0.257 [0.128, 0.378]), not how well. The deterministic layer is a competence detector, not a competence amplifier. A backbone swap bounds the claim: the prose finding reproduces, but prompt-requested abstention does not transfer; a code-enforced commit gate (predictions registered first; measured on a third backbone and held-out databases) makes no-evidence coverage 0.000 on every backbone. On a blind i2b2 clinical warehouse Rosetta decodes 95.5% of 134 real ICD-9 codes from values alone and abstains on all 44 NDC drug codes. The catalog supports calibrated abstention at query time: under full schema opacity a naive translator falls from 0.92 to 0.42 execution accuracy while our gate answers at 86% accuracy over 59% coverage. Negative results are reported plainly, including that our own authority ladder is not the mechanism behind the headline.

View source

Similar papers

Review Sep 2026

Specification-Driven Data Architecture Reconstruction: From Physical Code to Logical and Conceptual Specifications

Legacy database migrations often begin with incomplete or outdated documentation, leaving physical data definition language (DDL) as the principal evidence of data architecture. However, DDL does not fully encode conceptual intent, and model-generated completions can be plausible without being correct. This study propo...

Oleg Grynets, Olena Pochernina, V. Lyashkevych · 0 citations
Preprint Aug 2026

Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction

Doc2DB-Bench is introduced, a benchmark for Document-to-Database construction, containing 203 long-document instances across 42 schemas and seven domain groups, with 117 entity tables, 132 relationship tables, 7,341 rows, and 41,935 cells, which provides a testbed for reliable, auditable, and relationally faithful LLM-...

Zhuo-Wen Liang, Zhengxuan Zhang, Jia-Yang Wang et al. · 1 citation
Aug 2026

ATLAS: Adaptive Text-to-SQL with Lifecycle-Aware Self-Maintaining Context

Deploying Text-to-SQL in production is hampered by context-window limits on large schemas, metadata that goes stale as schemas evolve, and infrastructure sprawl from external vector stores. We demonstrate ATLAS , which addresses all three by co-locating schema metadata, semantic annotations, and vector embeddings e...

Qing Zhang, Shijing Hu, Zhihui Lu · 0 citations
Review Open access Sep 2026

Provenance-Native Audit Infrastructure for LLM-Maintained Wiki Knowledge Systems

LLM-maintained wiki systems accumulate structured knowledge by having a language model incrementally build and revise a persistent wiki layer over raw evidence sources. When such systems support decision-making in regulated domains, each output must be traceable to grounded evidence through an auditable chain. This pap...

Bai-Ling Zhang · 0 citations
Preprint Aug 2026

Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning

As"AI Scientists"emerge to drive research via the Model Context Protocol (MCP), systems relying on ephemeral scripts will fail. The sheer scale of stateful, interconnected evidence requires a machine-walkable warranty grounded in a purpose-built database architecture. Eigenius is an open-source, typed knowledge-graph D...

Hans-Martin Will, Allen L. Brown, M. Fuchs · 0 citations
Preprint Aug 2026

Beyond the Harness: End-to-End Optimization of Context Artifacts for Enterprise Text-to-SQL

In this ablation, retrieved knowledge-base context provides the largest marginal improvement when added to the full oracle graph, and a distillation procedure that turns historical query profiles into reusable SQL reference cards is optimized.

Kate Gwimm, Carson Eisenach · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.