Back to #generative ai
#generative ai Review Open access

A semantic firewall for proactive governance of synthetic tabular data in generative AI pipelines

Aug 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 28 references

Abstract

The generation of synthetic tabular data by unconstrained generative models has become a routine component of modern, data-centric machine-learning pipelines. Suéch generators reproduce the statistical structure of real data, yet statistical fidelity does not guarantee logical validity. A generator may emit records that are plausible with respect to every marginal and correlation of the training data while violating the inter-attribute dependencies that make a record meaningful—for instance, an individual simultaneously never-married and a husband. This failure mode, termed here semantic poisoning, is shown to pass undetected through the distributional monitoring on which pipelines ordinarily rely. Three contributions are made. First, a Semantic Integrity Risk Score (SIRS) is proposed: a learned, generator-agnostic measure of the degree to which a record violates the functional and logical dependencies mined from trusted data. Second, a Semantic Firewall applies SIRS at the point of data ingress to route each record to a pass, review, or quarantine decision. Third, an empirical study on real census data across three structurally distinct generators demonstrates that the proposed signal recovers injected semantic violations with an F1 of 0.86–0.95, whereas marginal drift detection remains near zero; that it generalizes to violations of dependencies withheld from any explicit rule set; and that it prevents downstream harm which aggregate-accuracy monitoring cannot perceive, at a bounded screening cost of approximately thirty milliseconds per thousand records. Semantic screening is thereby positioned as a practical, proactive complement to statistical monitoring rather than a replacement for it.

Read PDF

Similar papers

#generative ai Open access Sep 2026

The socio-ecological costs of AI: Toward socially responsible and sustainable communication practices

The adoption of generative artificial intelligence among communication practitioners and researchers surged after the launch of ChatGPT in November 2022, urging practitioners to critically engage in exploring pathways for fostering socially responsible and environmentally sustainable AI practices.

Emma Christensen · 4 citations · ⚡1
#generative ai Review Open access Sep 2026

Toward an AI-integrated nursing curriculum: A Kano model analysis of generative AI competency needs.

Clinical nurses' GenAI learning needs are currently oriented toward practical, application-focused skills, and curriculum development may benefit from a phased approach that prioritizes high-impact practical skills while progressively incorporating foundational, ethical, and advanced competencies.

Yeru Xia, Jingbang Liu, Kaili Wang et al. · 1 citation · ⚡1
#generative ai Aug 2026

AI and Bullshit

It is argued that both AI and bullshitters are untrustworthy informants, and for similar reasons, it is natural to describe AI’s informational outputs as bullshit, as it signals their distinctive kind of epistemic deficiencies, which they share with bullshit.

Duncan Pritchard · 1 citation

Related blog posts