Skip to content
Review

SynthGuard-ReleaseBench: Locked-Audit Evidence for Synthetic Tabular Data Releases

Aug 2026 · 0 citations · 41 references
Computer Science Mathematics

TL;DR

SynthGuard-ReleaseBench is introduced, an audit framework that locks the use, candidate panel, tolerances, and audit schedule before evaluation, and is a reproducible workflow for use-specific release evidence, not a claim that any generator is private, safe, or deployment-ready.

Abstract

Synthetic tabular data are often judged by realism, privacy, or downstream-task scores. Those scores do not answer whether a proposed release is supported for a named use, population, and threat model. We introduce SynthGuard-ReleaseBench, an audit framework that locks the use, candidate panel, tolerances, and audit schedule before evaluation. It compares real-trained and synthetic-trained workflows on protected data, gives simultaneous finite-sample bounds for bounded loss gaps, requires controls, and keeps utility, empirical privacy risk, mechanism claims, and human release authority separate. Across four American Community Survey studies, five non-ACS records, two chronological diagnostics, and a sealed prototype, the benchmark retains favorable, unfavorable, and excluded outcomes. Transparent baselines pass some locked audits; compact learned models fail under the declared budgets; a health-table case is excluded because its negative control passes. A post-audit scaling arm, repeated across three generation seeds, shows the same locked criterion admitting those learned models once they are fit on enough data while still rejecting a dependence-destroying control at every size, so the criterion discriminates rather than merely rejects; the same repetition withdraws a finer single-seed ordering. The theory adds a pre-audit sample-size rule, variance-adaptive and anytime-valid certificates that tighten the bound two to ten times on the same locked evidence, a temporal certificate for time-ordered audits, and two lower bounds: ordinary bounded queries reconstruct a protected audit once the query budget reaches its size, and the panel-size correction is necessary rather than conservative. The contribution is a reproducible workflow for use-specific release evidence, not a claim that any generator is private, safe, or deployment-ready.

View source

Similar papers

Preprint Aug 2026

ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets

Researchers often choose a proxy dataset from many releases, transformations, or seeds. Search can make an invalid release appear adequate, while one adequate release does not establish that its generator is reliable. ProxyGuard controls both errors using prespecified bounded risks and a sealed target set. Named-release mode corrects for multiplicity and certifies specific releases. Direct shared-target mode evaluates independent mechanism draws on a common target, lower-bounds their favorable-score rate, and subtracts a bound on favorable scores contributed by invalid releases. Conditional on the target, release scores are independent, yielding a finite-sample mechanism-reliability guarantee without independent target batches or assumptions on release-level $p$-value dependence. We show that the mean-only penalty is sharp and derive a smooth-score certificate with additive target concentration. In a registered three-requirement study, direct mode raises power from 5.6\% to 64.2\% at reliability 0.95, while named mode remains stronger under high-signal evidence. Prospective audits span full-pipeline Rice--TVAE, which retrains on every draw, and a non-tabular text mechanism.

Dipesh Mahato, Pramod Dhungana · 0 citations
#small language model Preprint Aug 2026

From Subjective Judgments to Auditable Standards:Protocol-Guided AI Auditing of Website Redundancy

Website redundancy does not have a single fixed meaning. The same repeated element may distract during one task and provide backup during another. We introduce CORA (Counterfactual, Observable Redundancy Audit), which measures repetition load, normal-use tax, and failure-domain recovery reserve separately. Each run retains screenshots, stable element identities, and task traces. A versioned vision-language model proposes the annotations. Typed validation and release checks then determine whether a calibrated dimension can be reported; failed or malformed outputs stay in the fixed denominator. On a transparent mechanistic testbed, the factorized CORA representation separated reserve from normal-use tax and predicted perturbed success more accurately than scalar-load baselines. The model studies then showed why repeatability is not enough: two small local vision-language models produced recurring outputs, but neither instrument met all release requirements. CORA therefore withheld automated scores from both instruments while retaining the raw responses and failure records. Separate checker fixtures confirmed that the typed validator and hardened release gates implement their specifications; these tests do not establish semantic grounding or accuracy on production sites. Taken together, the results position CORA as an auditable candidate procedure for the controlled benchmark studied here rather than a general standard. Human agreement, AI-versus-human accuracy, and validation on independent production sites remain open empirical questions.

G. Kong, Yongtong Cao · 0 citations
Open access Aug 2026

Auditing GenAI–Student Grade Claims on Public Datasets: Nested Controls, Frozen Thresholds, and Claim Labels

Public generative artificial intelligence (GenAI)–student datasets invite contested links between AI intensity, usage style, and grades, yet many analyses treat predictive accuracy or significant coefficients as sufficient evidence while skipping prior achievement, co-outcome leakage checks, and absolute effect-size thresholds. This paper presents a construct-audit protocol that treats associational claim survival as a reproducible labeling task: a feature-role taxonomy, forbidden-feature gates, nested out-of-fold change-in-R2 materiality thresholds, and operational labels (stable, vanished, artifact-born—the last defined but not positively observed here), with predictive models used as instruments rather than as the scientific product. On the public ai_student_impact_dataset, treated as a construct-audit sandbox (possibly synthetic or engineered; no campus-population or causal claims), a five-seed Ridge-primary run is used to validate those rules rather than to estimate GenAI effects: all eight primary intensity and style claims are non-material under locked absolute gates (AI joint change-in-R2≈0.0064 versus prior grade-point-average lift ≈0.859), while a kitchen-sink OLS significance foil stars 12/19 coefficients that the inventory does not promote. A report-only Random Forest check shows that style and joint-block clearance can depend on the modeling instrument; inventory labels remain Ridge-primary under the locked metric. The protocol can therefore withhold GenAI–GPA claims when absolute gates fail, and a six-step laptop workflow is specified so educational researchers can apply the same checks without reproducing the full validation schedule.

Ke-Fu Chen · 0 citations
#machine learning Preprint Aug 2026

Creation begins with understanding: LLMs as strategy designers for privacy-preserving tabular data synthesis

Sharing tabular data in high-stakes domains is constrained by privacy regulations. Synthetic data offer a promising alternative, but deep generative models are costly to train and difficult to audit, while LLM-based methods often serialize records as text, obscuring tabular structure and exposing sensitive data. We introduce Tabular Synthesis Strategy Designer (TabSSD), which uses an LLM to design synthesis procedures rather than directly generate records. TabSSD provides the LLM with tree-derived summaries of variable dependence rather than raw records, which produces Python programs for local execution and evaluation. Across twelve datasets, TabSSD strikes a favourable balance among statistical fidelity, predictive utility, and empirical privacy risk, achieving the best average rank across six metrics among ten methods. Moreover, it substantially reduces local computation and token consumption relative to the compared methods. By enabling human-guided refinement and eliminating user-side model tuning, TabSSD lowers the expertise and infrastructure barriers to transparent tabular data synthesis.

Jin-Meng Li, Quan Zhang, Hangting Ye et al. · 0 citations
Preprint Aug 2026

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research

ABE-Ralph is introduced, a reference-anchored auditing framework that represents claims, protocols, required components, baselines, and metrics as structured experimental constraints, guides implementation through an 8-step workflow, and performs quantitative, qualitative, and code-level verification.

Le-Zhi Yu, Xiaogang Xu, Yuhong Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.