Auditing Information Boundaries in AI Evaluation: Reproducibility Data and Code (v3.1)
Abstract
This record contains the research data, executable analysis code, and reproducibility materials associated with the manuscript “Auditing Information Boundaries in AI Evaluation: Evidence-Gated Claims, Counterfactual Testing, and Fault Injection.” The package supports an exploratory, locally controlled synthetic software experiment with 6,000 run-level results, spanning logistic regression, decision tree, and random forest classifiers, together with 20 model-by-condition summary groups. It also includes analytical four-choice binomial tests, controlled target-hint and fault-injection scenarios, pre-confirmation paired counterfactual probes, separately implemented numerical checks, dependency specifications, and six reproducible figures. The 6,000 evaluation runs reuse three fitted classifiers: logistic regression in 10 conditions with 400 runs per condition, and decision tree and random forest each in five conditions with 200 runs per condition. They are not 6,000 independently trained models. The pilot intervenes on a known artificial target-hint feature. The confirmation-only bypass is a programmed phase change, not a learned attacker. All targets and features are synthetic and publicly seeded. No human participants, protected information, operational language-model data, deployed system, secret target acquisition, or independent third-party validation are included. The method's inability to detect information leakage activated only during confirmation is explicitly documented. The 128-pair pilot is an exploratory retained reference comparison, not an independently preregistered protocol. Results do not constitute a security certification or establish real-world attack-detection rates. Read README_ZENODO.md and the archive's README.md before reproducing the results. Archive version: v3.1.