Artifact for "Memorised, Not Generated: Verbatim Recall of Published Fixtures in LLM-Generated Test Data, Measured Across Five Models and Eight Public Schemas"
Abstract
Complete artifact for the empirical study "Memorised, Not Generated: Verbatim Recall of Published Fixtures in LLM-Generated Test Data, Measured Across Five Models and Eight Public Schemas" (submitted to Empirical Software Engineering). Version 1.2.1 adds experiment E1c: the copy-rate measurement (exact-text match, near-match at Indel ratio 0.9, and column-wise overlap against the published human fixture) extended to five models (Claude Sonnet 4.6, Claude Sonnet 5.5, GPT-4o, GPT-5.6 Terra, Qwen3-235B-A22B-Instruct-2507 via OpenRouter) and eight public PostgreSQL schemas (Pagila, Chinook, Northwind, employees, Dell DVD Store 2, ClassicModels, Oracle HR, BikeStores), direct-emission arm, one run per cell, with every cell's generated rows and per-call logs (tokens, latency, stop reason, serving host, sampling settings), the three new subjects' PostgreSQL ports and provenance manifest, a check of which edition of the Oracle HR fixture each model reproduces (current v23.3 versus the 2015-2019 v19.2 edition), and a model-identity check recording, for every model identifier named in the paper, the provider's listing entry and the identifier echoed back by a live call. It retains the v1.1.0 content: (E1) direct emission vs plan-then-execute (SynthData 1.2.2) vs no-LLM vs Faker vs human fixtures on five schemas, two providers, two runs, scored by row-by-row insertion into PostgreSQL 16 with all constraints on and by column-wise fidelity metrics; (E1b) the same LLM arms on identifier-renamed schemas as a memorisation control; (E2, exploratory) the provisioned fixtures injected into the RAITG LLM test-generation pipeline (172 requirements, 4 services, 293 frozen mutants) in three arms plus the pre-registered E2c state-contract arm. Contains the pre-registered plan, frozen inputs with checksums, all scores and metric files, the scoring harness, the analysis scripts (stats_paper.py and copy_rate_stats_e1c.py print every statistic with the label used in NUMBER_TRACE.md), the manuscript source, and an MD5 manifest of every file. External inputs (the eight human fixture datasets, SynthData 1.2.2 on npm, the RAITG harness at 10.5281/zenodo.20285103) are referenced with pinned versions and checksums rather than redistributed.