Governing Copyright Responsibly in the Era of Generative AI
This dataset accompanies the study Governing Copyright Responsibly in the Era of Generative AI, which applies the Responsible Research and Innovation (RRI) framework to the governance of copyright in generative AI training data. The dataset consists of three components: (1) a corpus inventory of 38 primary-source policy, legislative, judicial, and soft-law documents covering the period 2019--2026 across five jurisdictions (international instruments, European Union, United Kingdom, United States, and China); (2) a codebook defining eight analytical categories drawn from the RRI literature, each with an operational definition and a set of coding cues; (3) a coded data matrix containing binary code indicators, character-offset pointers (quote_start/quote_end), and document group flags that support quantitative analysis of coding distributions; and (4) derived visualizations of code frequencies, category co-occurrences, jurisdictional distributions, temporal patterns, and document-level coding density. The dataset is designed to facilitate reproducibility and secondary analysis in STS, science policy, and legal scholarship on AI governance, intellectual property, and responsible innovation.