Skip to content
#generative ai Open access

Pegi1727/GenAI-Identity-Compression-Framework: `Dataset for the Study of Generative AI's Impact on LL2 Writing Identity and Agency: The Authenticity Gap Framework`

Aug 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Description: This dataset accompanies the research article titled "When AI Improves the Text but Changes the Voice: Generative AI, Authorial Identity, and Authenticity in EFL Writing." The repository contains a comprehensive set of anonymized data and analysis materials investigating the tension between AI-mediated linguistic optimization and authorial voice. Contents include: Anonymized Corpus: Paired L2 essays (original independent versions vs. GPT-4o revised versions) from 60 advanced EFL learners. Quantitative Metrics: Linguistic features including MTLD (Measure of Textual Lexical Diversity), T-unit analysis, and stance marker frequencies. Psychological Scales: Raw and processed scores from the Authenticity Gap Scale (AGS). Analysis Scripts: R and Python code used for Linear Mixed-Effects Models (LME) and stylistic homogenization analysis. Research Instruments: The standardized prompt used for AI revision and the semi-structured interview protocol. Keywords: Generative AI, L2 Writing, Authorial Identity, Authenticity Gap, Linguistic Agency, EFL.

View source

Similar papers

#generative ai Open access Sep 2026

The socio-ecological costs of AI: Toward socially responsible and sustainable communication practices

The adoption of generative artificial intelligence among communication practitioners and researchers surged after the launch of ChatGPT in November 2022, urging practitioners to critically engage in exploring pathways for fostering socially responsible and environmentally sustainable AI practices.

Emma Christensen · 4 citations · ⚡1
#generative ai Open access Aug 2026

Ten-Year Panel of Japanese Municipal Finance from the Local Government Financial Settlement Survey

This R script (make_kessan10_csv.R) converts the Local Government Financial Settlement Survey (市町村別決算状況調), published by the Ministry of Internal Affairs and Communications on its annual pages of local government financial status survey materials, into machine-readable CSV. The source workbooks are print-oriented Excel files with multi-row merged headers, issued as four separate files per fiscal year (overview and expenditure, for cities and for towns and villages). The script consolidates them into long-format panels carrying fiscal year and municipality type as columns, and also writes one file per fiscal year. The output of a run over ten fiscal years (FY2015–FY2024) is deposited alongside it: all 1,741 municipalities, with 33 overview indicators and 94 expenditure items classified by purpose, giving panels of 17,410 rows each. Every municipality and every year is checked for internal consistency: the components of each expenditure category sum to that category's total, and the sum of all categories matches the total expenditure reported in the overview table. All checks passed for all ten years. Amounts are in thousands of yen, as published; blank cells are left blank rather than filled with zero. The column structure of the source data does not change over the period covered. One definitional change affects the adjusted ratio of current expenditure to current revenue: for FY2020 and FY2021 the special bonds issued for deferred tax collection are removed from current general revenue as well. Four changes of municipality occurred: Tomiya and Nakagawa became cities in FY2016 and FY2018 respectively, each receiving a new municipality code; Sasayama was renamed Tamba-Sasayama in FY2019, and Aogashima was renamed in FY2018 in the written form of its name only, both keeping their codes. The code was written with generative AI: Claude (Anthropic) was used to write and revise it. The author has verified the output and takes responsibility for the content. Version 1.1 corrects the reading of the census population change column in the overview table, where a small negative rate written with the triangle sign used in Japanese official statistics was left blank instead of being read as a number. 56 cells across the ten years were affected; no other value changed.

Yasutoshi Moteki · 1 citation

Related blog posts