Open data portals face a persistent challenge: datasets and their metadata often remain incomprehensible to non-specialist audiences, limiting the democratic potential of open government data. While recent advances in large language models offer promising approaches to automated description generation, their application to multi-level accessible descriptions of structured datasets, particularly in non-English contexts, remains underexplored. This paper presents the design, implementation, and preliminary evaluation of a modular pipeline that generates dataset descriptions at three accessibility levels using locally deployed language models. The system integrates deterministic dataset profiling (Frictionless Data and ydata-profiling) with generative AI, employing Qwen3.5-9B for description generation and an initial two-model evaluation setup for automated quality assessment (Prometheus 2 and Mistral 7B), which was later critically examined against human expert judgement. We address practical constraints of public sector deployment, including reproducibility, digital sovereignty through local model execution, and alignment with operationalized criteria derived from German accessibility standards. Our implementation processes datasets from GovData.de, generating descriptions conforming to DIN 8581-1 (Plain Language German), DIN SPEC 33429 (Easy Language German), and Standard German for general public audiences. This work contributes both a reusable technical architecture and methodological insights for accessible metadata generation in open data ecosystems, with particular attention to cross-lingual prompting strategies and the limitations of the DCAT-AP.de metadata standard.
The adoption of generative artificial intelligence among communication practitioners and researchers surged after the launch of ChatGPT in November 2022, urging practitioners to critically engage in exploring pathways for fostering socially responsible and environmentally sustainable AI practices.
This R script (make_kessan10_csv.R) converts the Local Government Financial Settlement Survey (市町村別決算状況調), published by the Ministry of Internal Affairs and Communications on its annual pages of local government financial status survey materials, into machine-readable CSV. The source workbooks are print-oriented Excel files with multi-row merged headers, issued as four separate files per fiscal year (overview and expenditure, for cities and for towns and villages). The script consolidates them into long-format panels carrying fiscal year and municipality type as columns, and also writes one file per fiscal year. The output of a run over ten fiscal years (FY2015–FY2024) is deposited alongside it: all 1,741 municipalities, with 33 overview indicators and 94 expenditure items classified by purpose, giving panels of 17,410 rows each. Every municipality and every year is checked for internal consistency: the components of each expenditure category sum to that category's total, and the sum of all categories matches the total expenditure reported in the overview table. All checks passed for all ten years. Amounts are in thousands of yen, as published; blank cells are left blank rather than filled with zero. The column structure of the source data does not change over the period covered. One definitional change affects the adjusted ratio of current expenditure to current revenue: for FY2020 and FY2021 the special bonds issued for deferred tax collection are removed from current general revenue as well. Four changes of municipality occurred: Tomiya and Nakagawa became cities in FY2016 and FY2018 respectively, each receiving a new municipality code; Sasayama was renamed Tamba-Sasayama in FY2019, and Aogashima was renamed in FY2018 in the written form of its name only, both keeping their codes. The code was written with generative AI: Claude (Anthropic) was used to write and revise it. The author has verified the output and takes responsibility for the content. Version 1.1 corrects the reading of the census population change column in the overview table, where a small negative rate written with the triangle sign used in Japanese official statistics was left blank instead of being read as a number. 56 cells across the ten years were affected; no other value changed.
Yasutoshi Moteki· Zenodo (CERN European Organi...· 1 citation
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.