An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026
Excess vocabulary, a word's frequency above its pre-2023 trend, is how the change in scholarly English after 2022 has been measured. We adapt it to Korean with morphological units on 398,296 KCI abstracts (2018–August 2026), with 47,165 Vietnamese abstracts for comparison. Placebo floors are 0.1–2.2 points for the single-word statistic and at most 2.9 for the re-selected split-half set statistic. Korean abstracts show nothing in 2023, onset in late 2024, a rise through 2025 flattening in mid-2026: 시사하다 "suggest" appears in 21.4% of 2026 abstracts against 5.3% expected; plain verbs like 알아보다 "look into" fall to a quarter of trend. Under stated assumptions the single-word conditional lower bound on LLM-processed abstracts is 3.5%, 10.5% and 16.1% for 2024–2026 and a split-half set bound 7.8%, 20.6% and 33.0%. Holzwarth et al.'s estimator under the same discipline gives 41.9% and 72.1% for 2025–2026. Subject-matter controls reduce but do not remove it: restricting the set to lemmas three language-model annotators all call style leaves 14.7 of the 33.0 points, and pairing each 2026 abstract with its journal's closest base-period abstract leaves 34.1. Tested translation routes do not explain it: the surface marks of translated Korean fall as the markers rise. In the same articles' English abstracts the excess appears a year earlier; where the English side carries none, the Korean shift persists at 30 to 66% of the rate where it does. Control abstracts from three providers reproduce the rising words, with marker turnover consistent with model generations; implied prevalences are scenario-dependent. Working paper, version 8 (4 September 2026). Version 8 adds a declaration of language-model use (Section 9) that names every model that took part and its role, gives the agent pipeline with session and instruction counts, and states the fractions of analysis code, build scripts and English text produced by the coding agent; reimplements the mixture estimator of Holzwarth, González-Márquez and Kobak (2026) on the Korean abstracts under the split-half discipline, with the in-sample, cross-fitted and placebo values side by side (Section 5.15); and carries the version history in Appendix J. Version 7 is 10.5281/zenodo.22110398. The reproducibility package contains the analysis code, per-year document-frequency tables for Korean, English and Vietnamese, the generated control abstracts, the annotation, matching, translation-route and robustness results, the Holzwarth-estimator scripts and their validation against the released data, the results manifest and the manuscript consistency gate. A public Korean AI-style dictionary and checker built from the same measurements are at https://os.intframe.com/report/ai-style-dictionary-ko and https://os.intframe.com/report/ai-style-check-ko.