Using Large Language Models for Automated Corpus Annotation and Linguistic Analysis: A Critical Methodological Framework
Abstract
Large language models (LLMs) are increasingly used to classify, label, summarize, and interpret large text collections, creating new possibilities for corpus linguistics. Their capacity for zero-shot and few-shot instruction following could reduce the cost of linguistic annotation and extend analysis beyond the categories handled by conventional part-of-speech taggers, parsers, and dictionary-based tools. At the same time, LLM outputs are probabilistic, prompt-sensitive, model-dependent, and potentially biased, raising fundamental questions about measurement validity, annotation reliability, and reproducibility. This article critically synthesizes foundational corpus-annotation principles with recent evidence on LLM-based text annotation and develops a validated human–LLM workflow for corpus research. The framework distinguishes token-, span-, sentence-, document-, and discourse-level annotation; requires a human-coded gold sample before large-scale deployment; treats prompt design as part of the annotation manual; and evaluates accuracy, precision, recall, F1, inter-annotator agreement, stability across repeated runs, subgroup performance, and error types. Recent studies show that LLMs can approach or exceed crowd-worker performance on some well-specified classification tasks, but that performance varies substantially across datasets, languages, models, prompts, text lengths, and annotation complexity. Span-level annotation and context-dependent semantic or pragmatic coding remain particularly challenging. The article therefore argues against unvalidated full automation and proposes selective automation, disagreement-based human adjudication, model/version documentation, and preservation of raw outputs. For corpus linguistics, the strongest near-term use of LLMs is as flexible annotators within a transparent, theory-driven, and auditable pipeline rather than as replacements for linguistic expertise. The resulting framework supports scalable corpus annotation while preserving the empirical principles on which corpus-based linguistic inference depends.