Skip to content

Author

Maria Ibrar

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Using Large Language Models for Automated Corpus Annotation and Linguistic Analysis: A Critical Methodological Framework

Large language models (LLMs) are increasingly used to classify, label, summarize, and interpret large text collections, creating new possibilities for corpus linguistics. Their capacity for zero-shot and few-shot instruction following could reduce the cost of linguistic annotation and extend analysis beyond the categories handled by conventional part-of-speech taggers, parsers, and dictionary-based tools. At the same time, LLM outputs are probabilistic, prompt-sensitive, model-dependent, and potentially biased, raising fundamental questions about measurement validity, annotation reliability, and reproducibility. This article critically synthesizes foundational corpus-annotation principles with recent evidence on LLM-based text annotation and develops a validated human–LLM workflow for corpus research. The framework distinguishes token-, span-, sentence-, document-, and discourse-level annotation; requires a human-coded gold sample before large-scale deployment; treats prompt design as part of the annotation manual; and evaluates accuracy, precision, recall, F1, inter-annotator agreement, stability across repeated runs, subgroup performance, and error types. Recent studies show that LLMs can approach or exceed crowd-worker performance on some well-specified classification tasks, but that performance varies substantially across datasets, languages, models, prompts, text lengths, and annotation complexity. Span-level annotation and context-dependent semantic or pragmatic coding remain particularly challenging. The article therefore argues against unvalidated full automation and proposes selective automation, disagreement-based human adjudication, model/version documentation, and preservation of raw outputs. For corpus linguistics, the strongest near-term use of LLMs is as flexible annotators within a transparent, theory-driven, and auditable pipeline rather than as replacements for linguistic expertise. The resulting framework supports scalable corpus annotation while preserving the empirical principles on which corpus-based linguistic inference depends.

Maria Ibrar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.