Skip to content
Review

The Annotation Bottleneck in Persian Text NLP: Persian as an Annotation-Scarce Language

Aug 2026 · 0 citations · 58 references
Computer Science

Abstract

Persian (Farsi) is often described as a low-resource language in natural language processing, but that label collapses distinct shortages into a single category. This paper argues that Persian is more precisely described as annotation-scarce, provided that the term is understood as a property of its NLP resource ecology rather than an intrinsic property of the language. The review covers 34 representative Persian text resources available by July 2026 and adds three quantitative cross-checks. First, independent web measurements place Persian among roughly the twenty most visible content languages: W3Techs reports Persian on about 0.9% of websites with a known content language, while Common Crawl CC-MAIN-2026-30 identifies Persian as the primary language of 0.7039% of HTML pages. Second, a selective speech review shows a long resource trajectory from FARSDAT to recent corpora containing hundreds or thousands of hours of speech. Third, a matched Persian-English comparison normalizes task-specific annotation volumes by relative Common Crawl web presence. The resulting ratios vary sharply: Persian syntax and news NER are comparatively dense, whereas natural-language inference falls below the web-proportional baseline. The evidence therefore does not support a simple claim that Persian is globally deficient in labeled volume. Instead, annotation scarcity is expressed through uneven task and domain coverage, incompatible schemes, access and documentation friction, and limited supervision for specialist domains, preference data, and varieties beyond standard Iranian Persian.

View source

Similar papers

Preprint Aug 2026

PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

This work introduces PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words, and conducts the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms.

Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi et al. · 0 citations

Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition

Neyshekar is presented as an open Persian read-speech corpus designed for coverage of both formal and informal language, named entities, and longer utterances. In version 6, 62,279 validated recordings totalling 99.02 hours are provided from 190 contributors, with 34,541 distinct recorded prompts. The prompt pool was a...

Ahmad Amirivojdan, Farzad Nadiri, Abolfazl Alizadeh et al. · 0 citations
#natural language process... Preprint Sep 2026

Correction as Annotation: Bootstrapping a Dependency Parser for Documentary Medieval Latin

Medieval documentary sources remain inadequately served by existing natural language processing tools. None of the five readily available Latin treebank models attains usable performance on a collection of 160 inventories compiled in Marseille between 1258 and 1446. The best labelled attachment score is 0.62 and the be...

G. Pizzorno · 0 citations

Leveraging LLMs to Automatically Construct WordNets as Bilingual Resources

This paper proposes automated methods to construct high-quality WordNets using large language models (LLMs) to generate missing lemmas to address the synset shortfall in non-English and low-resource languages.

Johann Bergh, J. Waitelonis, Melanie Siegel · 0 citations
Aug 2026

A semi-automated LLM-based framework for word sense disambiguation in Serbian

LLM-assisted sense assignment with a Serbian WordNet-based custom inventory, iterative inventory expansion, and expert validation is combined with a constrained JSON-formatted output to support the practical construction and refinement of sense-annotated resources in a low-resource setting.

Saša Petalinkar, R. Stanković, Milica Ikonić Nešić et al. · 0 citations
Open access Aug 2026

NerAxom: a BIO-tagged NER dataset and hybrid neural-rule framework for Assamese

The NerAxom dataset is presented, a BIO-tagged NER dataset for Assamese comprising 4,173 sentences and 106,046 tokens annotated across seven entity categories, and a set of language-specific post-processing rules based on morphological suffixes and keyword cues are introduced.

Punam Sarmah, M. Lahkar, Shobhanjana Kalita et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.