Skip to content

Shared Task: Translating Historical Text to Contemporary Language for Improving Automatic Linguistic Annotation

· 0 citations · 40 references

TL;DR

This work focuses on improving part-of-speech tagging analysis of seventeenth-century Dutch, and finds the best system obtained an error reduction of 51% in comparison with the baseline of tagging unmodified text.

View source

Similar papers

Preprint Aug 2026

Assessing Quality of Experience in Natural Language Generation of German Text

The rapid advancement of Natural Language Generation (NLG) has made the reliable evaluation of generated text increasingly critical, as these systems, such as large language models (LLMs), are now widely deployed in real-world applications. However, traditional automatic metrics fail to capture the multifaceted nature of perceived quality. In this paper, we introduce TextQ-German, a novel dataset suite for human-centered evaluation of German NLG from a Quality of Experience (QoE) perspective, covering automatic text summarization and machine translation. Through crowdsourcing studies with German speakers, we collect human quality ratings and identify relevant perceptual quality dimensions for each task. We develop automatic QoE prediction models, including transformer-based, linguistic feature-based, and hybrid approaches. Hybrid models outperform pure transformer baselines in almost all experimental settings, while linguistic features alone can approach the performance of fine-tuned language models. The dataset is extended with LLM-generated outputs annotated with overall QoE scores. Final validation on held-out sets indicates generalization to unseen data. Our work contributes a publicly accessible resource for NLG evaluation and baselines for automatic QoE prediction, providing a foundation for developing NLG systems that better align with human quality perception.

Dinh Nam Pham, Shushen Manakhimova, Vivien Macketanz et al. · 0 citations
Open access Jul 2026

Gloss-to-Text Translation for Libras and Portuguese: Evaluating Pretrained and Fine-Tuned Encoder-Decoder Models

We evaluate encoder-decoder models for Gloss-to-Text translation from Brazilian Sign Language (Libras) glosses into Portuguese using a corpus derived from Libras-UFPel. The evaluated models are mT5-small, mT5-base, Flan-T5-base, and PTT5-v2-base. Experiments were conducted with 5-fold cross-validation and evaluated using BLEU and chrF. All models improved after supervised fine-tuning, with PTT5-v2-base achieving the best overall performance. The results suggest that Portuguese-specialized encoder-decoder models are a promising direction for Gloss-to-Text translation in low-resource settings.

J. Tomaszewski, B. S. Santana, Antonielle Martins et al. · 0 citations
Jul 2026

A POS Tier Is the Key to Automated Annotation for Low-Resource Language Documentation: Neural Interlinear Glossing for Irabu, a Southern Ryukyuan Language

This work implements a full neural annotation pipeline (morpheme segmentation, POS tagging, glossing) for Irabu Ryukyuan using deliberately small, transparent BiLSTM-CRF models, and concludes with a concrete recommendation for documentation practice: annotate quadrilinearly - text, POS, gloss, translation.

Michinori Shimoji · 0 citations
Open access Sep 2026

Bridging the linguistic divide: recent developments in machine translation for Indian languages

This paper analyses various recent state-of-the-art variants of large language models (LLMs) and neural machine translation (NMT) for Indian languages in comparison to statistical machine translation (SMT) and tackles key questions, such as idiomatic expressions, morphologically complex grammar or the scarceness of parallel corpora.

Jayanand A. Kamble, S. Jadhav, V. J. Kadam · 0 citations
Jul 2026

Parsing Middle High German: exploring cross-lingual NLP for treebank construction in low-resource historical languages

Building syntactically annotated corpora, such as treebanks, for historical languages is a challenging yet vital task in digital humanities, as it underpins linguistic analysis and facilitates a range of interdisciplinary research. However, the scarcity of annotated data and the need for extensive expertise in historical linguistics make this process particularly demanding. In this study, we explore the potential of cross-lingual natural language processing (NLP) techniques as a semiautomatic solution for treebank construction in low-resource historical languages. We use Middle High German (MHG) as a case study. Leveraging the linguistic continuity and structural similarities between MHG and Modern German (MG), we effectively utilize the extensive MG treebank resources to develop a constituency parsing system tailored for MHG. Specifically, to design a semiautomatic system that integrates automatic annotation with manual validation, we explore two cross-lingual transfer techniques: zero-shot transfer and delexicalization; the latter removes lexical information to focus on syntactic structure. In our experiments, we first train parsers on MG treebanks, and then transfer them to MHG using the two cross-lingual transfer techniques. The delexicalization method achieves a parsing performance of 67.3 per cent in terms of F1-score. This performance significantly surpasses the zero-shot cross-lingual method by a margin of 28.6 percentage points. These investigations validate the effectiveness and feasibility of cross-lingual transfer techniques for historical language treebank construction. This study highlights the potential of NLP tools to streamline the semiautomatic annotation process, reducing the reliance on extensive linguistic expertise and manual effort, and paving the way for broader applications in digital humanities research.

Ercong Nie, Siyao Peng, Helmut Schmid et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.