Skip to content

Adapting TBXTools to automatic terminology extraction with BERT

2026 · DETEC@MDTT · 1 citation · 5 references
Computer Science

TL;DR

Using the BERT model as a filtering mechanism applied to terminology extraction, the approach used for the DETECH 2026 shared task on monolingual term extraction achieved a significant improvement in both precision and recall.

View source

Similar papers

Aug 2026

Modelling specialized language through AI: challenges in training systems for technical domain translation

Modelling specialized language with artificial intelligence is a significant challenge for specialized translation, especially in fields where controlled terminology is essential, such as medicine, engineering, law, or the exact sciences. The performance of neural machine translation systems depends directly on the qua...

Parascovia Cozma · 0 citations
#natural language process... Preprint Sep 2026

SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers

Terminology-aware translation asks for more than a correct translation: the output must use the exact terms a glossary prescribes. The standard recipe, fine-tuning on glossary-annotated translation pairs, hides an inefficiency: for most examples the glossary prescribes exactly what the model would have produced anyway,...

Xi-Xian Liao, Maite Melero · 0 citations
#natural language process... Preprint Sep 2026

TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs

Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we term translation noise. Despite its prevalence, this problem lacks dedicated benchmarks and...

Shenbin Qian, Yves Scherrer · 0 citations
Open access Sep 2026

Using Large Language Models for Automated Corpus Annotation and Linguistic Analysis: A Critical Methodological Framework

Large language models (LLMs) are increasingly used to classify, label, summarize, and interpret large text collections, creating new possibilities for corpus linguistics. Their capacity for zero-shot and few-shot instruction following could reduce the cost of linguistic annotation and extend analysis beyond the categor...

Maria Ibrar · 0 citations
Open access Aug 2026

A unified framework and benchmark for generalizable biomedical knowledge extraction and applications with large language models

Results demonstrate that InfoFlowEX equips LLMs with robust adaptability, achieving consistent gains over baselines with minimal task-specific customization, highlighting InfoFlowEX for real-world biomedical applications.

Wuyang Lan, Siqi Zhang, Wenzheng Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.