Skip to content

The Impact of Tokenization Algorithms on Hungarian Language Model Performance

2026 · International Conference on Language Resources and Evaluation · pp. 2545-2556 · 1 citation · 45 references
Computer Science

TL;DR

Results show that BPE produces the most compact and morphologically aligned subword representations, while the modified Unigram LM achieved the best overall downstream performance across tasks, underscore that tokenizer choice and vocabulary design are critical determinants of language model efficiency and performance in morphologically rich languages.

View source

Similar papers

Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

TokEval is introduced, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics.

Clara Meister · 1 citation
#machine learning Preprint Aug 2026

TokEval: A Tokenizer Evaluation Suite

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framewor...

Clara Meister · 0 citations
#natural language process... Preprint Sep 2026

A Grapheme-Aware Indic Tokenizer for Tamil: Large-Scale Training and Intrinsic Evaluation

Tokenization forms the foundation of modern Natural Language Processing (NLP) systems by transforming raw text into discrete units that neural language models can process. The effectiveness of this process directly influences vocabulary efficiency, sequence length, computational cost, and downstream model performance....

V. HariKrishnanK, Sudarsun Santhiappan · 0 citations
Conference Open access Sep 2026

Subword Tokenization for Low- and Medium-Resource Languages: A Systematic Evaluation

Subword tokenization is a standard technique for pre-trained language models, mapping text into sequences of tokens from a fixed-size vocabulary. Despite its widespread use, the impact of tokenization algorithms and vocabulary sizes on downstream performance remains underexplored, particularly for low- and medium-resou...

J. Daðason, H. Loftsson · 0 citations
#natural language process... Preprint Sep 2026

To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs

Multilingual Large Language Models (LLMs) traditionally rely on a single vocabulary shared by all supported languages, which can lead to uneven compression across them. Moreover, their large embedding and output matrices increase memory usage and slow inference, notably for small-scale models. It is also wasteful as mo...

Franck Signe, Hippolyte Pilchen, François Yvon et al. · 0 citations
Preprint Aug 2026

What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?

This work analyzes what tokens are learned when tokenization is jointly optimized with language modeling, and finds tokenizer-free approaches optimize for contextual and computational efficiency rather than strict morphological structure, resulting in fundamentally different yet effective vocabularies for downstream NL...

Saketh Reddy Vemula, Parameswari Krishnamurthy · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.