Skip to content

Author

Zhibek Sarsen

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Tokenization Disparity in Kazakh: Cross-Tokenizer Measurement and Low-Cost Vocabulary Adaptation for Large Language Models

Large language models are billed by the token, and tokens are not distributed evenly across languages. We measure how many tokens ten widely deployed tokenizers require to encode identical content in Kazakh and English, using 900 sentence-aligned pairs from OPUS-100 with bootstrap confidence intervals over 10,000 resamples. Kazakh requires 3.16x more tokens than English under GPT-4's tokenizer and 2.62x under Qwen2.5-7B. We show this is not a property of the Kazakh language. On the same sentences XLM-R needs 1.32x, and GPT-4o — one generation after GPT-4, from the same provider — needs 1.75x rather than 3.16x, a 44% reduction achieved purely by changing the tokenizer. Vocabulary size does not predict the premium: mBERT achieves a lower ratio than Qwen2.5 with a smaller vocabulary. We then test vocabulary extension as a remedy. Adding 5,875 Kazakh tokens to Qwen2.5-0.5B with subword-composition initialisation, and training only the new embeddings, yields 1.718x compression together with a 14.12% improvement in bits per character, at 1.2% additional parameters and 24 minutes on a free-tier GPU. An ablation with random initialisation reaches only 7.38%. However, on out-of-domain Kazakh the compression gain persists (1.591x) while the quality gain reverses (+5.74% BPC), showing the improvement is substantially domain adaptation rather than general gain.

Alikhan Akim, Zhibek Sarsen · 0 citations
#large language models Open access Sep 2026

Tokenization Disparity in Kazakh: Cross-Tokenizer Measurement and Low-Cost Vocabulary Adaptation for Large Language Models

Large language models are billed by the token, and tokens are not distributed evenly across languages. We measure how many tokens ten widely deployed tokenizers require to encode identical content in Kazakh and English, using 900 sentence-aligned pairs from OPUS-100 with bootstrap confidence intervals over 10,000 resamples. Kazakh requires 3.16x more tokens than English under GPT-4's tokenizer and 2.62x under Qwen2.5-7B. We show this is not a property of the Kazakh language. On the same sentences XLM-R needs 1.32x, and GPT-4o — one generation after GPT-4, from the same provider — needs 1.75x rather than 3.16x, a 44% reduction achieved purely by changing the tokenizer. Vocabulary size does not predict the premium: mBERT achieves a lower ratio than Qwen2.5 with a smaller vocabulary. We then test vocabulary extension as a remedy. Adding 5,875 Kazakh tokens to Qwen2.5-0.5B with subword-composition initialisation, and training only the new embeddings, yields 1.718x compression together with a 14.12% improvement in bits per character, at 1.2% additional parameters and 24 minutes on a free-tier GPU. An ablation with random initialisation reaches only 7.38%. However, on out-of-domain Kazakh the compression gain persists (1.591x) while the quality gain reverses (+5.74% BPC), showing the improvement is substantially domain adaptation rather than general gain.

Alikhan Akim, Zhibek Sarsen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.