Preprint
Jul 2026
Language Models are not Equally Robust to Non-Canonical Tokenization across Languages
The study of tokenization robustness serves as a diagnostic of how tightly a model is coupled to its tokenizer, and demonstrates that tokenization robustness is not a universal property of language models, but depends strongly on the language and its interaction with the tokenizer.
Poulami Ghosh, P. Jyothi
· 0 citations