Evolution and Adaptation of Large Language Models for Bahasa Indonesia
This paper investigates the engineering methodologies of cross-lingual vocabulary adaptation, parameter initialization heuristics, and language-adaptive pre-training strategies designed to address text overfragmentation, representational misalignment, and tokenization cost inefficiencies in Bahasa Indonesia and its low-resource regional dialects.