When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages
This paper initializes byte embeddings directly from the subword representations of a frozen base model, applies a chunk alignment loss to project dynamically grouped byte chunks toward precomputed subword targets, and interleave lightweight part-of-speech supervision to guide boundary detection.