Informational Antilocality and the Locality Bias in LLMs
The ability of transformer-based language models to learn k-antilocal languages, i.e., languages that have no mutual information across any span of $k$ contiguous symbols, is considered, finding that LLMs trained on them achieve comparable cross-entropy loss regardless of antilocality, but converge more slowly on more antilocal languages.