This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 102M Chinese words. The models were evaluated on three tracks: natural language understanding, cognitive alignment, and Hanzi knowledge. There were no restrictions on tokenizers, model architectures, or the number of training epochs. Eighteen teams submitted 28 distinct models, generating 74 result files. The overall-winning team used a DeBERTa-v2 architecture and introduced an auxiliary pinyin-prediction objective during pretraining. Several submissions also explored curriculum-learning strategies and architectural innovations. Overall, the challenge provides a benchmark for advancing data-efficient and cognitively plausible approaches to Chinese language modeling.
Siyuan Song, Zhiheng Qian, Yunhao Zhang et al.· 0 citations
The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conjecturing a positive'quality-power'relationship, in which language models'(LMs') fit to psychometric data continues to improve as their ability to predict words in context increases. This is important because it might suggest that elements of LLM architecture reflect the architecture of the human sentence processing faculty, and that any inadequacies in predicting human reading time and brain imaging data may be attributed to insufficient model complexity, which recedes as larger models become available. But recent studies have shown this scaling inverts after a point, as LMs become excessively large and accurate, when information-theoretic surprisal is used as a predictor. Other studies propose the use of entire vectors from differently sized LLMs, still showing positive scaling, casting doubt on the value of surprisal as a predictor, but do not control for dimensionality expansion using untrained LLMs with more than 1.6B parameters. This study evaluates scaling of LLM vector predictors controlled using untrained LLMs with up to 66B parameters. Results show that inverse scaling obtains, and moreover the contribution of trained LMs over corresponding untrained LMs drops to zero at around a few billion parameters on most datasets.
Yi-Chien Lin, Hongao Zhu, William Schuler· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.