Skip to content

Babies Learn to Look Ahead: Multi-Token Prediction in Small LMs

· 1 citation · 33 references

TL;DR

The experimental results show that even 130M-parameter models benefit from including the MTP task in the pre-training objective, and hold even under severe data constraints, as demonstrated on both zero-shot benchmarks and downstream tasks.

View source

Similar papers

Preprint Aug 2026

Hierarchical Latent Prediction for Language Models

Hierarchical Latent Prediction (HiLP) is introduced, which introduces an auxiliary higher-level abstract latent to help reduce the error accumulation effect in latent-space rollouts and can lead to longer-horizon coherent belief state representation.

Changyan Shi, Tim Pearce, Manan Tomar et al. · 0 citations
Aug 2026

Instructing the Learning of Language Model with the Token Interpretation to Improve Language Understanding

Pretrained language models (PLMs) have established state-of-the-art performance across diverse natural language understanding (NLU) tasks. This study reveals that seman-tic-rich explanations of lexical units can effectively guide PLM learning processes. We propose a novel language understanding enhancement method with token interpretation (LUETI) that addresses two critical limitations in conventional PLMs: Incomplete token semantics caused by isolated contextual learning and insufficient semantic encoding in embedding matrices. LUETI operates through dual mechanisms, augmenting token represen-tations by integrating hidden states with corresponding token interpretations and refining embedding spaces using interpretation-derived semantic vectors for token prediction. LUETI, which is implemented as a plug-in module for standard architectures, demonstrates significant improvements on BERT and GLM, achieving average performance gains of 3.36% and 4.87% respectively on the SuperGLUE benchmark with equivalent parameters and training data. Note that LUETI-equipped models attain comparable performance to baseline PLMs using only 60% of pretraining data. Findings establish token interpretation as a computationally efficient but semantically powerful enhancement strategy for language model pretraining.

Tianyi Chen, Yashen Wang, Huan Chang et al. · 0 citations
Preprint Jul 2026

Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models

The results suggest that competitive MLLM can emerge from alignment alone, reducing multimodal extension to a lightweight projector-training problem that generalizes across modalities and adapts rapidly to each new LLM release.

Xuanru Zhou, Yiwen Shao, Jiahong Li et al. · 1 citation
Preprint Jul 2026

Self-Guided Test-Time Training for Long-Context LLMs

A simple method, Self-Guided TTT (S-TTT), which improves accuracy for both Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct, achieving up to a 15% relative improvement.

Xinyu Zhu, Zhenqin Xu, Xiaohan Wei et al. · 0 citations
Preprint Aug 2026

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction

AdaMTP is proposed, an adaptive training paradigm that dynamically aligns the prediction horizon with the intrinsic predictability of the sequence, and consistently outperforms standard MTP in both task performance and inference speedup.

Ziqiang Cui, Han Shi, Bowei He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.