Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as"On the table."is usually produced as four separate predictions for the preposition (On), article (the), noun (table),...
A multi-scale analysis of representations in transformers, SSMs, and hybrid architecture suggests that while transformers and SSMs induce different usage of latent space, they display a striking functional convergence at the level of local semantic manifolds.
This work proposes a self-distillation-based SFT method that facilitates effective factual learning while minimizing hallucinations w.r.t.~pre-existing knowledge by regularizing output-distribution drift and investigates the mechanism, contrasting capacity limitations, behavior cloning, and localized interference.
Guy Kaplan, Zorik Gekhman, Zhen Zhu et al.· arXiv.org· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.