Full Glyph Images Beat Token Embeddings: A Controlled Study for Transformers
A dual-branch controlled framework is constructed in which both a Vision-based model and an index-based baseline share an identical decoder backbone, training objective, optimizer, and data curriculum, suggesting that transformers are more modality-agnostic than commonly assumed, and that discrete tokenization is not a fundamental requirement for Chinese language modeling.