The role of grammar in transformer encoders: Does it really matter for language understanding process?
Abstract
Transformer-based methods achieve strong performance on language understanding tasks, yet the role of grammar in encoder representations remains insufficiently explained. While decoder-only models implicitly learn language structure through next-token prediction, encoder-only models still dominate many classification and matching tasks, where grammatical relations can affect the decision boundary. We therefore ask two research questions: (1) how are grammatical structures represented within transformer-based encoders, and (2) how can grammatical regulation improve language understanding in these encoders? We first design an interpretable analysis showing that grammar is weakly present in standard encoders but is not explicitly organised without regulation. We then propose a Semantics-Grammar Coupled Space (SGCS), a low-dimensional space in which token positions encode semantic information and spatial relations encode grammatical dependencies. SGCS is trained through sentiment-based spatial regulation and attention-grammar distribution alignment, and it is mapped back to the encoder representation through an invertible mapping network. Experiments on 12 language understanding datasets with BERT and RoBERTa backbones show that SGCS raises syntactic tree similarity from 23.84% to 80.10% and improves downstream performance on most evaluated tasks, with especially consistent gains on sentiment-oriented and grammar-sensitive benchmarks.