MOBI: Monolithic Graph-Language Modeling Beyond Modality Interference
Abstract
Graph-Language Models (GLMs) aim to endow LLMs with structure-grounded reasoning ability, yet existing solutions often struggle with modality interference : structural information can disrupt pretrained linguistic reasoning, while language cues can overwhelm structural signals. Mainstream modular GLMs with an external graph encoder attempt to mitigate the interference by separating graph encoding from language decoding. This separation fails to strike an effective balance between modality fusion and interference, exhibiting limited cross-modal interaction while leaving interference between modalities largely unresolved. To tackle the above challenges, we propose MOBI (Monolithic Graph-Language Modeling Beyond Modality Interference), a monolithic graph-language model that unifies graph encoding and language decoding within a single backbone for end-to-end graph-text fusion. Specifically, MOBI overcomes modality interference via (i) dual-pathway transformer that preserves pretrained linguistic knowledge while acquiring structural understanding, (ii) progressive interaction scheduling that suppresses cross-modal noise by dynamically regulating the information flow, and (iii) correlation-guided attribute perturbation that discourages textual shortcuts on text-attributed graphs. Extensive experiments on 11 datasets across various settings demonstrate that MOBI consistently outperforms modular baselines.