MGTA: Multi-scale Graph Tokens Alignment for CTR Prediction via Pre-trained Language Models
Click-through rate (CTR) prediction is a critical task in personalized recommender systems. Existing methods that align collaborative information from conventional CTR models with semantic information from pre-trained language models (PLMs) have demonstrated superior performance compared to approaches relying on a single information source. However, most of them perform alignment at the embedding level, which introduces noise from heterogeneous vector spaces and limits fine-grained semantic mapping. Moreover, these models highly depend on tabular features, thereby limiting their transferability. To address these challenges, we propose to conduct Multi-scale Graph Tokens Alignment (MGTA) for CTR prediction via pre-trained language models, which enables deep cross-modal information alignment while maintaining strong generalizability. Specifically, MGTA first captures multi-scale graph tokens rich in collaborative signals by decoupling and quantizing graph structures based on graph neural networks (GNNs), and then achieves token-level alignment between collaborative signals and semantic knowledge via PLM fine-tuning. To achieve efficient transfer with MGTA, we further introduce the Cross-domain Token Adapter that enables collaborative signals adaptation by mapping graph tokens from the target domain to the source domain, which necessitates only the injection of target-domain semantic knowledge, in turn reducing fine-tuning time. Extensive experiments on three real-world datasets demonstrate the effectiveness of MGTA compared to existing baselines.