Recommendation Algorithm Fusing Multimodal Auxiliary Information for Multi-view Social Recommendation
Abstract
Aiming at the problems of data sparsity, cold start and the limitation of single-modal auxiliary information in recommendation systems, this paper proposes a multi-view social recommendation algorithm that fuses multimodal auxiliary information. The algorithm consists of a user-consistent social recommendation module and a multimodal information fusion module. The former dynamically measures the preference similarity between the target user and social neighbors with an attention mechanism, selects preference-consistent neighbors through a differentiable gate for neighborhood aggregation, and extracts item features from both the knowledge graph view and the user-item interaction view, which are integrated by a cross-attention fusion unit. The latter encodes user raw features into low-dimensional dense vectors with multi-layer perceptrons, performs multi-layer information propagation with residual connections, and adopts attention-based late fusion for textual and visual modalities so that modality weights adapt to users and scenarios dynamically. Experiments are conducted on two public datasets, Yelp Open and LastFM-2k, compared with six baselines (BPR-MF, NeuMF, LightGCN, SocialMF, MMGCN, VBPR) using Recall, NDCG and Precision as evaluation metrics. Experimental results show that the proposed algorithm achieves the best performance on both datasets, and ablation studies verify the effectiveness of user consistency filtering, social aggregation, multi-view item feature extraction and dynamic multimodal fusion.