Semantic-Augmented Item Graph Denoising with Collaborative Signals for Multimodal Recommendation
Abstract
Multimodal recommender systems (MMRS) incorporate item-side content, such as images and texts, to complement limited interaction records and improve recommendation for less-observed users or items. However, existing methods still suffer from several limitations: (1) high-level semantic information in multimodal content is difficult to explicitly capture; (2) multimodal signals are often overwhelmed by ID-based collaborative signals during training; (3) item-item graphs constructed solely based on the multimodal similarity may introduce noisy edges due to the semantic-behavior inconsistency. Driven by the above observations, we propose SAGDRec, a Semantic-Augmented Graph Denoising approach for multimodal recommendation. Specifically, we leverage an Multimodal Large Language Model (MLLM) to extract the high-level semantics from the visual and textual content, thereby enriching the multimodal semantic features. Specifically, to prevent the multimodal semantics from being overshadowed by the collaborative signals, we introduce an additional multimodal semantic branch to strengthen the learning driven by semantic cues. Moreover, we build a semantic item-item graph and refine its edges with collaborative signals, so that behaviorally inconsistent semantic neighbors are suppressed during propagation. Experiments show that SAGDRec achieves competitive performance on three public datasets.