A novel method, dubbed Disentangle and Distillation-based Dynamic Ensemble for multi-modal Recommendation (D3ER), which introduces gradient boosting into MR for the first time to formalize the optimization objective for alternately learning HOI and HEI.
Abstract
Incorporating items'information shared among multiple modalities into a fused representation, multi-modal recommendation (MR) has demonstrated documented success than canonical unimodal recommendation. Although several attempts have been made to extract the discriminative information unique in each modality, existing methods suffer from a core limitation: the joint learning of modal-homogeneity discriminative information (HOI) and modal-heterogeneity discriminative information (HEI) tends to weaken their individual effectiveness. To remedy this deficiency, we propose a novel method, dubbed Disentangle and Distillation-based Dynamic Ensemble for multi-modal Recommendation (D3ER). We introduce gradient boosting into MR for the first time to formalize the optimization objective for alternately learning HOI and HEI. This design enables models dedicated to each type of information to focus on their proficient samples, thereby promoting specialized optimization. Furthermore, to mitigate the inherent high storage cost and risk of local optima in gradient boosting, we enhance our framework with knowledge distillation and a global correction regularization. Experiments on prevalent real-world datasets confirm the superiority of our proposed method on MR.
Sequential recommendation aims to predict users'future interests from their historical interactions. Although Large Language Models (LLMs) capture rich item semantics, existing methods often struggle to align collaborative signals with textual semantic knowledge. As a result, the learned item representations fail to ca...
Shih-Hong Chen, J. Ying, Vincent S. Tseng· 0 citations
Existing multimodal recommendation models using complex fusion mechanisms (e.g., attention) or multi-stage processes (e.g., early or late fusion) integrate different modalities. However, attention-based adaptive fusion is prone to shortcut learning, where dominant collaborative signals (ID) can overshadow other modalit...
Hang-Tong Xu, Yuanbo Xu, En Wang· Proceedings of the Thirty-Fi...· 0 citations
Recent studies in multimodal recommendation, which leverage diverse modal information to address data sparsity and enhance recommendation accuracy, have garnered significant interest. Two critical processes in this domain are modality fusion and representation learning. In representation learning, existing studies ofte...
Jin-Feng Xu, Zhe-Yu Chen, Wei Wang et al.· ACM Transactions on Recommen...· 0 citations
Multimodal Sequential recommendation alleviates the semantic insufficiency and data sparsity of item-ID-based models by incorporating side information such as text and images. However, multimodal systems face the dual challenges of feature-space heterogeneity and modality-specific noise, in addition to the dynamic evol...
Yu-Yin Meng, Ai-Xiang Cui, Jun-Lin Zhou et al.· Big Data and Cognitive Compu...· 0 citations
This work proposes Retrv-MoE, a unified retrieval architecture built upon sparse Mixture-of-Experts (MoE), and theoretically and empirically demonstrates that this conditional computation mechanism provides a structural remedy to optimization interference by decoupling the learning trajectories of conflicting tasks and...
Tongxu Lin, Jiayin Xiao· Proceedings of the 32nd ACM...· 0 citations
Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them fo...
Zirui Cheng, Xun Xu, Tiankai Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.