May 2026· IEEE International Conference on Data Engineering· pp. 3848-3860· 0 citations· 28 references
Computer Science
TL;DR
GALA is a three-stage pipeline extending existing learning frameworks, whose core innovation lies in an intermediate “generative RL alignment” stage that constructs multimodal pretraining data from user behavior and refines it via conversion-based rewards, effectively bridging the pretraining-fine-tuning gap to align with downstream objectives.
Abstract
Modern recommender systems in food delivery increasingly leverage multimodal signals-including images, text, and user interaction histories-to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging due to discrepancies in semantics, scale, and update frequency, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent. In mainstream two-stage approaches, the separation between content-semantic pretraining of image-text encoders and behavior-driven ranking models limits alignment between semantic understanding and user behavior patterns. To address these issues, we present GALA, a three-stage pipeline extending existing learning frameworks, whose core innovation lies in an intermediate “generative RL alignment” stage that constructs multimodal pretraining data from user behavior and refines it via conversion-based rewards, effectively bridging the pretraining-fine-tuning gap to align with downstream objectives. Specifically, GALA comprises three tightly coupled stages: first, behavior-aware triplet pretraining on query-image-text pairs from search logs to early capture user intent and content preferences; second, the novel intermediate stage, which refines multimodal embeddings through rewarddriven optimization (GRPO) to dynamically align them with user behavior and bridge the pretraining-fine-tuning gap; and finally, integration of multimodal and ID embeddings via adaptive gating with a hybrid loss, preserving multimodal contributions under long-term ID-dominant training. GALA has been deployed in the production environment at Taobao Shangou, serving over 200 million daily active users. Compared with state-of-the-art (SOTA) methods, it delivers consistent offline gains of +0.12/+0.20 AUC along with better PCOC metrics. Large-scale online A/B tests further report a 0.55% increase in order volume, confirming GALA's effectiveness at industrial scale and its robustness across diverse demand patterns.
Existing multimodal recommendation models using complex fusion mechanisms (e.g., attention) or multi-stage processes (e.g., early or late fusion) integrate different modalities. However, attention-based adaptive fusion is prone to shortcut learning, where dominant collaborative signals (ID) can overshadow other modalit...
Hang-Tong Xu, Yuanbo Xu, En Wang· Proceedings of the Thirty-Fi...· 0 citations
PALRec is proposed, a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed and consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge.
Hyunsoo Na, Minseok Gang, Sang-goo Lee et al.· ACM Transactions on Informat...· 0 citations
Multimodal Sequential recommendation alleviates the semantic insufficiency and data sparsity of item-ID-based models by incorporating side information such as text and images. However, multimodal systems face the dual challenges of feature-space heterogeneity and modality-specific noise, in addition to the dynamic evol...
Yu-Yin Meng, Ai-Xiang Cui, Jun-Lin Zhou et al.· Big Data and Cognitive Compu...· 0 citations
SODA, a plug-and-play alignment framework that adopts a BPR-style contrastive objective to align recommender representations with target-side distributional representations against negative ones, is developed and demonstrated that SODA consistently strengthens diverse generative recommendation architectures.
Zi-Qiu Xue, Ding-Xian Wang, Yi-Meng Bai et al.· Proceedings of the 20th ACM...· 0 citations
LARK (Latent-Aligned Reasoning frameworK), a two-stage latent reasoning framework with complementary alignment mechanisms within a single VLM, achieves state-of-the-art performance across multiple recommendation architectures, with controlled ablations confirming the distinct contribution of each component.
Jia-Rui Jin, An-Ya Ji· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.