Skip to content
Conference Open access

GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System

May 2026 · IEEE International Conference on Data Engineering · pp. 3848-3860 · 0 citations · 28 references
Computer Science

TL;DR

GALA is a three-stage pipeline extending existing learning frameworks, whose core innovation lies in an intermediate “generative RL alignment” stage that constructs multimodal pretraining data from user behavior and refines it via conversion-based rewards, effectively bridging the pretraining-fine-tuning gap to align with downstream objectives.

Abstract

Modern recommender systems in food delivery increasingly leverage multimodal signals-including images, text, and user interaction histories-to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging due to discrepancies in semantics, scale, and update frequency, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent. In mainstream two-stage approaches, the separation between content-semantic pretraining of image-text encoders and behavior-driven ranking models limits alignment between semantic understanding and user behavior patterns. To address these issues, we present GALA, a three-stage pipeline extending existing learning frameworks, whose core innovation lies in an intermediate “generative RL alignment” stage that constructs multimodal pretraining data from user behavior and refines it via conversion-based rewards, effectively bridging the pretraining-fine-tuning gap to align with downstream objectives. Specifically, GALA comprises three tightly coupled stages: first, behavior-aware triplet pretraining on query-image-text pairs from search logs to early capture user intent and content preferences; second, the novel intermediate stage, which refines multimodal embeddings through rewarddriven optimization (GRPO) to dynamically align them with user behavior and bridge the pretraining-fine-tuning gap; and finally, integration of multimodal and ID embeddings via adaptive gating with a hybrid loss, preserving multimodal contributions under long-term ID-dominant training. GALA has been deployed in the production environment at Taobao Shangou, serving over 200 million daily active users. Compared with state-of-the-art (SOTA) methods, it delivers consistent offline gains of +0.12/+0.20 AUC along with better PCOC metrics. Large-scale online A/B tests further report a 0.55% increase in order volume, confirming GALA's effectiveness at industrial scale and its robustness across diverse demand patterns.

Read PDF

Similar papers

Conference Open access Sep 2026

Three Minds, One Student: Online Multi-Teacher Knowledge Distillation for Multimodal Recommenders

Existing multimodal recommendation models using complex fusion mechanisms (e.g., attention) or multi-stage processes (e.g., early or late fusion) integrate different modalities. However, attention-based adaptive fusion is prone to shortcut learning, where dominant collaborative signals (ID) can overshadow other modalit...

Hang-Tong Xu, Yuanbo Xu, En Wang · 0 citations
#large language models Review Open access Sep 2026

PALRec: Large Language Model-Based Sequential Recommendation with Parameter-Preserving Augmentation

PALRec is proposed, a parameter-preserving augmentation framework that equips an LLM with recommendation capabilities while keeping its original parameters fixed and consistently outperforms fully fine-tuned counterparts in recommendation accuracy while preserving the LLM’s pre-trained knowledge.

Hyunsoo Na, Minseok Gang, Sang-goo Lee et al. · 0 citations
Open access Sep 2026

Cross-Modally Aligned and Temporally Gated Mixture of Experts for Multimodal Sequential Recommendation

Multimodal Sequential recommendation alleviates the semantic insufficiency and data sparsity of item-ID-based models by incorporating side information such as text and images. However, multimodal systems face the dual challenges of feature-space heterogeneity and modality-specific noise, in addition to the dynamic evol...

Yu-Yin Meng, Ai-Xiang Cui, Jun-Lin Zhou et al. · 0 citations
Book Open access Feb 2026

Distribution-Level Contrastive Supervision for Generative Recommendation

SODA, a plug-and-play alignment framework that adopts a BPR-style contrastive objective to align recommender representations with target-side distributional representations against negative ones, is developed and demonstrated that SODA consistently strengthens diverse generative recommendation architectures.

Zi-Qiu Xue, Ding-Xian Wang, Yi-Meng Bai et al. · 0 citations
#machine learning Preprint Sep 2026

Latent-Aligned Reasoning for Multimodal Recommendation

LARK (Latent-Aligned Reasoning frameworK), a two-stage latent reasoning framework with complementary alignment mechanisms within a single VLM, achieves state-of-the-art performance across multiple recommendation architectures, with controlled ablations confirming the distinct contribution of each component.

Jia-Rui Jin, An-Ya Ji · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.