Skip to content
Book Open access

CoSID: Concept-Conditioned Semantic-ID Decoding for Efficient Generative Recommendation

Sep 2026 · Proceedings of the 20th ACM Conference on Recommender Systems · 1 citation · 15 references

TL;DR

CoSID is introduced, a concept-conditioned SID decoder that encodes the history once into a compact next-item concept and delegates the entire beam search to a lightweight KV-cached decoder.

Abstract

Generative recommenders retrieve items by autoregressively decoding semantic IDs (SIDs). The standard autoregressive SID interface (SID-AR) represents each history item with K code tokens, expanding a T-item history to TK tokens, while trie-constrained beam search repeatedly re-enters the full backbone during generation. Composing each item’s K code embeddings into a single input token restores item-level history length, but every decoding step still runs through the full backbone. We introduce CoSID, a concept-conditioned SID decoder that encodes the history once into a compact next-item concept and delegates the entire beam search to a lightweight KV-cached decoder. Across four datasets under a global temporal split, CoSID matches or surpasses the baselines in accuracy, maintains comparable or broader catalog coverage, and delivers up to 6.2 × higher throughput at beam width 100 and 7.3 × at beam width 1000. Because autoregressive SID decoding no longer re-enters the backbone, throughput remains nearly independent of backbone depth.

Read PDF

Similar papers

Preprint Aug 2026

Preserving Item Semantics for Free: Rethinking Token Initialization in LLM-Based Generative Recommendation

This work proposes a simple, parameter-free intervention that initializes SID token embeddings directly from their corresponding centroids in the semantic embedding space, and shows that preserving SID geometry, beyond shared-prefix structure, provides a simple and effective semantic prior for LLM-based GR.

Donald Loveland, Liam Collins, B. Kumar et al. · 0 citations
Book Open access Sep 2026

Codebook-Based Semantic IDs in Generative Recommendation: Enabling Interface, Emerging Bottleneck

Codebook-based semantic IDs (SIDs), short discrete code sequences produced by quantization over item representations, made generative recommendation practical by turning catalog-scale retrieval into low-cardinality generation. Yet the same interface now concentrates the field’s hardest questions. We revisit the canonic...

Danil Gusak, Evgeny Frolov · 3 citations
#artificial intelligence Review Sep 2026

From Interests to Semantic IDs: Retrieval-Grounded Credit Assignment for Generative Recommendation

Semantic IDs (SIDs) encode each catalog item as a short token sequence, enabling generative recommenders to predict the next item autoregressively. Reasoning-enhanced variants, an increasingly common extension, first generate a textual trace and then decode a next-item SID by beam search. Such recommenders are commonly...

Meng-Dan Zhu, Yu-Fan Zhao, Yao Zhao et al. · 0 citations
Preprint Aug 2026

Rethinking Item Tokenization in Generative Recommenders: From Fixed Atoms to Semantic Subwords

In generative recommender systems, items are typically tokenized into fixed-length semantic ID sequences for autoregressive next-item prediction. However, for user-context modeling, this fine-grained representation triggers Intra-item Attention Overload: excessive attention is spent on low-level intra-item dependencies...

Xin-Rui Miao, Mingjia Yin, Jiaqing Zhang et al. · 0 citations
Preprint Aug 2026

Difficulty-Aware Semantic-ID Optimization for Generative Recommendation

Difficulty-Aware Semantic-ID Optimization (DASO), a tree-aware post-training method that addresses failure mode as an online rollout-allocation problem and improves over MiniOneRec-style GRPO on 11 of 12 metrics and achieves the best result on 9 of 12 metrics.

Xin Yu, Stephen Li, S. Aghaei et al. · 0 citations
Preprint Aug 2026

History-Conditioned Joint-Prefix Alignment for Generative Recommendation

Generative recommendation retrieves items by autoregressively generating semantic identifiers, but beam search may discard a target before its complete identifier is generated. Our preliminary analysis across three benchmarks shows that most missed targets are pruned within the first two decoding steps, highlighting th...

Hong-Liang Sun, Lian-Jie Li, Bolin Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.