Skip to content

Author

Wen-Jie Wang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

SETRec++: Scaling Order-agnostic Identifier for Large Language Model-based Generative Recommendation

Leveraging Large Language Models (LLMs) for generative recommendation has attracted significant research interest, where item tokenization is a critical step. It involves assigning item identifiers for LLMs to encode user history and generate the next item. Existing approaches leverage either token-sequence identifiers, representing items as discrete token sequences, or single-token identifiers, using ID or semantic embeddings. Token-sequence identifiers face issues such as the local optima problem in beam search and low generation efficiency due to step-by-step generation. In contrast, single-token identifiers fail to capture rich semantics or encode Collaborative Filtering (CF) information, resulting in suboptimal performance. To address these issues, we propose three fundamental principles for item identifier design: 1) integrating both CF and semantic information to fully capture multi-dimensional item information, 2) designing order-agnostic identifiers without token dependency, mitigating the local optima issue and achieving simultaneous generation for generation efficiency, and 3) disentangling semantics across different tokens, unlocking the potential of scaling order-agnostic identifier. Accordingly, we introduce a novel set identifier paradigm, representing each item as a set of order-agnostic tokens. To implement, we propose SETRec, which leverages CF and semantic tokenizers to obtain order-agnostic multi-dimensional tokens. To eliminate token dependency, SETRec uses a sparse attention mask for user history encoding and a query-guided generation mechanism for simultaneous token generation. We instantiate SETRec on T5 and Qwen (from 1.5B to 7B). To reinforce disentanglement between tokens, we propose SETRec++, which introduces two disentanglement strategies, i.e., disentangled regularization loss and token masking mechanism. Experiments on four datasets demonstrate its effectiveness across various scenarios (e.g., full ranking, warm- and cold-start ranking, and various item popularity groups). Moreover, results validate SETRec’s superior efficiency and scalability on cold-start items as model sizes increase, and SETRec++’s potential in the scalability of order-agnostic identifier.

Xinyu Lin, Chuan-Bo Zhang, Yu-Fan Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.