Session 1:

1A: Semantic IDs & Generative Interfaces

Date: Tuesday September 29, 10:45 – 12:15 CDT
Session Chair: Arnie Bhadury

  • RESEmpowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding
    by Yuxuan Hu, Yuhao Wang, Tianbo Huang, Chao Zhang, Ziwei Liu, Lihua Zhang and Xiangyu Zhao

    Cross-domain sequential recommendation (CDSR) aims to model users’ dynamic interest transitions and sequential patterns across multiple domains. Recently, generative recommendation (GR) has emerged, which first learns semantic identifiers (SIDs) using semantic information of items and models the recommendation task as autoregressive generation. However, it faces with two critical issues: 1) ignoring collaborative correlations across different domains in tokenization step and 2) adopting inefficient decoding strategies like beam search in generation step, which hinders GR’s application in real-time services. To address these issues, we propose GenCDSR, an effective and efficient generative framework for CDSR. Specifically, GenCDSR learns domain-aware SIDs through a cross-domain hybrid tokenization mechanism, which jointly incorporates domain-shared and domain-specific codebooks to capture both cross-domain commonalities and distinctions. Furthermore, we design a serial-parallel decoding strategy that partially parallelizes cross-domain generation while preserving generation consistency, thus significantly reducing inference latency. Experimental results on three public datasets validate that GenCDSR achieves a 1.5% improvement in accuracy and an 85.1% reduction in inference latency on average compared to SOTA baselines. The implementation code and datasets are available online: https://anonymous.4open.science/r/GenCDSR.

  • RESTopology-Aware Tokenization for Generative Recommendation
    by Yaokun Liu, Yifan Liu, Zhenrui Yue, Gyuseok Lee, Zelin Li, Ruichen Yao and Dong Wang

    Generative recommendation has emerged as a powerful paradigm by reframing sequential recommendation as an autoregressive generation task. Central to this paradigm is item tokenization, which quantizes continuous item embeddings into discrete semantic IDs for autoregressive item prediction. Despite its importance, we identify a critical yet overlooked issue in the tokenization process: topology distortion. Specifically, we observe that the intrinsic adjacency relationships of items in the continuous embedding space are significantly disrupted after quantization. This topology distortion misleads the model’s perception of item similarity, ultimately bottlenecking the accuracy of generative recommendations. To address this issue, we propose Hierarchical Topology Distillation (HiToD), a topology-aware tokenization framework that preserves item relational structure throughout the quantization hierarchy. Different from the prior monolithic supervision in tokenization, HiToD introduces a multi-level distillation scheme to progressively recover the topology from coarse to fine granularity: 1) Inter-Group Distillation to capture global cluster-wise relations; 2) Intra-Group Distillation to refine local structures within semantic clusters; and 3) Inter-Item Distillation to enforce fine-grained alignment at the individual item level. Extensive experiments on three benchmark datasets demonstrate that HiToD effectively alleviates topology distortion and consistently outperforms state-of-the-art tokenizers, achieving significant performance gains of up to 9.42% in Recall@5. Our code is available at https://anonymous.4open.science/r/HiToD.

  • RESBeyond Fixed Depths and Widths: Optimizing Textual Decoding Tries in LLM-based Generative Recommendation
    by Jingzhe Liu, Hanbing Wang, Jiliang Tang, Liam Collins, Tong Zhao, Neil Shah and Mingxuan Ju

    Generative recommendation (GR) is an increasingly popular paradigm in recommender systems, with a prominent line of work using LLMs as autoregressive backbones to predict the next item’s term IDs (e.g., titles or keywords). The success of autoregressive generation hinges on constrained beam search over a decoding trie to ensure that generated outputs correspond to valid items. However, current research predominantly focuses on generating more comprehensive term IDs to describe items, while largely neglecting the structural design of the decoding trie formed by these terms. This can lead to a trie that is poorly suited to beam search, which degrades performance. To address this, we examine the effectiveness of term IDs from the perspective of decoding trie optimization. Through empirical and theoretical analyses, we identify two desirable properties for a highly performant trie: (1) adaptive and variable ID length, enabling items with varying semantic richness to be represented by IDs of appropriate lengths, and (2) constrained branching factors, especially at shallow levels, which drastically improves the success rate of constrained beam search. Motivated by these properties, we introduce BONSAI: Branching-Optimized Node Structure for Adaptive Identifiers, a novel framework that co-designs textual term IDs and their underlying decoding trie. BONSAI extracts recommendation-informative words from item metadata and employs a minimum set cover formulation to recursively build a trie that satisfies the above properties. Experiments reveal that BONSAI achieves up to a 21.6% relative improvement over state-of-the-art baselines. Further analyses confirm the crucial role of our proposed properties, and demonstrate their generalizability to be applied to enhance the performance of other term ID methods.

  • RESHierarchical Semantic Tokenization for Generative Recommendation
    by Tianxin Wei, Xuying Ning, Xuxing Chen, Ruizhong Qiu, Yupeng Hou, Yan Xie, Shuang Yang, Zhigang Hua and Jingrui He

    Generative recommendation models next-item prediction as autoregressive generation over tokenized user histories, where each item is represented as a sequence of discrete tokens. However, existing methods typically construct these tokens by compressing heterogeneous item attributes, such as ID, category, title, and description, into a single latent representation before quantization, which obscures the hierarchical structure of item semantics and limits their ability to capture how user preferences evolve from broad interests to specific choices during web interactions. To address this issue, we propose NAME, a generative recommendation framework that explicitly incorporates Coarse-to-Fine semantic structure into both item tokenization and decoding. Specifically, NAME organizes item information into multiple semantic levels, spanning high-level categories, fine-grained textual content, and collaborative signals. Building on this design, we introduce the CoFiRec Tokenizer, which tokenizes each semantic level independently while preserving their structural order, thereby better reflecting how users refine their preferences and enabling more structured generation. During autoregressive decoding, the language model generates item tokens progressively from coarse to fine, allowing the recommendation process to better capture the natural refinement of user intent. Extensive experiments on multiple public benchmarks and backbone models demonstrate that NAME consistently outperforms existing baselines, and our theoretical analysis further shows that structured hierarchical tokenization reduces the expected dissimilarity between generated items and ground-truth targets.

  • RESNot All Branches Are Equal: Adaptive Semantic ID Construction for Generative Recommendation
    by Guy Hadad, Haggai Roitman and Bracha Shapira

    Generative recommendation models represent items as sequences of discrete tokens known as semantic IDs. Existing approaches derive these identifiers using residual or hierarchical vector quantization, assuming that tokens carry independent semantic meaning across quantization levels. Revisiting this assumption, we argue that semantic IDs should behave as hierarchical paths where each token’s meaning is conditioned on its prefix. This highlights a mismatch between conventional fixed-codebook quantization and the hierarchical structure implicitly learned by generative recommenders. Motivated by this insight, we propose Adaptive Silhouette Tree (AST), a top-down divisive clustering method that constructs semantic IDs as paths within a tree structure. Unlike prior methods with fixed branching factors, AST dynamically determines the number of branches at each node by maximizing the silhouette coefficient, enabling finer partitions in dense regions and compact groupings in homogeneous ones. We further introduce constrained variants that enforce minimum probability mass during splitting to improve supervision for long-tail items. Extensive experiments on five Amazon Review datasets of varying sizes and domains show that AST significantly outperforms strong baselines. These results demonstrate that aligning semantic ID construction with the inherent hierarchical structure of generative recommenders leads to more effective recommendations.

  • PPFCodebook-Based Semantic IDs in Generative Recommendation: Enabling Interface, Emerging Bottleneck
    by Danil Gusak and Evgeny Frolov

    Codebook-based semantic IDs (SIDs), short discrete code sequences produced by quantization over item representations, made generative recommendation practical by turning catalog-scale retrieval into low-cardinality generation. Yet the same interface now concentrates the field’s hardest questions. We revisit the canonical two-stage SID lineage (RQ-VAE, R-KMeans, and close descendants) as a past-present-future story. Our thesis is that these identifiers succeeded as an enabling interface, but not yet as a stable theory of item identity for recommendation: one code must simultaneously preserve behaviorally useful similarity, protect identity under head-tail skew, survive catalog drift, align with language-model generators, and remain decodable under production constraints. We organize the present literature around five recurring fractures – objective mismatch, identity-versus-sharing tension, static codes in dynamic settings, LLM alignment tax, and serving feedback into identifier design – and argue that SID utility is regime-dependent: model capacity, catalog scale, and interaction density jointly determine whether semantic structure helps or is redundant. We close with a research agenda centered on recommendation-native tokenization, adaptive identity, explicit alignment protocols, and co-design of tokenization with serving.

  • RESAddressing Cross-Stage Decoupling of Semantic and Collaborative Signals in Generative Recommendation
    by Jiayi Dan, Weijian Li, Yongqi Liu and Kaiqiao Zhan

    Generative recommendation reformulates sequential recommendation as autoregressive generation by encoding items into semantic tokens, enabling improved scaling capability and cross-domain generalization. However, existing generative recommender systems typically follow a two-stage pipeline, where item tokenization is largely dominated by textual semantics with limited incorporation of collaborative signals and interaction similarity, leading to code assignments that are misaligned with downstream generation. Conversely, the generation stage tends to overlook the original semantic information, as the code sequences are re-embedded based on interaction data. This cross-stage information decoupling limits semantic coherence and recommendation accuracy. To address this issue, we propose SCRec, a general framework that enhances cross-stage coherence through bidirectional information supplementation. Specifically, we introduce (i) collaborative-enhanced tokenization to explicitly inject textualized collaborative signals into semantic tokenization, without introducing additional alignment task, (ii) semantic-guided generation to dynamically recalibrate semantic priors with learnable code embeddings in generation stage, and (iii) manifold alignment to reconcile the geometric mismatch between the embedding space of discrete codebook indices and the dense continuous semantic space. These interrelated components form a general framework that aligns semantic and collaborative signals and enhances cross-stage information coherence, with minimal additional training and inference costs. Extensive experiments demonstrate the effectiveness, robustness, and generalizability of our proposed framework.

Back to program