Session 1:
1A: Semantic IDs & Generative Interfaces
Date: Tuesday September 29, 10:45 – 12:15 CDT
Session Chair: Arnie Bhadury
- RESEmpowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding
by Yuxuan Hu, Yuhao Wang, Tianbo Huang, Chao Zhang, Ziwei Liu, Lihua Zhang and Xiangyu ZhaoCross-domain sequential recommendation (CDSR) aims to model users’ dynamic interest transitions and sequential patterns across multiple domains. Recently, generative recommendation (GR) has emerged, which first learns semantic identifiers (SIDs) using semantic information of items and models the recommendation task as autoregressive generation. However, it faces with two critical issues: 1) ignoring collaborative correlations across different domains in tokenization step and 2) adopting inefficient decoding strategies like beam search in generation step, which hinders GR’s application in real-time services. To address these issues, we propose GenCDSR, an effective and efficient generative framework for CDSR. Specifically, GenCDSR learns domain-aware SIDs through a cross-domain hybrid tokenization mechanism, which jointly incorporates domain-shared and domain-specific codebooks to capture both cross-domain commonalities and distinctions. Furthermore, we design a serial-parallel decoding strategy that partially parallelizes cross-domain generation while preserving generation consistency, thus significantly reducing inference latency. Experimental results on three public datasets validate that GenCDSR achieves a 1.5% improvement in accuracy and an 85.1% reduction in inference latency on average compared to SOTA baselines. The implementation code and datasets are available online: https://anonymous.4open.science/r/GenCDSR.
- RESTopology-Aware Tokenization for Generative Recommendation
by Yaokun Liu, Yifan Liu, Zhenrui Yue, Gyuseok Lee, Zelin Li, Ruichen Yao and Dong WangGenerative recommendation has emerged as a powerful paradigm by reframing sequential recommendation as an autoregressive generation task. Central to this paradigm is item tokenization, which quantizes continuous item embeddings into discrete semantic IDs for autoregressive item prediction. Despite its importance, we identify a critical yet overlooked issue in the tokenization process: topology distortion. Specifically, we observe that the intrinsic adjacency relationships of items in the continuous embedding space are significantly disrupted after quantization. This topology distortion misleads the model’s perception of item similarity, ultimately bottlenecking the accuracy of generative recommendations. To address this issue, we propose Hierarchical Topology Distillation (HiToD), a topology-aware tokenization framework that preserves item relational structure throughout the quantization hierarchy. Different from the prior monolithic supervision in tokenization, HiToD introduces a multi-level distillation scheme to progressively recover the topology from coarse to fine granularity: 1) Inter-Group Distillation to capture global cluster-wise relations; 2) Intra-Group Distillation to refine local structures within semantic clusters; and 3) Inter-Item Distillation to enforce fine-grained alignment at the individual item level. Extensive experiments on three benchmark datasets demonstrate that HiToD effectively alleviates topology distortion and consistently outperforms state-of-the-art tokenizers, achieving significant performance gains of up to 9.42% in Recall@5. Our code is available at https://anonymous.4open.science/r/HiToD.
- RESBeyond Fixed Depths and Widths: Optimizing Textual Decoding Tries in LLM-based Generative Recommendation
by Jingzhe Liu, Hanbing Wang, Jiliang Tang, Liam Collins, Tong Zhao, Neil Shah and Mingxuan JuGenerative recommendation (GR) is an increasingly popular paradigm in recommender systems, with a prominent line of work using LLMs as autoregressive backbones to predict the next item’s term IDs (e.g., titles or keywords). The success of autoregressive generation hinges on constrained beam search over a decoding trie to ensure that generated outputs correspond to valid items. However, current research predominantly focuses on generating more comprehensive term IDs to describe items, while largely neglecting the structural design of the decoding trie formed by these terms. This can lead to a trie that is poorly suited to beam search, which degrades performance. To address this, we examine the effectiveness of term IDs from the perspective of decoding trie optimization. Through empirical and theoretical analyses, we identify two desirable properties for a highly performant trie: (1) adaptive and variable ID length, enabling items with varying semantic richness to be represented by IDs of appropriate lengths, and (2) constrained branching factors, especially at shallow levels, which drastically improves the success rate of constrained beam search. Motivated by these properties, we introduce BONSAI: Branching-Optimized Node Structure for Adaptive Identifiers, a novel framework that co-designs textual term IDs and their underlying decoding trie. BONSAI extracts recommendation-informative words from item metadata and employs a minimum set cover formulation to recursively build a trie that satisfies the above properties. Experiments reveal that BONSAI achieves up to a 21.6% relative improvement over state-of-the-art baselines. Further analyses confirm the crucial role of our proposed properties, and demonstrate their generalizability to be applied to enhance the performance of other term ID methods.
- RESHierarchical Semantic Tokenization for Generative Recommendation
by Tianxin Wei, Xuying Ning, Xuxing Chen, Ruizhong Qiu, Yupeng Hou, Yan Xie, Shuang Yang, Zhigang Hua and Jingrui HeGenerative recommendation models next-item prediction as autoregressive generation over tokenized user histories, where each item is represented as a sequence of discrete tokens. However, existing methods typically construct these tokens by compressing heterogeneous item attributes, such as ID, category, title, and description, into a single latent representation before quantization, which obscures the hierarchical structure of item semantics and limits their ability to capture how user preferences evolve from broad interests to specific choices during web interactions. To address this issue, we propose NAME, a generative recommendation framework that explicitly incorporates Coarse-to-Fine semantic structure into both item tokenization and decoding. Specifically, NAME organizes item information into multiple semantic levels, spanning high-level categories, fine-grained textual content, and collaborative signals. Building on this design, we introduce the CoFiRec Tokenizer, which tokenizes each semantic level independently while preserving their structural order, thereby better reflecting how users refine their preferences and enabling more structured generation. During autoregressive decoding, the language model generates item tokens progressively from coarse to fine, allowing the recommendation process to better capture the natural refinement of user intent. Extensive experiments on multiple public benchmarks and backbone models demonstrate that NAME consistently outperforms existing baselines, and our theoretical analysis further shows that structured hierarchical tokenization reduces the expected dissimilarity between generated items and ground-truth targets.
- RESNot All Branches Are Equal: Adaptive Semantic ID Construction for Generative Recommendation
by Guy Hadad, Haggai Roitman and Bracha ShapiraGenerative recommendation models represent items as sequences of discrete tokens known as semantic IDs. Existing approaches derive these identifiers using residual or hierarchical vector quantization, assuming that tokens carry independent semantic meaning across quantization levels. Revisiting this assumption, we argue that semantic IDs should behave as hierarchical paths where each token’s meaning is conditioned on its prefix. This highlights a mismatch between conventional fixed-codebook quantization and the hierarchical structure implicitly learned by generative recommenders. Motivated by this insight, we propose Adaptive Silhouette Tree (AST), a top-down divisive clustering method that constructs semantic IDs as paths within a tree structure. Unlike prior methods with fixed branching factors, AST dynamically determines the number of branches at each node by maximizing the silhouette coefficient, enabling finer partitions in dense regions and compact groupings in homogeneous ones. We further introduce constrained variants that enforce minimum probability mass during splitting to improve supervision for long-tail items. Extensive experiments on five Amazon Review datasets of varying sizes and domains show that AST significantly outperforms strong baselines. These results demonstrate that aligning semantic ID construction with the inherent hierarchical structure of generative recommenders leads to more effective recommendations.
- PPFCodebook-Based Semantic IDs in Generative Recommendation: Enabling Interface, Emerging Bottleneck
by Danil Gusak and Evgeny FrolovCodebook-based semantic IDs (SIDs), short discrete code sequences produced by quantization over item representations, made generative recommendation practical by turning catalog-scale retrieval into low-cardinality generation. Yet the same interface now concentrates the field’s hardest questions. We revisit the canonical two-stage SID lineage (RQ-VAE, R-KMeans, and close descendants) as a past-present-future story. Our thesis is that these identifiers succeeded as an enabling interface, but not yet as a stable theory of item identity for recommendation: one code must simultaneously preserve behaviorally useful similarity, protect identity under head-tail skew, survive catalog drift, align with language-model generators, and remain decodable under production constraints. We organize the present literature around five recurring fractures – objective mismatch, identity-versus-sharing tension, static codes in dynamic settings, LLM alignment tax, and serving feedback into identifier design – and argue that SID utility is regime-dependent: model capacity, catalog scale, and interaction density jointly determine whether semantic structure helps or is redundant. We close with a research agenda centered on recommendation-native tokenization, adaptive identity, explicit alignment protocols, and co-design of tokenization with serving.
- RESAddressing Cross-Stage Decoupling of Semantic and Collaborative Signals in Generative Recommendation
by Jiayi Dan, Weijian Li, Yongqi Liu and Kaiqiao ZhanGenerative recommendation reformulates sequential recommendation as autoregressive generation by encoding items into semantic tokens, enabling improved scaling capability and cross-domain generalization. However, existing generative recommender systems typically follow a two-stage pipeline, where item tokenization is largely dominated by textual semantics with limited incorporation of collaborative signals and interaction similarity, leading to code assignments that are misaligned with downstream generation. Conversely, the generation stage tends to overlook the original semantic information, as the code sequences are re-embedded based on interaction data. This cross-stage information decoupling limits semantic coherence and recommendation accuracy. To address this issue, we propose SCRec, a general framework that enhances cross-stage coherence through bidirectional information supplementation. Specifically, we introduce (i) collaborative-enhanced tokenization to explicitly inject textualized collaborative signals into semantic tokenization, without introducing additional alignment task, (ii) semantic-guided generation to dynamically recalibrate semantic priors with learnable code embeddings in generation stage, and (iii) manifold alignment to reconcile the geometric mismatch between the embedding space of discrete codebook indices and the dense continuous semantic space. These interrelated components form a general framework that aligns semantic and collaborative signals and enhances cross-stage information coherence, with minimal additional training and inference costs. Extensive experiments demonstrate the effectiveness, robustness, and generalizability of our proposed framework.
1B: Personalized Search, Retrieval & Ads
Date: Tuesday September 29, 10:45 – 12:15 CDT
Session Chair: Sergio Oramas
- RESSPEAR: Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval for Community Search
by Wenbin Wu, Yuzhong Wu, Yufan Xu, Kuan Fang, Xing Xu, Cheng Ye and Xiaobin HuQuery reformulation bridges user intent and retrieval in e-commerce search, yet production systems optimize rewrite quality and retrieval effectiveness separately, leaving the two stages structurally misaligned. Path-based architectures unify them end-to-end but were designed for personalization, where relevance is not an explicit constraint—search additionally requires the rewrite to remain faithful to the user’s stated query intent. Transplanted directly, these models learn a shortcut we term the generalization-word dominance effect: they favor generic rewrites that score well on paths but drift from query intent. To address this, we propose SPEAR (Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval), which integrates three components each targeting one failure mode: (1) a dual-embedding backbone with gradient isolation that shields recall-side semantics from being eroded by CTR-driven ranking signals; (2) a multiplicative gating aggregator that lets a rewrite score high only when both its confidence and item relevance are strong, eliminating the generic-rewrite shortcut; (3) a relevance-aligned auxiliary loss that steers the selector toward rewrites faithful to the original query intent, directly enforcing rewrite relevance. Offline evaluation on 100K held-out industrial search sessions shows that the proposed framework improves rewrite semantic similarity@10 by +18.2% and click recall@10 by +99.5% over the production baseline. In online A/B testing, SPEAR achieves +0.259% in query-view CTR and +0.733% in average reading depth, confirming that improved rewrite selection translates into stronger retrieval and deeper user engagement.
- RESOn Reranking Space for Multi-Tenant Retrieval with Adapted Queries
by Jun Woo Chung and Weijie ZhaoMulti-tenant dense retrieval systems increasingly employ shared compressed indexes where individual tenants adapt embeddings via fine-tuning (e.g., LoRA). While query-side projection adapters bridge the resulting embedding mismatch, a critical design choice remains for the optional reranking stage: should distances be computed in the original index space (source-space) or the adapted query space (target-space)? Contrary to the intuition that the calibrated target-space should perform better, we find the opposite to be true. Across 23 LoRA-adapted dataset-tenant pairs, four adapter architectures, and four index configurations (ranging from PQ-16 to HNSW), source-space reranking consistently outperforms target-space reranking, improving nDCG@10 by more than 10 percentage points in some cases. We further evaluate distance blending between these signals, finding that it provides robust gains on coarse indexes (e.g., PQ-16) when the reverse adapter is structurally sound, while adding minimal latency. Our results offer a straightforward heuristic for multi-tenant platforms: maintain the shared index, project queries using existing adapters, rerank in the source-space, and apply distance blending as latency permits when working with low-precision indexes.
- RESDart: Adaptive Tweedie Likelihood Mitigates the σ Escape Route in Cross-Platform Popularity Prediction
by Tomohiro MimuraAlthough predicting social media popularity is crucial for modern recommender systems, it remains a significant challenge due to heavy-tailed, zero-inflated, and non-negative target distributions. While architectural innovations have advanced the field, standard objectives such as Gaussian negative log-likelihood and mean squared error are often poorly matched to these distributions, creating a critical performance bottleneck. In this paper, we diagnose one concrete manifestation of this mismatch, which we refer to as the “sigma escape route”: under Gaussian negative log-likelihood, the noise parameter sigma absorbs prediction errors on heavy-tailed data, preventing the model from improving its mean prediction mu. To address this issue, we propose replacing Gaussian negative log-likelihood with an adaptive Tweedie likelihood, which ties variance directly to the mean (var(Y)=phi mu^p) and naturally handles non-negative, zero-inflated, and heavy-tailed targets in a single parametric family. Notably, our empirical results demonstrate that the Tweedie likelihood consistently improves the Spearman’s rho independent of the model architecture. We also introduce Dart, a purpose-built retrieval-augmented system that achieves state-of-the-art ranking across multiple platforms. Dart excels in cold-start scenarios, substantially outperforming meta-learning baselines without requiring target-platform data. Furthermore, our downstream evaluations confirm that these ranking improvements translate directly into practical recommender-system tasks.
- INDMESH: Scaling Up Retrieval with Heterogeneous Content Unification
by Jiaxing Qu, Yilin Chen, Junpeng Hou, Jinfeng Rao, Olafur Gudmundsson, Sai Xiao and Huizhong DuanOptimizing large-scale retrieval hinges on the ability to efficiently surface candidates across diverse content tiers. However, to capture segments such as fresh and long-tail content, modern systems typically resort to a fragmented “zoo” of specialized retrieval models. This operational complexity is attributed to a fundamental challenge in heterogeneous retrieval systems — the Scaling Bias of Heterogeneity — where model capacity gains do not apply equally across diverse content tiers. To bridge this gap, we propose MESH as a unified retrieval scaling framework that mitigates this bias through a modularized architecture integrated with gated bias correction. By partitioning the feature space into independent domains, MESH enforces a structural inductive bias that reduces interference between sparse-item signals and high-frequency engagement features. This protected gradient path leads to improved scaling behavior for sparse content, empirically validated by a 14× improvement in the power-law scaling exponent for fresh items. In online evaluations on Pinterest’s Related Pins platform — a billion-scale item-to-item recommendation system — these improvements translate into a +5.5% lift in fresh-item repins, alongside with 55% improvement in funnel efficiency and +0.46% improvement in user retention. Finally, our asynchronous serving strategy ensures production viability by delivering a 2.87× improvement in system throughput. Our findings suggest MESH as a promising paradigm for consolidating fragmented retrieval infrastructures into more scalable and ecosystem-aware backbones.
- INDFrom Exploitation to Balance: Efficient Retargeting Recommendation via DPO-Guided Slot Allocation
by Jiangwei Deng, Xiruo Shi, Lifang Deng, Dan Wang, Linke Zhao, Yujing Song and Xiaoyi ZengIn e-commerce recommender systems, items that users have previously interacted with are frequently re-recommended, a practice known as retargeting recommendation. While retargeted items offer significantly higher conversion efficiency by leveraging verified user interests, their inherent predictability induces a systemic bias in conventional recommendation pipelines. Specifically, retargeted items are prone to redundant retrieval across multiple recall channels and systematic score inflation by ranking models that over-rely on historical behavior signals. This unchecked dominance results in recommendation homogenization and user fatigue. We term this phenomenon the retargeting homogenization trap, which ultimately erodes long-term platform value. To address this problem, we propose RetargetRec, a novel industrial framework that formulates retargeting recommendation as an independent optimization problem. RetargetRec introduces a retargeting-aware recall module to implement explicit scale control and a Direct Preference Optimization guided by Item-level Reward (DPOIR) paradigm at the ranking stage. By training on item-level preference pairs derived from a composite reward, DPOIR enables optimization beyond immediate feedback signals without the computational overhead of conventional reinforcement learning. Additionally, a preference-based slot allocation mechanism then jointly governs the number, positions, and ordering of retargeted items. Extensive online A/B tests on the Lazada platform demonstrate that RetargetRec effectively suppresses overexposure while delivering significant gains in both transactions (+2.3\%) and GMV (+4.4\%).
- INDEmbedding Subspace Partitioning for Dynamic Multi-Objective Retrieval
by Shaobo Zhang, Alice Leung, Yunxiang Ren, Ping Liu, Yuchin Juan, Qianqi Shen, Benjamin Le, Jianqiang Shen, Chengming Jiang, Ko-Cheng Wang, Vidya Krishnamurthy, Caleb Johnson, Fedor Borisyuk, Luke Simon, Jingwei Wu and Wenjing ZhangModern industrial recommender systems must optimize across competing objectives, balancing semantic relevance with business metrics such as engagement and revenue. While bi-encoders dominate large-scale retrieval due to their efficiency, they collapse these heterogeneous signals into a single static embedding space. This design creates a fundamental limitation: once trained, the retriever cannot adapt to shifting objective priorities at serving time without retraining. Moreover, joint optimization with multi-objective losses often induces interference between objectives, leading to suboptimal trade-offs. We propose \emph{Embedding Subspace Partitioning} (\emph{ESP}), a retrieval framework that decomposes the embedding into task-aware subspaces and replaces the single dot product with a weighted sum of per-subspace similarities, whose weights are tunable at serving time. For Transformer bi-encoders, ESP uses the model’s native end-of-sequence token as a segment delimiter, with segment-aware attention masking and position encoding resets to guarantee subspace isolation in a single forward pass. Serving is performed via GPU-accelerated exhaustive kNN over one concatenated index, eliminating the need for per-objective Approximate Nearest Neighbor (ANN) infrastructure required by multi-head approaches. We evaluate ESP on an open-source benchmark built from MS~MARCO~\cite{nguyen2016ms}. A single ESP model traces a broad Pareto frontier, consistently outperforming strong multi-task baselines across diverse operating points. In LinkedIn’s job matching platform (70M+ weekly users), ESP enabled dynamic retrieval reconfiguration and delivered significant key business metric lifts.
- INDSMART: LLM-Augmented Hybrid Retrieval for Dynamic Product Ads
by Congfei Zhang, Jingxiao Ma, Xiaodong Liu, Hsiang-Wei Chao, Siman Wang, Ge Liu, Shantanu Aggarwal, Vincent Zhang, Xiao Bai, Yunzhi Zhou, Yajun Wang, Zhe Liu, Jinchao Li, Yu Zhang, Rachel Liao and Meghana MissulaDynamic Product Ads (DPA) require retrieving relevant items from multi-million product catalogs, balancing two competing ob- jectives: retargeting (re-surfacing known interests) and prospecting (discovering new categories). While Large Language Models (LLMs) capture semantic intent better than traditional embedding models, deploying them at scale introduces prohibitive inference costs and lexical mismatch issues. Through controlled experiments on mil- lions of users, we demonstrate a critical retrieval decomposition: rule-generated queries excel at retargeting on a lexical BM25 in- dex, while LLM-generated queries excel at prospecting on a dense ANN index. Building on this, we propose SMART (SeMantic-aware Adaptive ReTrieval). To manage costs, a lightweight quality gate identifies coverage gaps in initial keyword results, adaptively rout- ing only the ∼10% of users who benefit from semantic prospecting to the LLM path. Offline evaluation demonstrates that this gated approach captures the bulk of semantic prospecting gains in Rel- evance Score while maintaining competitive re-targeting performance at a 90% reduction in LLM costs. Finally, in a 2-week online A/B test at Snap, SMART improved the ad conversion rate by +27.6% over a strong embedding-based baseline.
RecSys 2026 (Minneapolis)
- About the Conference
- Registration
- Program at Glance
- Program
- Call for Contributions
- Challenge
- Keynotes
- Accepted Contributions
- Presenter Instructions
- Workshops
- Tutorials
- Committees
- Inclusion
- Student Volunteers
- Women in RecSys
- Visa Information
- Addressing Attendance Issues
- Location / Hotel
- Lasting Impact Award
- 60 Milestones for 20 Years




















