Session 8:
8A: Generative Recommendation: Retrieval, Ranking & Content Generation
Date: Thursday October 01, 10:30 – 12:15 CDT
Session Chair: Cataldo Musto
- RESSAGE: Sequence-level Adaptive Gradient Evolution for Generative Recommendation
by Yu Xie, Xingkai Ren, Qi Ying, Di Jia and Yao HuGenerative recommender systems hold the promise of jointly optimizing accuracy, content diversity, and creator exposure fairness. However, current reinforcement learning–based optimizers such as Gradient-Bounded Policy Optimization (GBPO) exhibit a Symmetric Conservatism failure mode: symmetric update bounds suppress learning from rare positive signals (e.g., cold-start items), static negative-sample constraints fail to prevent diversity collapse under rejection-dominated feedback, and group-normalized multi-objective rewards produce low-resolution training signals. These limitations directly harm user experience by reinforcing information cocoons and reducing new-creator visibility. We propose SAGE (Sequence-level Adaptive Gradient Evolution), a unified optimizer for list-wise generative recommendation. SAGE introduces (i) a geometric-mean importance ratio for sequence-level signal alignment, (ii) asymmetric adaptive bounding—a Positive Boost for cold-start slates and an Entropy-Aware Penalty for low-diversity failures—and (iii) a decoupled multi-objective advantage estimator. On three Amazon Product Reviews datasets and the large-scale RecIF-Bench, SAGE consistently improves top-K accuracy while delivering +89% to +101% cold-start recall recovery and +11% diversity gains relative to GBPO. Beyond-accuracy evaluation confirms that SAGE substantially reduces intra-list similarity and broadens catalog coverage, suggesting that asymmetric, sequence-aware policy optimization is an effective approach to improving both recommendation quality and content ecosystem health.
- RESGenerating Personalized Images for Sparse-Interaction Users with Uncertainty-Aware Retrieval and Dense Knowledge Guidance
by Yuting Zhang, Ying Sun, Dazhong Shen, Ziwei Xie, Feng Liu, Changwang Zhang, Xiang Liu, Jun Wang and Hui XiongPersonalized image generation aims to synthesize target images tailored to individual preferences based on users’ historical interaction data. Existing methods typically inject features from historical interaction records to guide personalized generation. However, such methods encounter two critical challenges when serving sparse-interaction users: (1) Preference Misalignment: Sparse interactions tend to lack target-semantic preference information, causing the direct injection of interaction features to misalign with users’ true target preferences. (2) Lack of Reliable Supervision: Mining preferences from sparse interactions requires sufficient supervision signals, yet personalized generation with sparse data inherently lacks direct feedback or ample preference signals for generated outputs. To this end, we propose Uncertainty-aware retrieval with Dense guidance for Sparse personalized Image Generation (UDSIG). For preference alignment, we first retrieve reference images that exhibit low-uncertainty matching with user sparse preferences from the entire dataset. For reliable supervision, we propose a dense-to-sparse scheme that incorporates a reward model derived from active users’ dense interaction data to drive personalized generation in sparse scenarios. Extensive experiments and human evaluations across three public datasets confirm the superiority of our model, along with its strong generalization to dense interaction scenarios.
- RESCoarse-to-Fine Long-term Interest Modeling for Generative Recommendation
by Shiteng Cao, Junda She, Ji Liu, Bin Zeng, Chengcheng Guo, Kuo Cai, Qiang Luo, Ruiming Tang, Han Li, Kun Gai, Zhiheng Li and Cheng YangLeveraging long-term user behavioral patterns is a key trajectory for enhancing the accuracy of modern recommender systems. Due to the quadratic complexity of attention mechanisms, existing GR models are typically confined to short interaction sequences. While pioneer works have attempted to adapt Search-based Interest Models (SIM) to the generative context, they typically overlook the inherent hierarchical distinction of SIDs. GR is fundamentally a coarse-to-fine generation task, where the initial SIDs (prefix) determine the broad semantic category and the subsequent SIDs (suffix) pinpoint the specific item. Thus, our core insight is that the prefix and suffix of SIDs require distinct long-term signal injections. To bridge this gap, we propose GLASS, a Generative recommendation framework that integrates Long-term user interests into the generative process via SIDTier and Semantic Search. For the generation of SID prefix, we introduce SID-Tier, a module that maps long-term interactions into a unified interest vector to enhance the prediction of the initial SID token. SID-Tier leverages the compact nature of the semantic codebook to incorporate cross features between the user’s long-term history and candidate semantic codes. Furthermore, for the generation of SID suffix, we present semantic hard search, which utilizes generated coarse-grained semantic ID as dynamic keys to extract relevant historical behaviors, which are then fused via an adaptive gated fusion module to recalibrate the trajectory of subsequent fine-grained tokens. Extensive experiments on two large-scale real-world datasets, TAOBAO-MM and KuaiRec, demonstrate that method outperforms state-of-the-art baselines. A two-week online A/B test on a short-video platform demonstrate that GLASS achieves significant gains in recommendation quality. Our codes are publicly available at this anonymous link to facilitate further research in generative recommendation.
- RESUniRec: A Unified Expressive-Aligned Generative Recommendation Framework for E-commerce
by Ziliang Wang, Gaoyun Lin, Xuesi Wang, Shaoqiang Liang, Yili Huang, Weijie Bian, Li Zhang, Mingchen Cai, Jian Dong and Guanxing ZhangTraditional discriminative recommendation pipelines suffer from objective misalignment and error propagation across stages, motivating a shift toward generative recommendation (GR). However, existing GR methods decode over compact Semantic ID (SID) tokens without access to item-side features, lacking the explicit user–item feature crossing that discriminative models rely on. Combined with the inherent one-to-many nature of recommendation, this absence of item-side signals significantly amplifies generation uncertainty, making the generative paradigm widely regarded as having a lower performance ceiling than its discriminative counterpart. We propose UniRec, a unified expressive-aligned GR framework that unifies the multi-stage pipeline into a single generative model and aligns its expressive power with the discriminative counterpart. Via Bayes’ theorem, we show that any practical gap stems from feature coverage rather than modeling asymmetry, motivating Chain-of-Attribute (CoA), an expressive-alignment mechanism that pre-generates item attributes before decoding SIDs, recovering item-side feature crossing and yielding measurable per-step entropy reduction. Beyond CoA, Capacity-constrained SID enforces exposure-weighted load balancing to suppress token collapse, and Conditional Decoding Context (CDC) injects scenario-conditioned signals to stabilize multi-scenario decoding and Cartesian-product-based structured summaries of generated tokens to reinforce conditional dependence across decoding layers. A joint Reward-Driven Fine-tuning (RFT) and Direct Preference Optimization (DPO) framework further aligns the model with business objectives. Deployed on a large-scale e-commerce platform, online A/B tests confirm significant gains in page-view click-through rate (PVCTR, +5.37%), orders (+4.76%), and gross merchandise volume (GMV, +5.60%).
- INDTokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in Recommendation Systems
by Baolei Li, Yiping Yuan, Yilin Zheng, Likang Yin, Ling Liu, Fabio Soldo, Romer Rosales, Xinyang Yi and Lichan HongLarge-scale recommendation systems face “Memory Wall” bottlenecks due to massive, dense embedding tables. While generative retrieval uses discrete tokens for IDs, high-dimensional context still relies on inefficient dense formats. Inspired by computer vision data compression, we propose Dual-purpose Semantic IDs to achieve LLM-level I/O efficiency. Our methodology uses hierarchical quantization to condense continuous embeddings into discrete Semantic IDs performing two concurrent roles: (1) Collaborative Identity: modeling user-item interactions via learnable embedding table; and (2) Content Reconstruction: using a lightweight Semantic Decoder for on-the-fly embedding approximation. This approach replaces massive vector storage with on-demand reconstruction, reducing system overhead and data footprints. We demonstrate the efficacy of our framework through offline evaluations and successful online deployment in production-scale ranking and retrieval systems at a major video sharing platform, showing that discrete tokens are indeed all you need for highly efficient, content-rich recommendation.
- INDGenPage: Towards End-to-End Generative Homepage Construction at Netflix
by Lequn Wang, Jiangwei Pan and Linas BaltrunasWe present GenPage, an end-to-end generative approach to Netflix homepage construction that replaces the traditional multi-stage recommender stack with a single transformer. GenPage represents the user and request context as a prompt, and autoregressively generates the entire structured, multi-row homepage as the response. We adapt the LLM training recipe: pretraining on positively engaged production pages, followed by post-training via weighted binary classification (WBC) or reinforcement learning (RL). For industry-scale deployment, we introduce techniques addressing cold start, model freshness, business-rule enforcement, and serving efficiency. In online A/B tests against a mature, highly optimized production homepage recommender, the WBC variant of GenPage delivered a +0.24% lift on the core engagement metric we use for launch decisions (p < 0.001), while reducing end-to-end serving latency by 20%. Offline experiments yield two findings worth highlighting: enriching the prompt yields a larger improvement than scaling model capacity in our current regime, and RL post-training increases homepage diversity even though diversity is not part of the objective.
- INDUniPinRec: Unifying Generative Retrieval and Ranking at Pinterest Scale
by Hanyu Li, Yi-Ping Hsu, Aditya Mantha, Prabhat Agarwal, Laksh Bhasin, Jialu Wang, Hongtao Lin, Bella Huang, Yaxin Li, Xinyi Li, Chuxi Wang, Kousik Rajesh, Hooshmand Shokri Razaghi, Shunyao Li, Zongyue Qin, Jaewon Yang, James Li, Dhruvil Deven Badani, Jiajing Xu and Charles RosenbergModern recommendation systems predominantly train retrieval and ranking as separate models despite both increasingly relying on large transformers encoding the same user behavior data, duplicating parameters, compute, and serving cost. Prior work unifies the model architecture but not the full pipeline: input formats, training procedures, and serving stacks remain fragmented across stages. We present UniPinRec, which achieves full-stack unification of retrieval and ranking at Pinterest: one input format, one model, one training stage, deployed within existing serving infrastructure. A shared transformer encodes the user action sequence into candidate-independent representations that branch into retrieval (ANN dot-product) and ranking (cross-attention) via task-specific heads. Three ideas make this work: (1) Masked Action Modeling (MAM) eliminates interleaving, enabling weight sharing without doubling context length; (2) Blended training examples pair action sequences with feedview impression slates to satisfy both objectives jointly; (3) Cross-stage KV cache sharing reuses user-history computation from retrieval for ranking, reducing total FLOPs versus serving two independent models. Deployed in the Pinterest core surfaces, UniPinRec delivers approximately +1% online engagement lift while cutting end-to-end serving latency by 11.1% and lifting QPS by 63.6%. To our knowledge, this is the first full-stack unification of retrieval and ranking, covering inputs, model, training and serving, deployed in a production recommendation system.
8B: Evaluation, Reproducibility & Benchmarks
Date: Thursday October 01, 10:30 – 12:15 CDT
Session Chair: Antonela Tommasel- RESGeneralized Position-Based Model: Rethinking Position Weights in Ranking Off-Policy Evaluation
by Norman Knyazev, Vito Bellini, Huseyin Yurtseven and Ben LondonOPE estimates the performance of new recommendation policies using logged data, thus enabling fast, safe and inexpensive iteration prior to costly A/B tests. To evaluate ranking policies, existing OPE estimators all make structural assumptions about user behavior, leading to a spectrum of trade-offs between bias and variance. The recently proposed INTERPOL estimator navigates these trade-offs through a window system that defines how clicks at different positions are combined. However, this approach has two key limitations: the window configuration must be specified a priori, which can impede practical use, and all positions within a window are weighted uniformly, regardless of their relative utility, potentially limiting accuracy. To address these gaps, we introduce GPBM, an estimator that learns position-specific weights by minimizing an approximate upper bound on the estimation error. GPBM retains the unbiasedness guarantees of INTERPOL while significantly simplifying its hyperparameter selection. Our experiments demonstrate that GPBM provides more accurate and robust estimates across a wide range of position bias misestimation levels, logging policies, and data sizes.
- RESOn the Convergent Validity of Offline Evaluation Designs for Recommender Systems
by Sushobhan Parajuli, Samira Vaez Barenji and Michael EkstrandOffline evaluation on historical interaction logs is the most common evaluation methodology for recommender systems. However, such evaluations depend on sparse, incomplete or biased data which raises concerns about whether commonly used evaluation setups reliably reflect true user preferences. In this work, we study how offline evaluation design choices affect the validity of recommender system comparisons. We evaluate a set of recommendation models across many evaluation configurations that vary key factors including data filtering thresholds, feedback binarization versus graded relevance, candidate set construction, train-test splitting strategies, and evaluation metrics. To assess the validity of these configurations, we measure the correlation between model rankings obtained from conventional train-test splits on sparse interaction data and rankings from evaluations based on dense ground-truth relevance judgments. We use this agreement as an evidence of their validity with respect to true user preferences. Using KuaiRec and extended MovieLens-32M datasets that provide such ground-truth data, we analyze which evaluation setups produce results that better align with ground-truth performance.
- RESA Control Function Framework for Mitigating Position Bias in Learning to Rank Systems
by Md Aminul Islam, Kathryn Vasilaky and Elena ZhelevaLearning-to-rank (LTR) systems commonly depend on implicit feedback, such as user clicks, because it is easy to collect and can serve as a valuable signal of user preferences. However, directly optimizing ranking models using implicit feedback data often yields suboptimal performance because such data is inherently skewed by systematic biases. Among these biases, position bias is particularly pervasive: items ranked higher tend to receive disproportionately more interactions, regardless of their actual relevance. To address this, we introduce a novel two-stage framework based on control functions. In the first stage, we utilize exogenous variation from the residuals of the ranking process, which are then incorporated into a second stage click model to account for position-dependent distortions. In contrast to existing methods, our approach avoids explicit propensity estimation, supports nonlinear ranking models, and can be flexibly incorporated into any state-of-the-art ranking algorithm for position bias correction. We also propose a debiasing strategy for validation clicks that enables reliable hyperparameter tuning in the absence of unbiased validation data. Empirical results show that our method outperforms state-of-the-art position bias correction methods on both benchmark and real-world industrial datasets.
- REPRPersonalised Drug Recommender Systems based on Electronic Health Records: Reproducibility and External Validity
by Archie Carpenter and Dietmar JannachRecommender systems have achieved success in delivering personalised experiences across large-scale consumer platforms, particularly in domains such as entertainment and social media. Recently, there has been increasing interest in applying recommender system methodologies to domains with positive social impact, including healthcare. Although machine learning has demonstrated strong performance across a variety of clinical tasks, comparatively less attention has been given to recommender-style decision support for personalised care. The introduction and increasing adoption of Electronic Health Records (EHRs) has enabled the development of models that leverage patient histories to generate personalised clinical recommendations. In particular, EHR-based drug recommendation aims to predict safe and effective medication combinations using diagnoses, procedures and past visit information. In this work, we examine the current state of research in this area through an analysis and reproducibility study of representative models from the literature. We show that the field has grown rapidly in recent years and increasingly relies on complex deep learning architectures. However, our study reveals a stark gap between code availability and executability: despite nearly 88\% of papers providing public code, half of the executable pipelines failed entirely. We further demonstrate that only 2 out of 41 surveyed papers are fully reproducible, and we identify major questions regarding the real-world applicability of the proposed models. By highlighting these issues, we aim to provide a clearer understanding of the field and provide a foundation for more robust and clinically meaningful research in healthcare recommender systems.
- RSCscikit-rec: A Unified, Extensible Recommendation Library with scikit-rec-agent, a Conversational Interface
by Shankar Sankararaman, Ivelin Angelov, Tin Nguyen, Jingyuan Zhang, Shivani Gowrishankar, Jaspreet Singh, Bowen Long, Qingbo Hu and Parvez AhammadWe present scikit-rec and scikit-rec-agent, two complementary open-source Python packages forming a composable, agentic recommendation framework. scikit-rec is a scikit-learn-style library built on a three-layer architecture (Recommender → Scorer → Estimator) that separates business logic, scoring strategy, and the underlying ML model, unifying five paradigms — ranking, contextual bandits, uplift modeling, sequential, and hierarchical sequential — under a single training and evaluation API. To our knowledge, it is the only library exposing IPS, SNIPS, DR, DM, and Replay-Match through a single evaluate() method shared across ranking, bandit, uplift, and sequential recommenders on the same dataset object (seven evaluators in total — six off-policy plus an on-policy baseline — and nine metrics), with implementation correctness verified by floating-point agreement against Open Bandit Pipeline on IPS, SNIPS, DR, and DM. Supported estimators span XGBoost, LightGBM, scikit-learn, and deep models (NCF, Two-Tower, NFM, DCN, SASRec, HRNN), with GPU optional. New paradigms plug in as a single Recommender subclass; we demonstrate this with Goal-Conditioned Supervised Learning (GCSL) for multi-objective recommendation, which compares directly against ranking, bandit, and uplift policies through the same evaluation API. scikit-rec-agent wraps the library as a conversational agent over eleven structured tools with a provider-agnostic LLM backend and deterministic URL-echo and import-scope checks on model output, turning plain-English data and goal descriptions into trained, evaluated pipelines without the user writing code. An internal version of the library has been in production at Intuit for 3+ years, powering 15 use cases across TurboTax and QuickBooks under sub-100 ms real-time latency contracts. Both packages are released under Apache 2.0 at https://github.com/intuit/scikit-rec and https://github.com/intuit/scikit-rec-agent, with reproducibility notebooks on MovieLens-1M and Amazon Books.
- RSCBinge Watch: Reproducible Multimodal Benchmarks Datasets for Large-Scale Movie Recommendation on MovieLens-10M and 20M
by Giuseppe Spillo, Alessandro Petruzzelli, Cataldo Musto, Marco de Gemmis, Pasquale Lops and Giovanni SemeraroAs Multimodal Recommender Systems gain interest, high-quality datasets with multimedia side information (text, images, audio, video) have become essential. However, much of the current literature relies on small-scale, undocumented, or non-public datasets. In this paper, we introduce M3L-10M and M3L-20M, two large-scale, fully documented and reproducible datasets that enrich MovieLens-10M and MovieLens-20M with multimodal features. Following a documented pipeline, we collect movie plots, posters, and trailers, extracting features via several state-of-the-art encoders. We publicly release raw data mappings, extracted features, and complete datasets to foster reproducibility and advance the field. Qualitative and quantitative analyses demonstrate our datasets’ utility across multiple perspectives. This work establishes a foundational resource for large-scale, multimodal movie recommendation. Our resource is available at: https://zenodo.org/records/18499145, with source code at https://github.com/giuspillo/M3L_10M_20M.
- INDA Framework for Value-Aligned Modeling and Robust Experimentation
by Krishna Harsha Reddy Kothapalli, Angela Liu, Kamal Nayan Reddy Challa, Pavithra Seshadrivijayakrishnan, Yongpeng Yang and Mustafa IspirWhen developing large-scale industrial recommender systems, a critical challenge is misalignment between offline model evaluation metrics and the final business objective. This disconnect often leads to inefficient development cycles where promising offline improvements fail to translate into online performance gains. This paper introduces a systematic framework to bridge this gap by applying novel value-aligned model evaluation metrics and reducing systematic errors in experiment setups at each decision-making stage. We focus on three experiment stages of the development process: offline experiments, low-traffic online A/B experiments, and high-traffic online A/B experiments. Our results from incorporating advertiser-defined value into a weighted offline objective improves Delivered Value by 4.7% on an O($B) surface and strengthens correlation between offline and low-traffic online results from a baseline of 0.1 (negligible) to now 0.65 (substantial) correlation; and our methods for reducing systematic biases in low-traffic A/B experimentation setups further improve low-traffic to high-traffic metric correlation from 0.53 to 0.93. Validated with data from over 100 large-scale offline and online experiments in both conversion (pConvs) and click (pCTR) prediction models, our framework provides a comprehensive guide for aligning machine learning development with core business objectives, leading to more accurate decision-making and significant (2x) improvements in engineering productivity.
- INDUniPinRec: Unifying Generative Retrieval and Ranking at Pinterest Scale
RecSys 2026 (Minneapolis)
- About the Conference
- Registration
- Program at Glance
- Program
- Call for Contributions
- Challenge
- Keynotes
- Accepted Contributions
- Presenter Instructions
- Workshops
- Tutorials
- Committees
- Inclusion
- Student Volunteers
- Women in RecSys
- Visa Information
- Addressing Attendance Issues
- Location / Hotel
- Lasting Impact Award
- 60 Milestones for 20 Years
About this site
This site contains information about the ACM Recommender Systems community, the annual ACM RecSys conferences, and more.
RecSys 2026
- RecSys 2027 (Hawaii)
- RecSys 2026 (Minneapolis)
- RecSys 2025 (Prague)
- RecSys 2024 (Bari)
- RecSys 2023 (Singapore)
- RecSys 2022 (Seattle)
- RecSys 2021 (Amsterdam)
- RecSys 2020 (Online)
- RecSys 2019 (Copenhagen)
- RecSys 2018 (Vancouver)
- RecSys 2017 (Como)
- RecSys 2016 (Boston)
- RecSys 2015 (Vienna)
- RecSys 2014 (Silicon Valley)
- RecSys 2013 (Hong Kong)
- RecSys 2012 (Dublin)
- RecSys 2011 (Chicago)
- RecSys 2010 (Barcelona)
- RecSys 2009 (New York)
- RecSys 2008 (Lausanne)
- RecSys 2007 (Minnesota)
- Workshops
- Submission Statistics
© 2012-
RecSys Community. All rights are reserved.




















