Session 8:

8A: Generative Recommendation: Retrieval, Ranking & Content Generation

Date: Thursday October 01, 10:30 – 12:15 CDT
Session Chair: Cataldo Musto

  • RESSAGE: Sequence-level Adaptive Gradient Evolution for Generative Recommendation
    by Yu Xie, Xingkai Ren, Qi Ying, Di Jia and Yao Hu

    Generative recommender systems hold the promise of jointly optimizing accuracy, content diversity, and creator exposure fairness. However, current reinforcement learning–based optimizers such as Gradient-Bounded Policy Optimization (GBPO) exhibit a Symmetric Conservatism failure mode: symmetric update bounds suppress learning from rare positive signals (e.g., cold-start items), static negative-sample constraints fail to prevent diversity collapse under rejection-dominated feedback, and group-normalized multi-objective rewards produce low-resolution training signals. These limitations directly harm user experience by reinforcing information cocoons and reducing new-creator visibility. We propose SAGE (Sequence-level Adaptive Gradient Evolution), a unified optimizer for list-wise generative recommendation. SAGE introduces (i) a geometric-mean importance ratio for sequence-level signal alignment, (ii) asymmetric adaptive bounding—a Positive Boost for cold-start slates and an Entropy-Aware Penalty for low-diversity failures—and (iii) a decoupled multi-objective advantage estimator. On three Amazon Product Reviews datasets and the large-scale RecIF-Bench, SAGE consistently improves top-K accuracy while delivering +89% to +101% cold-start recall recovery and +11% diversity gains relative to GBPO. Beyond-accuracy evaluation confirms that SAGE substantially reduces intra-list similarity and broadens catalog coverage, suggesting that asymmetric, sequence-aware policy optimization is an effective approach to improving both recommendation quality and content ecosystem health.

  • RESGenerating Personalized Images for Sparse-Interaction Users with Uncertainty-Aware Retrieval and Dense Knowledge Guidance
    by Yuting Zhang, Ying Sun, Dazhong Shen, Ziwei Xie, Feng Liu, Changwang Zhang, Xiang Liu, Jun Wang and Hui Xiong

    Personalized image generation aims to synthesize target images tailored to individual preferences based on users’ historical interaction data. Existing methods typically inject features from historical interaction records to guide personalized generation. However, such methods encounter two critical challenges when serving sparse-interaction users: (1) Preference Misalignment: Sparse interactions tend to lack target-semantic preference information, causing the direct injection of interaction features to misalign with users’ true target preferences. (2) Lack of Reliable Supervision: Mining preferences from sparse interactions requires sufficient supervision signals, yet personalized generation with sparse data inherently lacks direct feedback or ample preference signals for generated outputs. To this end, we propose Uncertainty-aware retrieval with Dense guidance for Sparse personalized Image Generation (UDSIG). For preference alignment, we first retrieve reference images that exhibit low-uncertainty matching with user sparse preferences from the entire dataset. For reliable supervision, we propose a dense-to-sparse scheme that incorporates a reward model derived from active users’ dense interaction data to drive personalized generation in sparse scenarios. Extensive experiments and human evaluations across three public datasets confirm the superiority of our model, along with its strong generalization to dense interaction scenarios.

  • RESCoarse-to-Fine Long-term Interest Modeling for Generative Recommendation
    by Shiteng Cao, Junda She, Ji Liu, Bin Zeng, Chengcheng Guo, Kuo Cai, Qiang Luo, Ruiming Tang, Han Li, Kun Gai, Zhiheng Li and Cheng Yang

    Leveraging long-term user behavioral patterns is a key trajectory for enhancing the accuracy of modern recommender systems. Due to the quadratic complexity of attention mechanisms, existing GR models are typically confined to short interaction sequences. While pioneer works have attempted to adapt Search-based Interest Models (SIM) to the generative context, they typically overlook the inherent hierarchical distinction of SIDs. GR is fundamentally a coarse-to-fine generation task, where the initial SIDs (prefix) determine the broad semantic category and the subsequent SIDs (suffix) pinpoint the specific item. Thus, our core insight is that the prefix and suffix of SIDs require distinct long-term signal injections. To bridge this gap, we propose GLASS, a Generative recommendation framework that integrates Long-term user interests into the generative process via SIDTier and Semantic Search. For the generation of SID prefix, we introduce SID-Tier, a module that maps long-term interactions into a unified interest vector to enhance the prediction of the initial SID token. SID-Tier leverages the compact nature of the semantic codebook to incorporate cross features between the user’s long-term history and candidate semantic codes. Furthermore, for the generation of SID suffix, we present semantic hard search, which utilizes generated coarse-grained semantic ID as dynamic keys to extract relevant historical behaviors, which are then fused via an adaptive gated fusion module to recalibrate the trajectory of subsequent fine-grained tokens. Extensive experiments on two large-scale real-world datasets, TAOBAO-MM and KuaiRec, demonstrate that method outperforms state-of-the-art baselines. A two-week online A/B test on a short-video platform demonstrate that GLASS achieves significant gains in recommendation quality. Our codes are publicly available at this anonymous link to facilitate further research in generative recommendation.

  • RESUniRec: A Unified Expressive-Aligned Generative Recommendation Framework for E-commerce
    by Ziliang Wang, Gaoyun Lin, Xuesi Wang, Shaoqiang Liang, Yili Huang, Weijie Bian, Li Zhang, Mingchen Cai, Jian Dong and Guanxing Zhang

    Traditional discriminative recommendation pipelines suffer from objective misalignment and error propagation across stages, motivating a shift toward generative recommendation (GR). However, existing GR methods decode over compact Semantic ID (SID) tokens without access to item-side features, lacking the explicit user–item feature crossing that discriminative models rely on. Combined with the inherent one-to-many nature of recommendation, this absence of item-side signals significantly amplifies generation uncertainty, making the generative paradigm widely regarded as having a lower performance ceiling than its discriminative counterpart. We propose UniRec, a unified expressive-aligned GR framework that unifies the multi-stage pipeline into a single generative model and aligns its expressive power with the discriminative counterpart. Via Bayes’ theorem, we show that any practical gap stems from feature coverage rather than modeling asymmetry, motivating Chain-of-Attribute (CoA), an expressive-alignment mechanism that pre-generates item attributes before decoding SIDs, recovering item-side feature crossing and yielding measurable per-step entropy reduction. Beyond CoA, Capacity-constrained SID enforces exposure-weighted load balancing to suppress token collapse, and Conditional Decoding Context (CDC) injects scenario-conditioned signals to stabilize multi-scenario decoding and Cartesian-product-based structured summaries of generated tokens to reinforce conditional dependence across decoding layers. A joint Reward-Driven Fine-tuning (RFT) and Direct Preference Optimization (DPO) framework further aligns the model with business objectives. Deployed on a large-scale e-commerce platform, online A/B tests confirm significant gains in page-view click-through rate (PVCTR, +5.37%), orders (+4.76%), and gross merchandise volume (GMV, +5.60%).

  • INDTokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in Recommendation Systems
    by Baolei Li, Yiping Yuan, Yilin Zheng, Likang Yin, Ling Liu, Fabio Soldo, Romer Rosales, Xinyang Yi and Lichan Hong

    Large-scale recommendation systems face “Memory Wall” bottlenecks due to massive, dense embedding tables. While generative retrieval uses discrete tokens for IDs, high-dimensional context still relies on inefficient dense formats. Inspired by computer vision data compression, we propose Dual-purpose Semantic IDs to achieve LLM-level I/O efficiency. Our methodology uses hierarchical quantization to condense continuous embeddings into discrete Semantic IDs performing two concurrent roles: (1) Collaborative Identity: modeling user-item interactions via learnable embedding table; and (2) Content Reconstruction: using a lightweight Semantic Decoder for on-the-fly embedding approximation. This approach replaces massive vector storage with on-demand reconstruction, reducing system overhead and data footprints. We demonstrate the efficacy of our framework through offline evaluations and successful online deployment in production-scale ranking and retrieval systems at a major video sharing platform, showing that discrete tokens are indeed all you need for highly efficient, content-rich recommendation.

  • INDGenPage: Towards End-to-End Generative Homepage Construction at Netflix
    by Lequn Wang, Jiangwei Pan and Linas Baltrunas

    We present GenPage, an end-to-end generative approach to Netflix homepage construction that replaces the traditional multi-stage recommender stack with a single transformer. GenPage represents the user and request context as a prompt, and autoregressively generates the entire structured, multi-row homepage as the response. We adapt the LLM training recipe: pretraining on positively engaged production pages, followed by post-training via weighted binary classification (WBC) or reinforcement learning (RL). For industry-scale deployment, we introduce techniques addressing cold start, model freshness, business-rule enforcement, and serving efficiency. In online A/B tests against a mature, highly optimized production homepage recommender, the WBC variant of GenPage delivered a +0.24% lift on the core engagement metric we use for launch decisions (p < 0.001), while reducing end-to-end serving latency by 20%. Offline experiments yield two findings worth highlighting: enriching the prompt yields a larger improvement than scaling model capacity in our current regime, and RL post-training increases homepage diversity even though diversity is not part of the objective.

  • INDUniPinRec: Unifying Generative Retrieval and Ranking at Pinterest Scale
    by Hanyu Li, Yi-Ping Hsu, Aditya Mantha, Prabhat Agarwal, Laksh Bhasin, Jialu Wang, Hongtao Lin, Bella Huang, Yaxin Li, Xinyi Li, Chuxi Wang, Kousik Rajesh, Hooshmand Shokri Razaghi, Shunyao Li, Zongyue Qin, Jaewon Yang, James Li, Dhruvil Deven Badani, Jiajing Xu and Charles Rosenberg

    Modern recommendation systems predominantly train retrieval and ranking as separate models despite both increasingly relying on large transformers encoding the same user behavior data, duplicating parameters, compute, and serving cost. Prior work unifies the model architecture but not the full pipeline: input formats, training procedures, and serving stacks remain fragmented across stages. We present UniPinRec, which achieves full-stack unification of retrieval and ranking at Pinterest: one input format, one model, one training stage, deployed within existing serving infrastructure. A shared transformer encodes the user action sequence into candidate-independent representations that branch into retrieval (ANN dot-product) and ranking (cross-attention) via task-specific heads. Three ideas make this work: (1) Masked Action Modeling (MAM) eliminates interleaving, enabling weight sharing without doubling context length; (2) Blended training examples pair action sequences with feedview impression slates to satisfy both objectives jointly; (3) Cross-stage KV cache sharing reuses user-history computation from retrieval for ranking, reducing total FLOPs versus serving two independent models. Deployed in the Pinterest core surfaces, UniPinRec delivers approximately +1% online engagement lift while cutting end-to-end serving latency by 11.1% and lifting QPS by 63.6%. To our knowledge, this is the first full-stack unification of retrieval and ranking, covering inputs, model, training and serving, deployed in a production recommendation system.

Back to program