Day 2 Posters: Industry + R&P Notes + Short Papers
Date: Wednesday September 30
Posters are available from 10 am to 4 pm, and are presented by the authors during the coffee breaks.
Industry
- Spot A1RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation
by Renzhi Wu, Zikun Cui, Junjie Yang, Tai Guo, Hong Li, Xian Chen, Li Yu, Ke Pan, Sri Reddy, Mahesh Srinivasan, Nipun Mathur, Haomin Yu and Hong YanGraph-based retrieval at billion-node scale requires jointly solving three tightly coupled problems—graph construction, representation learning, and real-time serving—yet existing work addresses each in isolation. We present RankGraph-2, a framework deployed at Meta that co-designs all three lifecycle stages for similarity-based retrieval (U2U2I and U2I2I), where each stage’s requirements shape the others. Serving requires a co-learned cluster index to avoid expensive online KNN—this pushes index co-training into the training objective. Training benefits from the observation that similarity-based retrieval tolerates pre-computed neighborhoods, eliminating online graph infrastructure—this requires construction to produce self-contained data. Construction must also support hour-level refresh for item coverage. Acting on these cascading requirements, RankGraph-2 reduces hundreds of trillions of edges to hundreds of billions via subsampling with popularity bias correction, pre-computes multi-hop neighborhoods via personalized PageRank, and co-learns a residual-quantization cluster index that reduces serving computational cost by 83\%. This lifecycle co-design enables a simple architecture to achieve 3.8$\times$ higher recall than a GAT + Deep Graph Infomax model on a bipartite graph and 2.1$\times$ higher than PyTorch-BigGraph on item retrieval. RankGraph-2 delivers up to +0.96\% CTR and +2.75\% CVR, and has powered \textbf{20+ retrieval launches} across major surfaces.
- Spot A2PROMISE: Process Reward Models for Unlocking Test-Time Scaling Laws in Generative Recommendations
by Chengcheng Guo, Kuo Cai, Yu Zhou, Qiang Luo, Ruiming Tang, Han Li, Kun Gai and Guorui ZhouGenerative Recommendation has emerged as a promising paradigm, reformulating recommendation as a sequence-to-sequence generation task over hierarchical Semantic IDs. However, current approaches face a severe challenge that we define as Semantic Drift, where errors in early, high-level tokens irreversibly divert the generation trajectory into irrelevant semantic subspaces. Inspired by Process Reward Models (PRMs) that enhance reasoning in Large Language Models, we propose Promise, a novel framework that integrates dense, step-by-step verification into generative models. Our framework utilizes a lightweight PRM to assess the quality of intermediate inference steps, and a PRM-guided beam search strategy that leverages dense feedback to dynamically prune erroneous branches. Most importantly, this method unlocks Test-Time Scaling Laws in recommender systems, demonstrating that by increasing inference compute, smaller models can match or surpass larger models. Extensive offline experiments and online A/B tests on a large-scale platform demonstrate that Promise, effectively mitigates Semantic Drift, significantly improving recommendation accuracy while enabling efficient deployment.
- Spot B1POEM: Partial-Order Enhanced Real-Time Sequential Modeling for Recommendation
by Linxiao Che, Sun Yijia, Siyuan Lou, Shanshan Huang, Qiang Luo, Ruiming Tang, Han Li and Kun GaiAbstract Real-time recommendation systems face the challenge of dynamically changing user interests and contextual environments. Traditional sequential recommendation models rely on static historical click sequences, which struggle to capture real-time preference shifts and ignore the structured information embedded in the system’s internal ranking logic. This paper proposes POEM (Partial-Order Enhanced Modeling), a novel real-time sequential modeling framework that leverages the partial-order relations inherent in the recommendation cascade. POEM innovatively utilizes the real-time multi-task ranking scores (e.g., predicted CTR, watch time) from the preceding ranking stage as supervisory signals to construct dynamic partial-order sequences, thereby achieving fine-grained real-time interest modeling and aligning system objectives with user behavior. Our contributions are threefold: a partial-order guided sequence construction paradigm that augments traditional temporal sequences with a dynamically grouped and sampled sequence based on real-time ranking scores, enabling per-request interest reassessment; (2) a multi-objective score fusion mechanism that integrates various ranking signals through normalized rank-weighting into a unified quintuple representation; and (3) a hierarchical sample learning strategy combining system-preferred items (top-ranked) and user feedback (e.g., long-play videos) as positive samples, enhanced by graph-retrieved hard negative samples and a margin-based pairwise loss. Deployed in Kuaishou’s platform, POEM achieves significant online gains: +0.249% and +0.213% in average viewing time per user on the KS Single Page and KS Lite Page, respectively. Extensive ablation studies validate the effectiveness of each component, demonstrating POEM’s superiority in real-time responsiveness, recommendation accuracy, and content diversity.
- Spot B2Space Efficient Item Embeddings for Large Recommender Systems
by Duy Nguyen, Devanshu Jain, Mike Lawrence, Steffen Rendle, Kun Su and Anushya SubbiahEmbedding-based methods are central to modern recommender systems, but storing item embeddings at scale poses a significant memory challenge, particularly with catalogs of millions or billions of items. Common solutions, such as uniformly reducing embedding dimensions or truncating the vocabulary to popular items, can degrade recommendation quality, limit item discovery, and exacerbate popularity bias, especially harming long-tail item visibility. This paper introduces a memory-efficient hybrid embedding scheme that allocates representation capacity based on item popularity. We represent popular items with high-dimensional embeddings to capture detailed user-item interactions, while less popular items use a compressed representation. This approach allows for full vocabulary coverage within a strict memory budget. Experiments on public datasets and large-scale proprietary data demonstrate that our method significantly improves recommendation quality, particularly for long-tail items, compared to baselines with equivalent memory footprints. Our models achieve better coverage without sacrificing overall accuracy, offering a practical solution for building scalable and diverse recommender systems.
- Spot C1Structuring and Tokenizing Distributed User Interest Context for Generative Recommendation
by Ruizhong Qiu, Yinglong Xia, Dongqi Fu, Hanqing Zeng, Ren Chen, Xiangjun Fan, Hong Li, Hong Yan and Hanghang TongGenerative recommendation is an emerging paradigm that has shown promise in industrial recommendation systems, aiming to predict the next interactions of users based on their historical behavior in an autoregressive manner. At the core of generative recommendation is item tokenization to bridge item semantics and the recommendation model. However, existing methods often struggle to effectively organize and insert complex user-behavioral and item-semantic contexts simultaneously into the recommendation model: On the one hand, existing graph-based integration methods (graph serialization, graph neural networks, etc.) either suffer from scalability issues or only leverage local graph information; On the other hand, existing semantic tokenization methods typically rely on heuristics and lack a supervision signal, which may not guarantee accurate semantic representations. To address these critical limitations in user interest context modeling, we fundamentally propose G2Rec, i.e., a scalable framework that bridges holistic graph-based user co-engagement interest modeling and semantic tokenization to boost the industry-scale generative recommendation. First, we propose to construct a sparsified item-item co-engagement graph of size O(MlogM) as the item schema, where M is the total number of interactions. Second, we design a scalable ”soft” clustering algorithm with time complexity O($\rho$MlogM) per iteration to extract the distributed interest prototypes from the constructed graph, where $\rho$ is a small constant representing the sparsity of the soft cluster membership distribution (but not a hard-assigned one-to-one exclusive membership). Third, using the item interest profiles extracted from the soft clustering, we tokenize them together with the user’s interested items to train the generative sequential recommendation model. In all, our scalable framework, G2Rec, enables the recommendation model to capture holistic and semantic user interest prototypes without requiring ground-truth interests of users, providing a more comprehensive and accurate modeling of user behavior contexts in industrial sequential recommendation. Online deployment on product surfaces and extensive experiments on public datasets demonstrate the superiority of G2Rec over existing methods.
- Spot C2UniShare: A Unified Framework for Joint Video and Receiver Recommendation in Social Sharing
by Caimeng Wang, Li Chong, Dongxu Liu, Xu Min and Jianhui BuSharing behavior on short-video platforms constitutes a complex ternary interaction among the user (sharer), the video (content), and the receiver. Traditional industrial solutions often decouple this into two independent tasks: video recommendation (predicting share probability) and receiver recommendation (predicting whom to share with), leading to suboptimal performance due to isolated modeling and inadequate information utilization. To address this, we propose \textbf{UniShare}, a novel unified framework for joint sharing prediction on both video and receiver recommendation. UniShare models the share probability through an enhanced representation learning module that incorporates pre-trained GNN and multi-modal embeddings, alongside explicit bilateral interest and relationship matching. A key innovation is our joint training paradigm, which leverages signals from both tasks to mutually enhance each other, mitigating data sparsity and improving bilateral satisfaction. We also introduce \textbf{K-Share}, a large-scale real-world dataset constructed from Kuaishou platform logs to support research in this domain. Extensive offline experiments demonstrate that UniShare significantly outperforms strong baselines on both tasks. Furthermore, online A/B testing on the Kuaishou platform confirms its effectiveness, achieving significant improvements in key metrics including the number of shares (+1.95%) and receiver reply rate (+0.482%).
- Spot E1UniTraj: Cross-Domain Long-Sequence Modeling for Commercial Recommendation
by Xian Hu, Ming Yue, Zhixiang Feng, Junwei Pan, Junjie Zhai, Ximei Wang, Xinrui Miao, Qian Li, Xun Liu, Shangyu Zhang, Letian Wang, Hua Lu, Zijian Zeng, Chen Cai, Wei Wang, Fei Xiong, Pengfei Xiong, Jintao Zhang, Zhiyuan Wu, Chunhui Zhang, Anan Liu, Jiulong You, Chao Deng, Yuekui Yang, Shudong Huang, Dapeng Liu, Haijie Gu and Jie JiangLong-sequence modeling is increasingly important in recommender systems for capturing users’ evolving and long-term interests. In advertising, however, user interaction histories are often highly sparse due to limited exposure opportunities, making ad-only behavior sequences insufficient for effective long-sequence recommendation. To address this limitation, we propose UniTraj, a practical framework that extends sequence construction beyond the advertising domain by incorporating behaviors from content-consumption scenarios, forming unified commercial trajectories across domains and scenarios. Such unified trajectories provide richer behavioral context, but also introduce substantial heterogeneity in feature taxonomy, behavioral semantics, and optimization targets. In particular, they raise three key challenges: interference among fields from different domains and scenarios, target-specific conflicts in temporal and semantic patterns, and complex high-order dependencies across heterogeneous behavioral signals. To tackle these challenges, UniTraj adopts a two-stage design. In the first stage, it combines hierarchical hard search with a decoupled embedding-based soft search module to retrieve relevant behaviors under complex feature hierarchies while reducing conflicts between retrieval and representation learning. In the second stage, it introduces several decoupled sequence modeling components, including Decoupled Side Information Temporal Interest Networks for mitigating cross-field interference, target-decoupled positional encoding and target-decoupled SASRec for capturing target-aware temporal dynamics, and Deep TIN for modeling high-order behavioral correlations. We deploy UniTraj in a large-scale online advertising system and observe consistent improvements in online business metrics across multiple commercial scenarios. The results demonstrate the effectiveness of unified cross-domain behavior modeling for long-sequence recommendation in sparse advertising environments.
- Spot E2UNIQUE: A Unified Retrieval and Ranking System for Large-Scale Feed Recommendation
by Zhuang Liu, Yongkang Fu, Zuodong Yang, Zonggang Wu, Yuqi Lu, Shouke Qin, Shantao Li, Guangxing Chen and Maolin WangIndustrial mobile feed systems rely on a retrieval-ranking pipeline to serve large-scale, heterogeneous, and fast-changing content under strict latency constraints. However, existing pipelines still suffer from two critical issues: hierarchical quantization instability in candidate retrieval and information loss between separated retrieval and ranking stages. These issues hurt long-tail and cold-start recommendation and complicate efficient serving. To address them, we present UNIQUE, a unified retrieval and ranking recommendation framework with single-layer flat quantization. UNIQUE integrates generative code-based retrieval and target-aware ranking into one early-fusion architecture, enabling end-to-end training under a shared representation while preserving efficient candidate generation. A balanced quantization mechanism is further introduced to mitigate codebook imbalance and improve long-tail representation. Offline experiments evaluate UNIQUE from both retrieval and ranking perspectives, while codebook analysis shows more balanced resource allocation than hierarchical quantization. We deploy UNIQUE in the homepage feed, discovery-page, and short-video recommendation scenarios of Mobile Baidu, serving large-scale real-world traffic. Online A/B tests achieve a 0.96% gain in total watch duration and a 1.08% gain in total distribution volume, with notable improvements for new users and highly active users. Serving measurements show 89 ms P99 latency and 44.23% online inference MFU. These results show that UNIQUE provides a stable, efficient, and production-ready framework for unified retrieval and ranking in industrial recommendation.
- Spot F1The Text on the Creative: An Under-Exploited Ranking Modality for Short-Form Video Ads
by Shubham Goel, Angli Liu, Antoine Simoulin, Himanshu Thakur, Qin Huang, Benjamin Au, Guy Lebanon, Sagar Chordia, Andrew Treadway, Wendy Jiang, Yiding Wen, Jianing Fu, Alan Li, Haibo Zhang, Selahattin Akkas, Calvin Ma, Zhen Zeng and Yunyu HeVideo advertisements contain rich textual signal overlaid on the creative itself – promotional copy, prices, calls-to-action, brand and product mentions — that current production ranking systems largely under-exploit. We argue this is a high-yield, low-cost ranking modality and describe a production deployment on Instagram’s ads ranking model at billion-ad daily scale. Making this practical requires three components: a specialized OCR backbone delivering near-VLM word quality at ~80× the throughput; a smart frame-extraction stage that recovers more OCR-relevant frames at half the per-video frame budget; and a feature stack, which consists of TF–IDF tokens plus 3-byte RQ-VAE semantic IDs over a domain-tuned sentence encoder, that integrates into the ranker’s existing sparse embedding tables with no dedicated dense tower. The launched bundle delivers 0.055–0.06% offline relative NE gain (≈0.12% on Reels-class surfaces) and 0.06–0.08% NE gain in a multi-week 80%-traffic online A/B.
- Spot F2Versioned Late Materialization for Ultra-Long Sequence Training in Recommendation Systems at Scale
by Guo Liang, Ge Song, Litao Deng, Jianhui Sun, Chufeng Hu, Lu Zhang, Zhen Ma, Shouwei Chen, Weiran Liu, Sarang Sreeshylan and Xiaoxuan MengModern Deep Learning Recommendation Models (DLRMs) follow scaling laws with sequence length, driving the frontier toward ultra-long User Interaction History (UIH). However, the industry-standard “Fat Row” paradigm, which pre-materializes these sequences into every training example, creates a storage and I/O wall where data infrastructure usage exceeds GPU training capacity due to data redundancy that is amplified in multi-tenant environments where models with vastly different sequence length requirements share a union dataset. We present a versioned late materialization paradigm that eliminates this redundancy by storing UIH once in a normalized, immutable tier and reconstructing sequences just-in-time during training via lightweight versioned pointers. The system ensures Online-to-Offline (O2O) consistency through a bifurcated protocol that prevents future leakage across both streaming and batch training, while a read-optimized immutable storage layer provides multi-dimensional projection pushdown for heterogeneous model tenants. Disaggregated data preprocessing with pipelined I/O prefetching and data-affinity optimizations masks the latency of training-time sequence reconstruction, keeping training throughput compute-bound by GPUs. Deployed on production DLRMs, the system reduces training data infrastructure resource usage while enabling aggressive sequence length scaling that delivers significant model quality gains, serving as the foundational data infrastructure for modern recommendation model architectures, including HSTU and ULTRA-HSTU.
- Spot G1MoR: An Adaptive Retrieval Allocation Balancing Long-Term Interest and Short-Term Evidence
by Yan Fu, Roni Peled, Brian Biesel, Werner Mostert, Raghuram Krishnaswami, Fei Wang and Qilin QiModern recommender systems often rely on multiple retrieval sources to generate candidate items, yet determining optimal quota allocation across these sources remains challenging. We present Mixture of Retrieval (MoR) bandit, a novel framework that dynamically optimizes retrieval allocation by balancing long-term user interests with short-term behavioral signals. Unlike traditional static approaches, our method employs a modified Thompson Sampling algorithm that combines a user’s affinity scores to third-party video channels with their recent behavioral signals, using exponential decay to prevent over-reliance on stale signals. The framework enhances personalization while promoting exploration of potentially valuable but under-exposed content channels. In a large-scale A/B test on Prime Video, MoR bandit achieved a significant 4% increase in third-party channel subscriptions and substantial improvements in user engagement metrics. At the time of writing this paper, MoR has been fully launched in production worldwide serving hundreds of millions of customers. Our approach provides a generalizable solution for multi-source retrieval optimization in large-scale recommender systems.
- Spot G2DRanker: A Transformer-Based Multi-Task Ranking Model for the Disney+ Homepage
by Priya Nirmal Singh Khokher, Daniel Nemirovsky, Benjamin Whitesell, Ethan Sukrae Lee and Eric WulfWe present DRanker (Disney+ Ranker), a transformer-based multitask ranking model serving personalized recommendations on the Disney+ homepage. We describe the encoder-only architecture over enriched engagement-history sequences paired with a lightweight causal debiasing mechanism, and the inference stack that enables transformer-based ranking at scale. Through iterative development, we systematically evaluated more complex alternatives and found that simpler designs consistently outperformed them on offline metrics, qualitative behavioral evaluation, and in production. Deployed globally, DRanker drives over 1% lift in our key engagement metric and 0.5% lift in daily active users across all experiences powered by the ranker. We share practical lessons on data engineering, debiasing, and architectural choices that shaped our ranking system at scale.
- Spot H1Generative Spatiotemporal Intent Sequence Recommendation via Implicit Reasoning in Amap
by Sicong Wang, Ruiting Dong, Yue Liu, Bowen Zheng, Jun Meng, Jie Li, Shuaijun Guo, Yu Gu, Fanyi Di and Xin LiReal-world user behaviors are rarely isolated atomic actions but exhibit intent flows with spatiotemporal causal chains. To provide holistic service schemes, we focus on the task of Generative Spatiotemporal Intent Sequence Recommendation (GSISR), which aims to generate intent sequences that are both logically coherent and physically executable within complex spatiotemporal contexts. While Large Language Models (LLMs) offer strong reasoning potential for GSISR, their industrial deployment is hindered by significant inference latency and spatiotemporal hallucinations. To bridge this gap, we propose a generative framework GPlan that internalizes LLM reasoning into lightweight models through two key innovations. First, to achieve reasoning under strict latency constraints, we introduce Progressive Implicit CoT Distillation, which compresses explicit reasoning processes into special think tokens, allowing small models to inherit complex planning logic without generating long reasoning text. Second, to address the disconnect between general knowledge and real-world constraints, we design Spatiotemporal Counterfactual DPO. By aligning the model with counterfactual scenarios, we improve its responsiveness to spatiotemporal context and reduce context-mismatched plans. Offline experiments and online A/B testing demonstrate that our approach improves sequence coherence and context responsiveness. Our implementation and the anonymized GSISR dataset are available at https://github.com/alibaba/GPlan.
- Spot H2From Placements to Pages: Modeling Recommendation Module Interactions with Sequential Transformer Neural Bandits
by Xingming Qu, Peijun Zou, Jing Shao, Xin Yin, Albert Hsiung, Kathy Hu, Nitish Dhinaharan, Dan Schonfeld, Kevin Liu, Shawn Zhou, Binbin Li, Adam Ilardi, Jesse Lute and Julie ChengModern recommender systems often populate multiple placements on a page using independent bandit or ranking models, implicitly assuming that placement decisions are independent. This assumption overlooks interactions among recommendation modules and can lead to suboptimal page-level outcomes. We reframe multi-placement recommendation as a sequential decision problem in which each module selection is conditioned on previously selected modules. We propose the Sequential Transformer Neural Bandit (STNB), a production framework that learns sequence-aware page representations for contextual bandit optimization. For latency-aware deployment, STNB uses a decoupled serving architecture with configurable autoregressive truncation, applying strict sequential inference to high-visibility placements and serving the remaining placements with a single batch scoring pass. STNB further incorporates a multi-objective bandit layer to align serving-time decisions with both engagement and downstream commerce outcomes. Offline evaluation shows improved predictive quality over an independent-placement neural contextual bandit baseline, and large-scale online A/B tests on eBay View Item surfaces demonstrate consistent gains in promoted-listings revenue while preserving broader marketplace outcomes. Ablation and latency analyses show that multi-objective optimization and truncated sequential serving are important for real-world deployment.
- Spot I1Multi-Objective Ranking for Live-Streaming: Balancing Fresh and Delayed Signals with Segment-Aware Targeting
by Xiaoyi Gu, Julia Tavares, Eder Santana, Carlos Mendoza-Cardenas, Nikita Mishra and Saad AliOne of the most challenging problems entertainment live-streaming services face in recommendation systems is that user behaviors are sparse and delayed, and interaction data exhibits bias for different user segments. Unlike e-commerce applications where user actions follow linear sequences, live-streaming viewers engage in multiple concurrent behaviors of watching, chatting, following, and spending, each occurring with varying delays. We address these challenges through three key contributions: 1) a delayed window approach that extends feedback collection beyond immediate responses, 2) a multi-model architecture that combines fresh and delayed signals, and a segment-aware targeting module that optimizes ranking scores differently across user lifecycle stages, and 3) Multi-gate Mixture-of-Experts (MMoE) integration that jointly models correlated targets while reducing model parameters by 41.9% compared to independent models. Online A/B testing demonstrates significant improvements, including a +0.09% increase in Daily Active Viewers (DAV), generating millions more annual active viewer days, and +0.56% increase in highly engaged viewers’ capped Average Revenue Per User (ARPU). Viewer-segment targeting achieved an additional +0.15% DAV improvement for newer and less engaged viewers, while MMoE enhancement added +0.08% overall DAV and +0.27% new follows. The experimental system processes ranking requests with low latency, providing a scalable approach for balancing multiple business objectives across diverse user populations. In addition, we tested the multi-model architecture on the mobile livefeed application and achieved a +1.12% increase in positive user-channel interactions (clicks, follows, and likes), demonstrating applicability beyond the primary use case.
- Spot I2RECAP: Feedback-Driven Streaming Semantic User Profiles for Short-Video Recommendation
by Ziyi Zhao, Xiaoyou Zhou, Xiao Lv, Yangyang Li, Chubo He, Zhao Liu, Jiayao Shen, Yuqi Liu, He Li, Chengyi Zhang, Jian Liang, Ming Li, Chongming Gao, Fuli Feng, Ruiming Tang and Han LiLanguage-based user profiles convert long behavioral histories into explicit semantic representations for recommendation. However, most profile generators are optimized in an open loop: they may summarize past behavior fluently, but are not directly trained to improve future recommendation. We study this problem in real-world short-video recommendation, where user behaviors continuously arrive as streams and profiles must be incrementally updated under limited capacity. This requires maintaining a consistent bounded profile state and constructing profile-targeted semantic feedback from industrial implicit behavior logs. We propose RECAP, an offline closed-loop framework for optimizing streaming structured semantic profiles with historical recommendation feedback. RECAP maintains each profile as a bounded structured memory by combining LLM-based semantic updates with deterministic lifecycle and capacity control. RECAP constructs profile-targeted semantic feedback by filtering label-consistent behavior pairs with an LLM judge and training a dual-tower evaluator whose matching score serves as a GRPO reward. Experiments on Kuaishou short-video data show that RECAP improves uAUC by 0.0084 and Recall@2000 by about 4.9% over the base generator. Further analyses confirm the benefits of feedback construction and policy optimization, and show more grounded refinement and user-level abstraction in profile updates.
- Spot O1NextGen: A Multi-Objective Generative Re-ranking Framework for Taobao Recommendation
by Longxiang Xiong, Chaoqun Hou, Zihao Zhu, Cheng Guo, Tong Liu and Bo ZhengReranking is a critical stage in e-commerce recommendation that reorders items to optimize the final list as a whole. Existing genera- tive reranking methods suffer from two limitations: (1) they rely on next-item exposure prediction as supervision, neglecting page- level multi-objective signals such as total clicks or transactions; (2) autoregressive decoding incurs prohibitive latency, making full- candidate scoring infeasible under real-time constraints. We pro- pose Next-Gen, a generative reranking framework for page-aware multi-objective recommendation under strict latency budgets. Next- Gen introduces three key designs: (1) a full-candidate autoregressive generative framework via a global context-aware Encoder-Decoder architecture that jointly optimizes page-level multiple objectives while autoregressively selecting items across the entire candidate set, rigorously constraining inference latency for real-time deploy- ment; (2) a residual connection scheme for incremental reranking that feeds upstream ranking scores into both the item embedding layer and the candidate selection module, so the model only needs to predict nonlinear gains relative to ranking scores, improving con- vergence stability; (3) LLM-enhanced semantic embeddings from Taobao’s multimodal foundation model that enrich item represen- tations with domain knowledge at no additional inference cost. Online A/B tests on Taobao Miaosha show that Next-Gen achieves +18.85% GMV and +12.35% order volume over the strongest baseline. The system is deployed in production, serving hundreds of millions of users daily.
- Spot O2Melo: A Production LLM-Powered Music Recommendation Agent
by Shijia Wang, Da Guo, Qiang Xiao, Fanghui Bi, Weisheng Li, Dongjing Wang and Chuanjiang LuoWe describe Melo, an LLM-powered music recommendation agent deployed on NetEase Cloud Music. Melo is structured as a deterministic five-node state graph over heterogeneous tools, with a prompt- and state-machine-driven orchestration policy rather than a fine-tuned controller. Two production failure modes drove the design: entity hallucination, where the agent commits to interpretations unsupported by the live catalog or user-behavior index, and long-tail degradation, where over-constrained requests collapse to generic popular fallbacks. We address them with two complementary mechanisms. Inference-time entity grounding repurposes the production search index as a verification primitive that gates entity decisions before they propagate downstream. Reflective retry verbalizes failure reasons from a broken tool chain and feeds them into the next planning step, so the system can relax or revise constraints rather than fall back blindly. A one-month online A/B test within NetEase Cloud Music’s playlist business unit reports an over 2 pp absolute lift in a primary playlist retention metric and a lift of over one minute in a core playlist engagement metric; offline ablation isolates a 7.8 pp absolute reduction in entity misidentification from the three-layer grounding stack on our evaluation set, and production triggered-session analysis shows reflective retry firing on 5.8% of sessions with 59% processlevel recovery (non-empty, error-free); constraint preservation is reported as a separate open evaluation question. The takeaway from this deployment is that, at production scale, progress on LLMpowered music recommendation hinges less on stronger language understanding than on runtime mechanisms that detect and recover from its mistakes.
- Spot P1LLM-Based User Personas for Recommendations at Scale
by Haoting Wang, Haokai Lu, Zheyun Feng, Jenny Huang, Yifat Amir, Gregory Hinkson, Ben Most, Zelong Zhao, Yixin Kelly Cui, Rein Zhang, Fabio Soldo, Yu Xia, Nihar Bhupalam, Minmin Chen, Konstantina Christakopoulou, Lichan Hong and Ed H. ChiLarge Language Models (LLMs) offer unprecedented potential for enhancing recommendation systems through their world knowledge and reasoning capabilities. However, existing approaches often rely on structured IDs or offline processing, limiting semantic richness, real-time adaptability, and user-facing interpretability. In this paper, we introduce a novel framework that enables real-time generation of LLM-based user interest personas for a large-scale commercial video recommendation platform. Our method generates natural-language user interest personas that address the exploitation-exploration trade-off by combining the summarization of existing interests with novel topics, directly during serving. To overcome the computational challenges of online LLM inference at a billion-user scale, we design a cost-efficient architecture leveraging knowledge distillation, asynchronous inference, and input optimization via semantically clustered video representations. Extensive offline evaluations, user studies, and live A/B tests demonstrate significant improvements in viewer value. This work bridges the gap between high-level semantic understanding and industrial-scale recommendation, paving the way for more dynamic, explainable, and satisfying personalized experiences.
- Spot P2PLAIN: An Explainable Generative Search System Enhanced by Multi-granularity Semantic Alignment
by Guoliang Zhang, Weifan Wang, Junyao Zhao, Zhuo Li, Hongjing Zhang, Xiaobo Guo, Zhixin Zhai, Yonghui Zhao, Zhihao Wang, Jiayang Liu, Yingjie Cui, Jiwei Tan and Xuanping LiIndustrial search platforms must efficiently retrieve relevant items from billions of candidates while satisfying both query relevance and user preferences. Generative Search (GS) has emerged as a transformative paradigm that reformulates traditional indexing and matching as an autoregressive generation task. However, most existing generative search models suffer from two critical deficiencies: (1) generated Semantic IDs (SIDs) often lack explicit semantic correspondences, undermining codebook interpretability; (2) unstructured codebooks impose a fully-connected search space, forcing the generator to navigate a highly entangled decoding path. To address these limitations, we propose PLAIN, which integrates Multi-stage Codebook Construction (MCC) and Unified Generative Retrieval (UGR). MCC leverages LLM-generated taxonomy and metadata labels, applying hard assignment for closed-set taxonomy levels and soft assignment for open-set metadata levels, transforming unstructured codebooks into interpretable hierarchical topic paths. UGR operationalizes the MCC schema by employing Symmetric Context Encoders (SCE) that align both query and item representations to the structured label space via knowledge distillation and semi-supervised hierarchical quantization, enabling consistent end-to-end generative retrieval. Extensive experiments and online A/B testing in Kuaishou’s live search system demonstrate significant improvements in user engagement and content consumption.
- Spot V1Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops
by Enrico Palumbo, Alexandre Tamborrino, Victor Ode, Ben Lacker, Adrià Casas Escoda, Jeremy Hopple, Marcus Better, James Leoni, Hugo Galväo, Hugues Bouchard, Mounia Lalmas, Jose Luis Redondo-García, Abenezer Abebe, Ann Clifton, Anton Blomberg, Henrik Lindström, Dani Doro and Christine Doig CardetConversational recommendation agents are emerging as a new paradigm for content discovery, enabling users to express complex intents through natural language (e.g., “recommend Italian indie artists I haven’t heard before” or “explain why they fit my taste”). A central challenge in building such agents is optimizing agent planning, i.e., deciding how to select, sequence, and invoke tools. This challenge is particularly acute in cold-start settings, where real user interactions are not yet available. We introduce a pipeline for multi-turn synthetic data generation and a self-improvement loop to address the lack of interaction data and the difficulty of optimizing agent planning in cold-start set- tings. The synthetic data pipeline transforms single-turn prompts into realistic multi-turn user–agent conversations, enabling sys- tematic evaluation of conversational capabilities before launch. The self-improvement loop then uses this data and evaluation feedback to combine variance-based contrastive optimization with iterative refinement through a coding agent, allowing the system to auto- matically identify and fix planning and tool-use errors. Our approach provides fine-grained insights into conversational capabilities, uncovers issues before deployment, and improves qual- ity by +8% on top of a highly optimized manual prompt, automat- ically resolving several planning and tool-use errors. The system has been productionized and significantly accelerated iteration cy- cles for the launch of a conversational recommendation agent at Spotify. Online A/B tests demonstrate its effectiveness, with +14% user listening, +5% increase in weekly active users, and a 5% reduc- tion in skip rate compared to a prior experience that only supports session refinement. Overall, this work provides a practical framework for acceler- ating the development of conversational recommendation agents in industry, addressing challenges that are becoming increasingly central to recommender systems and agentic applications as natural- language interfaces reach widespread adoption.
- Spot V2Breaking the Loop: An Empirical Comparison of Strategies for Novelty and Freshness in YouTube Music
by Srivaths Ranganathan, Zihuan Diao, Bernardo Cunha, Joshua L. Moore, Robin Dumas, Murat Goksedef, Yanwei Song, Mukai Lu, Gergo Varady and Tracy PesinContinuously trained ranking models in music recommenders fall into feedback loops where previously consumed items dominate recommendations. This suppresses two distinct content classes: new releases (temporal freshness) and unlistened catalog items (novelty). Industry practitioners have a wide menu of interventions available, ranging from serving-time heuristics, training-data reweighting, architectural debiasing, to uncertainty-driven exploration, each of which are well understood in academic settings. But live systems offer challenges with continuously ingested content, interconnected components, and practical limitations that counteract the findings from academic research. We report results from off-policy online A/B tests for six interventions and a combination experiment across four conceptual layers (serving, training, architecture, exploration) on the YouTube Music homepage. All interventions modify the ranking model or the serving layer that consumes its scores; candidate generation and other upstream components are held fixed. We discuss key takeaways from our results: first, serving-time interventions on continuously trained systems are neutralized by the learning loop. Second, architectural debiasing reduces popularity dominance and improves diversity but does not create discovery, while carrying hidden integration costs. Finally, uncertainty-driven exploration interventions with a Spectral-normalized Neural Gaussian Process (SNGP) head produce the largest new-release lift, though they come with a measurable engagement or diversity tradeoff. We close with recommendations on which layer to intervene at, and the hidden costs of each choice.
- Spot W1Mosaic: A Fleet of User Embedding Specialists for Recommendation at Meta
by Zhiyuan Zheng, Xian Sun, Xiangyang Mou, Yujunrong Ma, Christina You, Michael He, Hrishikesh Paranjape, Aakarsha Agarwal and Hong LiUser representation is one of the highest-leverage modeling problems in industrial recommendation systems: a single advancement in how users are encoded can propagate across retrieval, ranking, and integrity tasks at platform scale. Prior industrial user representation work builds either a single user model that emits one or more embedding vectors or a shared backbone with task-specific adaptation. In this paper, we present Mosaic, a foundational user modeling platform that employs a fleet of specialists to learn user embeddings. The fleet comprises four architecturally diverse model families – memorization-driven, dense-heavy, sequential-based, and CoTrain models – each focusing on a distinct facet of user behavior. We developed MRM (Multi-task Relations Mining) and CRL (Cosine Redundancy Loss) techniques to maximize the marginal information contribution of each new specialist. We also introduce CoEval and User Tower Zero-Out, new logging-free embedding evaluation framework that improves development velocity while preserving downstream-aligned accuracy. Our hybrid CPU/GPU, online-and-offline serving stack allows each specialist to choose the adequate serving strategy to meet the freshness, latency, and computational requirements. Mosaic delivers consistent and significant offline NE improvements in addition to online gains.
- Spot W2A Unified Generative Re-ranking Framework with Adaptive Multi-objective Fusion
by Heng Zhang, Yifan Gu, Wei Xu, Lei Cheng, Chuan Yuan and Hengrui ZhangIn e-commerce recommendation, effective re-ranking must jointly optimize multiple business objectives such as clicks, add-to-carts, and purchases. Existing methods typically either employ multiple per-objective generators combined with static fusion through hand- crafted rules, or adopt non-autoregressive sequence generation with pre-defined objective fusion. Despite their differences, these approaches share a critical limitation: they treat multi-objective signals as statically fused, ignoring the fact that users have diverse and evolving intents in real time. To address this gap, we propose MAMGR, a unified generative re- ranking framework that enables end-to-end, sequence-level learn- ing of multiple objectives with dynamic weighting. MAMGR fea- tures a Multi-objective Adaptive Module (MAM) that learns person- alized objective weights in real time and shares them between the generator and evaluator, and a Dual-channel Context-aware Mod- ule (DCM) that jointly models local pairwise interactions and global sequence-level dependencies. Extensive offline and online experi- ments show that MAMGR consistently outperforms state-of-the-art methods, achieving +0.78% CTR and +1.43% GMV improvements while reducing CPU usage by 23.53% and incurring only +1.12% la- tency overhead. MAMGR has been deployed in Alipay’s production advertising system, serving hundreds of millions of users daily.
R&P Notes
- Spot J1In-Batch Negatives Can Silently Cripple LLM-Encoded Sequential Recommenders
by Younggue BaeLLM-based sequential recommenders increasingly train user/item encoders with in-batch InfoNCE negatives, a convention inherited from contrastive representation learning that keeps training tractable when batch size is constrained by encoder memory cost. We report an early finding from developing an LLM-encoded temporal memory recommender: with batch size 16 (15 in-batch negatives), training stalled at a validation NDCG@10 of roughly 0.035 for dozens of epochs, far below a strong ID-based baseline. Replacing the in-batch objective with full-catalog softmax cross-entropy over the item catalog (≈245K items), combined with a whitened, mixture-of-experts-adapted item encoder, nearly doubled validation NDCG@10 and closed the gap with a state-of-the-art semantic sequential recommender, reaching statistically indistinguishable NDCG@10 while the two methods split leads across the remaining ranking metrics. We describe the fix, report preliminary results, and surface an open question about which architectural components actually explain the remaining gain once the training objective itself is corrected—finding, via a completed sensitivity sweep, that contrastive temperature rather than any tested architectural component is the dominant remaining lever, and that its optimal setting reverses once the negative-sampling objective is fixed.
- Spot J2Calibrating Recommendations on Ordinal Attributes
by Marta Moscati, Varvara Toloknova, Oleg Lesota and Markus SchedlA recommendation list is said to be calibrated with respect to an item attribute if the distribution of the attribute over recommended items matches the distribution over the user’s consumed items. Calibration techniques were introduced for non-ordinal attributes such as genres, and later applied to ordinal ones like popularity. However, when the support of the distributions is ordinal, i.e., when the order of bins matters, current calibration techniques may fall short in capturing distribution displacements along the ordinal axis. To address this limitation, we propose to use order-sensitive measures of divergences between distributions. With quantitative experiments on movie and music recommendation, we show that standard, order-insensitive calibration metrics are not able to fully capture the miscalibration of ordinal attributes in recommendation lists. We then propose order-sensitive objectives for post-processing calibration and compare their impact on recommendations with that of order-insensitive approaches. We show that order-sensitive approaches reach a better accuracy-calibration tradeoff. With this work we aim to start the discussion on how to appropriately model order-sensitive item attribute for calibrated recommendations. Code for experiments: https://anonymous.4open.science/r/ordinal-calibration-experiments Code for dashboard: https://anonymous.4open.science/r/ordinal_calibration
- Spot K1Scaling Sequence Learning under Production Latency Constraints
by Zhirong Chen, Howard Cheng and Allen LinCapturing complex user behavior patterns has been shown to be key for prediction models in ads ranking. However, scaling user history modeling in online ranking is constrained by inference cost. While complexity can be offloaded to upstream models, transfer efficiency to the online model is limited, making it critical to scale the online model directly. Most prior work on scaling transformer-based sequence models targets retrieval or upstream settings; scaling the online ranking model under production latency constraints remains underexplored. We introduce two user-side model components, self-attention over user event histories and a compressed hypernetwork conditioned on user features, and show that both can be placed in a user request-only (RO) subnetwork evaluated once per request. On a large-scale production ads ranking dataset, scaling these components yields NE loss improvements with neutral serving QPS, demonstrating that prediction accuracy can be improved without increasing inference cost.
- Spot K2Intent-Description Anchoring Bias in LLM-as-a-Judge Evaluation of Recommendation Systems
by Himan Abdollahpouri, Kyle Kretschman and Mounia LalmasLarge Language Models (LLMs) are increasingly used to evaluate recommendation systems, but are known to exhibit systematic biases. We study whether LLM judges are influenced by descriptions of a recommendation algorithm’s optimization objective, even when evaluating identical recommendation outputs. We find that they are, a phenomenon we term intent-description anchoring bias, where an algorithm’s stated objective influences judgments beyond what is supported by the recommendations themselves. Across four frontier LLMs from three commercial providers, providing algorithm descriptions in the evaluation prompt increased diversity scores for identical recommendations by up to 0.89 points (Cohen’s d=1.82, p<10^-18), with substantial variation across models. Our results show that contextual information unrelated to recommendation quality can bias LLM-based evaluation, motivating protocols that hide algorithm metadata from LLM judges or apply mitigation strategies.
- Spot L1PCP-Decoding: Training-Free Personalized Category-Relative Popularity Calibration for Generative Retrieval
by Yuyang Qin, Congcong Liu, Yiming Sun, Cai Shang, Wenlong Chen, Changping Peng and Ching LawGenerative retrieval decodes hierarchical item identifiers under a prefix-constrained trie, where popular descendants can raise a shared-prefix score and prune lower-popularity siblings before the leaf level. We call this failure mode prefix equivalence-class collapse. We propose PCP-Decoding, a training-free method that calibrates next-token logits using category-relative prefix popularity and a user-specific strength () ∈ [0, ]. Masked normalization preserves comparability across prefixes, while the underlying popularity potential telescopes along a selected path. On three days of industrial evaluation traffic, exposure Gini decreases monotonically as increases, and Tail Node R@256 peaks at an intermediate setting. These preliminary results motivate broader offline and online evaluation.
- Spot L2Monte Carlo Power Analysis for Small-System A/B Trials: Detecting Differences in Newsletter Engagement
by Michael Ekstrand, Daniel Kluver, Karl Higley and Bart KnijnenburgGood experimental practice and predicting the scientific usefulness of an experiment or experimental platform require power analysis: estimating, a priori, how likely the experiment is to detect the intended effect if indeed it exists. While simple experimental designs admit well-understood power analysis methods, more sophisticated experimental designs and settings often require bespoke techniques to avoid either over- or under-estimating experimental power. We present the Monte Carlo method we use to estimate the power of experiments intended to increase user engagement with personalized e-mail newsletters of recommended news articles.
- Spot M1Sparse Autoencoders as Semantic Domain Adapters for Recommender Systems
by Vojtěch Vančura, Giacomo Medda, Martin Spišák and Ladislav PeškaPretrained text embeddings enable content-based cold-item recommendation, but their domain-agnostic objectives may underrepresent distinctions that matter within a particular recommendation domain. We investigate lightweight autoencoder-based adapters that specialize these embeddings without retraining the underlying encoder. Their parameters are learned solely from item-description embeddings, without interaction data; interactions are used only for validation-based model selection and evaluation. We compare a denoising autoencoder, a β-VAE, a sparse autoencoder, and its denoising variant across three pretrained embedding models and 13 recommendation datasets. Across 78 dataset-encoder-metric settings, a sparse variant achieves the best result in 70. Qualitative inspection further suggests that some sparse dimensions capture coherent domain-specific concepts, including automotive parts, baby products, and office supplies. We hypothesize that sparse autoencoders reorganize general-purpose embeddings into domain-specific concept spaces better suited to recommendation. These preliminary findings motivate sparse semantic adaptation as a distinct component of recommender systems.
- Spot M2Do Sequential Recommendation Benchmarks Really Require Higher-Order Sequence Modelling?
by Aleksandr V. Petrov, Praveen Chandar, Paul Bennett, Hugues Bouchard and Mounia LalmasSequential recommenders increasingly use language-model architectures designed to capture complex, context-dependent interactions. Yet it remains unclear whether widely used benchmarks actually require this modelling capacity. We investigate this question using two simple, recency-weighted pairwise probes that do not learn higher-order sequence representations: Sequential Rules (SeqRules) and our Probabilistic Collaborative Transition Model (PCTM). Using the evaluation protocol of eSASRec, at least one probe exceeds our eSASRec reproduction by 15–38% on three Amazon datasets and by 4.4% on MovieLens-1M, but trails it by 27.3 on MovieLens-20M. On the four remaining datasets, at least one probe also outperforms our sampled-softmax SASRec reproduction by 9–28%, suggesting that these widely used benchmarks are poorly suited to measuring gains from higher-order sequence modelling. More broadly, comparing Transformer-based models against strong recency-weighted pairwise probes provides a concrete test of whether a benchmark can meaningfully measure gains from higher-order sequence modelling
- Spot N1Don’t Waste the Last Token: A Quantizer-Agnostic Quality Boost from Semantic ID Collision Resolution
by Anna Volodkevich, Anton Klenitskiy, Artem Fatkulin, Darya Denisova and Alexey VasilevGenerative recommenders represent each item as a Semantic ID: a short sequence of discrete tokens. Popular quantization algorithms, such as RQ-VAE and residual K-means, can map several items to the same token sequence, producing collisions. Collisions must be resolved before training a single-stage generative recommender, since the model cannot distinguish items that share an identifier. The widely adopted disambiguation technique appends one extra codebook and uses the last token as an arbitrary item counter within each collision group. We argue that this last codebook can be used more effectively, as a source of additional content, popularity, or collaborative signal. Keeping the quantizer, the generative model, and the Semantic ID length unchanged, we replace the ordinal counter with tokens that both resolve collisions and carry a useful signal. Across several datasets, the collision resolution choice alone yields an easy-to-implement, quantizer-agnostic gain in recommendation quality (up to +12.7\% in NDCG@10). The code is available at https://anonymous.4open.science/r/sid-collisions-boost.
- Spot N2When Should Recommender Systems Not Act?
by Julia Neidhardt, Thomas Elmar Kolb, Ahmadou Wagne, Gwendolyn Rippberger and Ricardo Baeza-YatesRecommender systems (RSs) can produce useful outcomes, but they can also cause harm. This raises a basic operational question for responsible recommendation: under what conditions should a system withhold the action it would otherwise select? At each interaction, the system may respond, recommend, nudge, ask for clarification, defer, or take no action. Related forms of non-action appear in selective prediction, LLM refusal, clarification policies, and human handoff, but are typically studied in isolation. We examine justified non-action as restraint behavior in which an otherwise selected action is withheld because a different response or no response at all is judged to better protect users, reduce harm, or satisfy normative constraints. Rather than proposing a new abstention mechanism, we argue that justified non-action deserves explicit attention as a question of responsible recommendation. We organize the discussion around questions of detection, governance, alternative actions, and evaluation, and illustrate them through three ongoing projects on refusal and handoff in a university assistant, non-intervention in shopping assistance, and action selection in a conversational RS. Together, these projects motivate further research on when RSs should not act.
Short Papers
- Spot Q1Beyond Index-Only RoPE: Integrating Time and Order for Generative Recommendation
by Xiaokai Wei, Jiajun Wu, Daiyao Yi, Reza Shirkavand and Michelle GongIn recommendation, however, position has two distinct meanings: an event’s order in the sequence and its wall-clock time. Existing approaches usually inject temporal information through auxiliary embeddings or relative attention biases, while vanilla RoPE captures only order. We revisit RoPE for generative recommendation and ask how to encode both time and order without losing the benefits of rotary attention. We present Time-and-Order RoPE (TO-RoPE), a simple family of designs that uses both signals in rotary embeddings with three lightweight variants: early fusion, split-by-dimension, and split-by-head. The key insight is that early fusion can cause interference between time and order inside the same rotary plane, whereas split allocation provides a cleaner inductive bias. Across public dataset and a proprietary industrial dataset, TO-RoPE consistently outperforms absolute, relative-bias, and single-source RoPE baselines.
- Spot Q2Impression Share Prediction: An Offline Evaluation Task for Ads Ranking Systems
by Mohsen Malmir, Houssam Nassif, Danish Nasir Shaikh, Taher Rahgooy and Murat BayirOffline evaluation is the main gateway before deploying ads ranking models to A/B testing in production. Standard offline metrics measure predictive accuracy, but are only a surrogate for advertiser value-the total conversions and revenue generated through the impressions advertisers receive. Advertiser value depends not only on prediction quality but on how impressions are distributed across objective buckets (advertiser campaigns organized by optimization goal, such as clicks, purchases, or video views). A model can show superior offline performance while shifting this distribution in ways that degrade advertiser value once deployed. No existing offline evaluation method provides visibility into these impression share shifts before deployment. We propose impression share prediction as an offline evaluation task: given a candidate ranking model, predict the distribution of impressions it would produce across objective buckets. The task is inherently counterfactual-the candidate has never served live traffic, pacing controllers have not adapted to it, and budget dynamics still encode the prior model’s equilibrium. We propose a structural causal model that captures how model predictions, advertiser budgets, and pacing jointly determine impression allocation, and show the counterfactual effect is identified from observational data. Building on this, we develop a statistical learning framework that predicts impression shares from offline model signals and current market state, trained on historical model deployments. On production data from multiple ranking model families, a Random Forest predictor reduces L1 prediction error by 49% over a constant baseline for models seen during training. For models held out from training we evaluate by time since the candidate first appeared in the live system; the first hour is the closest empirical analog to true pre-deployment, since the candidate has barely interacted with the system. In this regime the Random Forest falls below the constant baseline because the marketplace budget state still reflects the prior model’s equilibrium. An encoder-conditioned architecture that simulates a 2-hour rollout over recent marketplace dynamics recovers +22% L1 improvement in this regime.
- Spot R1Support Gap: Selecting Fixed-K Candidate Sets for Retained Personalized Headroom
by Teresa ZhangTwo-stage recommenders often compare fixed-size candidate sets before running an expensive reranker or online experiment. Recall at K and spread diagnostics describe exact-target coverage or set geometry, but they do not answer a fixed-budget bottleneck question: once retrieval has selected K items, how much state-contingent choice remains for a richer downstream ranker? We study retained personalized headroom, the gap between the value of fine-state-specific choices and the value of the best single coarse-state choice within the same candidate set. We introduce Support Gap, or SG, a candidate-local witness estimate of this quantity. Under a candidate-local uniform approximation condition, SG admits an explicit sufficient-condition error bound for retained headroom. Empirically, we use this theorem as an interpretive lens rather than as a certificate for the evaluated sets: witness audits show informative within-set rankings and reasonable central fidelity, but worst-case error remains large and the uniform condition is not verified. Across four offline constructions on public datasets, selecting candidate sets by SG lowers retained-headroom regret relative to Recall@K, RidgeProxy, and HardProxy. Pooled normalized regret is 0.2839 for SG, 0.3902 for Recall@K, 0.3327 for RidgeProxy, and 0.3311 for HardProxy.
- Spot R2A Causal Transformer Multi-Touch Attribution with Dual Debiasing and Explainable Visualization for Advertising Recommendation
by Jiarong Zhang and Jing GaoMulti-touch attribution (MTA) aims to estimate the causal effect of advertising touchpoints on user conversions, providing essential support for budget allocation and advertising recommendation. Existing causal MTA methods often struggle to capture complex temporal dependencies in long touchpoint sequences and fail to adequately address confounding biases arising from both static user profiles and dynamic behaviors. To address these challenges, we propose CT-MTA, an end-to-end causal Transformer framework for multi-touch attribution. CT-MTA integrates a causally masked Transformer to model sequential dependencies, together with a Variational Autoencoder (VAE) and a Gradient Reversal Layer (GRL) to mitigate dual confounding biases. Furthermore, we design a counterfactual attribution module with a Mixture-of-Experts (MoE) architecture to generate personalized and interpretable attribution estimates. Experiments on the public Criteo datasets show that CT-MTA improves conversion prediction Area Under the ROC Curve (AUC) by 1.2% over the state-of-the-art method. In downstream budget allocation tasks, CT-MTA consistently reduces Cost Per Acquisition (CPA) and improves Conversion Rate (CVR) under limited budgets. Qualitative analysis further shows that CT-MTA alleviates over-attribution to historically favored touchpoints, leading to more efficient and fair budget allocation.
- Spot S1Position Bias Induces Inconsistent Rankings in Listwise LLM-based Recommendation
by Ethan Bito, Yongli Ren and Estrid HeLarge language models (LLMs) have emerged as promising listwise rerankers for recommender systems, but their reliability remains unclear. We show LLM-based rankings are not permutation-invariant, as reordering the same candidate set can change both the final ranked list and the pairwise preferences between items. We argue this behavior reflects a deeper structural problem, rather than simple output variability. Position bias perturbs local pairwise decisions, so the probability one item is ranked above another depends on input position. When aggregated across permutations, these local distortions can produce a pairwise preference system that is not consistent with any single global ranking. Instability in the final ranked lists is therefore a consequence of this underlying inconsistency. To study this phenomenon, we propose a multi-level evaluation framework that measures reliability at three levels, local instability in position-conditioned pairwise preferences, global inconsistency in aggregated pairwise preferences, and listwise instability across full rankings. We also consider a simple sequential inference strategy, Stochastic Greedy Selection, that reduces dependence on any single input ordering. Experiments across multiple models, datasets, and list lengths show mitigation methods do not improve reliability uniformly across levels. Methods that improve effectiveness or stabilize final rankings can still exhibit substantial inconsistency in the underlying preference structure. Our results show stability at the level of final rankings does not guarantee a coherent ranking process, and motivates evaluating on whether LLM-based recommenders define a consistent ranking function.
- Spot S2Choosing What Matters: Query-Aware Multimodal Routing for Conversational Recommendation
by Piao Huilin, Seoin Choi, Junbo Shim and Hayoung OhConversational recommender systems often benefit from evidence beyond dialogue context, and recent multimodal approaches have explored structured knowledge, textual semantics, and visual information to enrich user preference modeling. Building on MSCRS, which integrates multimodal graph representations with prompt learning, we propose QAMR-CRS (Query-Aware Multimodal Routing for Conversational Recommendation), a routing-based framework that explicitly estimates which modality-specific evidence is more relevant to the current query. QAMR-CRS forms a query representation from dialogue context and mentioned entities, uses it to route knowledge graph, co-occurrence, text similarity, and image similarity signals, and combines the routed evidence with a static multimodal prior before converting it into layer-wise prompt prefixes for the GPT-based recommender. We further apply entropy regularization to mitigate excessive dependence on a single modality. Experiments on ReDial and INSPIRED show that QAMR-CRS improves over MSCRS across major ranking metrics, with ReDial results averaged over three random seeds. Routing analysis further shows that entropy regularization mitigates single-modality collapse and promotes distributed routing.
- Spot T1Enriching Graph-Based Modeling with Sequential Signals for Explainable Course Recommendation
by Md Akib Zabed Khan, Dongsheng Luo and Agoritsa PolyzouCourse recommendation systems (CRS) play a crucial role in guiding university students through their academic journey by assisting in course selection. At the same time, models that function as black boxes and offer little transparency in their decision-making process limit students’ trust and adoption of CRS. To address this, we propose Twiner, a TWofold sequential INtegration with knowledge graphs for Explainable next-basket Recommendation, that captures sequential patterns at two complementary levels: within the knowledge graph itself and through a recurrent model. This dual infusion enables richer contextual modeling for domains where the sequential nature of the data plays an important role. Our approach integrates graph attention networks (GAT) to model complex relationships (including sequential ones between items) and gated recurrent units (GRU) to capture sequential patterns in students’ course-taking behavior across semesters. To enhance explainability, we leverage the attention weights from the GAT module to provide path-reasoning-based justifications. We conduct experiments on four real-world datasets in the domains of course recommendation in higher education and online learning platforms, as well as grocery shopping. We show that our Twiner model performs similarly to or better than existing state-of-the-art models while also enabling richer explanations. Our findings highlight a transparent and effective solution where explainability does not come at the cost of performance.
- Spot T2Covariance-Aware Newton-Schulz Orthogonalization for Noise-Robust Sequential Recommendation
by Jinxin Hu, Hao Deng, Haibo Xing and Lingyu MuAutoregressive sequential recommendation trains transformer models on tokenized user behavior via cross-entropy. The Muon optimizer projects gradient matrices onto the Stiefel manifold through Newton-Schulz orthogonalization, ensuring uniform utilization of all parameter directions. However, we identify that this projection equalizes all gradient singular values to unity, collapsing the natural energy separation between signal and noise subspaces. In recommendation, where long-tail distributions, noisy implicit feedback, and exposure bias produce corrupted gradients, this equalization promotes noise directions to the same magnitude as signal directions—a directional amplification that magnitude-based techniques such as gradient clipping and SAM cannot address. We propose Covariance-Aware Newton-Schulz (CovNS), which estimates gradient directional stability via an EMA of the gradient outer product and attenuates volatile directions before orthogonalization. Experiments on a public benchmark and an industrial dataset show that CovNS outperforms Muon and all baselines across all retrieval metrics on both datasets with minimal overhead.
- Spot U1Comparability in Recommender System Evaluation
by Stefania Ionescu, Ilia Shilov and Florian DörflerThe increased awareness of the limitations of single-dimensional, user-centric evaluation motivated new emergent research in multi-metric and co-designed recommender systems. A key challenge here is integrating a multitude of potentially contradicting stakeholder perspectives and using them for meaningful comparisons between models. To this end, our work considers the full utility vector of stakeholders under alternative models, and uses social choice for aggregating individual benefits into a societal evaluation via social welfare functions. This allows us to bring entropy-based indexes used in algorithmic fairness that measures both group and individual unfairness as inequalities in benefits. Moreover, we use social choice theory to analyze comparability issues (i.e., the benefit variations of one stakeholder not being comparable with those of another) and suggest analytically informed solutions. We find that considering relative gains can restore comparability for fairness analysis while Nash social welfare provides a natural aggregation rule when social evaluation must remain invariant to stakeholder-specific rescaling. We also apply this approach to a real-world dataset and address limitations in practical implementations.
- Spot U2Uncertainty-Aware Gated Context Fusion for Next POI Recommendation
by Soyoung Jang and Jaekwang KimNext Point-of-Interest (POI) recommendation requires modeling both sequential user behavior and rich contextual signals such as venue category and visit timing. While sequential models effectively capture visit patterns, they often underutilize these cues or rely on fixed-weight fusion that fails to account for varying informativeness across categories. We propose a modular gated fusion framework that integrates item, category, and temporal embeddings into any sequential backbone, with an uncertainty-aware gating mechanism that adaptively controls the contribution of contextual information based on category-level embedding variance. Experiments on Foursquare NYC/TKY and Yelp demonstrate consistent HR@K and NDCG@K improvements over strong baselines across multiple sequential recommendation architectures. We further analyze the learned gates to reveal when and how contextual signals contribute to next-POI prediction.
RecSys 2026 (Minneapolis)
- About the Conference
- Registration
- Program at Glance
- Program
- Call for Contributions
- Challenge
- Keynotes
- Accepted Contributions
- Presenter Instructions
- Workshops
- Tutorials
- Committees
- Inclusion
- Student Volunteers
- Women in RecSys
- Visa Information
- Addressing Attendance Issues
- Location / Hotel
- Lasting Impact Award
- 60 Milestones for 20 Years




















