Session 5:
5A: Industrial Ranking, CTR & Scalable Models
Date: Wednesday September 30, 14:00 – 15:30 CDT
Session Chair: Kim Falk
- INDSample Is Feature: Beyond Item-Level, Toward Sample-Level Tokens for Unified Large Recommender Models
by Shuli Wang, Junwei Yin, Changhao Li, Senjie Kou, Chi Wang, Yinqiu Huang, Yinhua Zhu, Haitao Wang and Xingxing WangScaling industrial recommender models has followed two parallel paradigms: \textbf{sample information scaling}—enriching the information content of each training sample through deeper and longer behavior sequences—and \textbf{model capacity scaling}—unifying sequence modeling and feature interaction within a single Transformer backbone. However, these two paradigms still face two structural limitations. Firstly, sample information scaling methods encode only a subset of each historical interaction into the sequence token, leaving the majority of the original sample context unexploited and precluding the modeling of sample-level, time-varying features. Secondly, model capacity scaling methods are inherently constrained by the structural heterogeneity between sequential and non-sequential features, preventing the model from fully realizing its representational capacity. To address these issues, we propose \textbf{SIF} (\emph{Sample Is Feature}), which encodes each historical Raw Sample directly into the sequence token—maximally preserving sample information while simultaneously resolving the heterogeneity between sequential and non-sequential features. SIF consists of two key components. The \textbf{Sample Tokenizer} quantizes each historical Raw Sample into a Token Sample via hierarchical group-adaptive quantization (HGAQ), enabling full sample-level context to be incorporated into the sequence efficiently. The \textbf{SIF-Mixer} then performs deep feature interaction over the homogeneous sample representations via token-level and sample-level mixing, fully unleashing the model’s representational capacity. Extensive experiments on a large-scale industrial dataset validate SIF’s effectiveness, and we have successfully deployed SIF on the Meituan food delivery platform.
- INDAlignment + Accuracy: The Cascade Reward Representation for Preranking
by Hedi Xia, Dylan Zhou, Yali Bian, Yichu Zhou, Zili Li, Tianyou Wang, Bella Huang, Hongbo Deng, Piyush Maheshwari, Dafang He, Darren Reger, Bowen Deng and James LiPrerankers in large-scale recommender systems select candidates for a downstream ranker under strict latency constraints. In practice, teams combine accuracy metrics with alignment losses to train and evaluate prerankers, but what these quantities should target—and how to combine them—remains ad hoc. We derive the \emph{Cascade Reward Representation}: under mild assumptions on a fixed-retrieval, fixed-ranker pipeline, the expected change in user reward for preranker swaps with controlled overlap shift admits a calibrated first-order two-term representation $\mathbb{E}[\mathbf{R}^E – \mathbf{R}^0] = \alpha\,\mathbb{E}[\Delta\hat O] + \beta\,\mathbb{E}[\Delta N] + \mathcal{R}$, where $\Delta\hat O$ is a logged top-fraction overlap shift (\emph{alignment}: agreement with the main ranker’s selections), $\Delta N$ is a threshold-conditioned precision shift (\emph{accuracy}: engagement above a shared ranker threshold), and the remainder is bounded by $O(\mathbb{E}[\Delta^2]) + O_P(1/\sqrt{n})$. Both proxies are computable from production logs without running the main ranker on the full retrieval pool. This representation motivates a calibrated offline metric and a matching two-branch training loss for the model family studied in our production system. We validate the representation in a large-scale industrial recommender system. A calibrated linear combination of the two metrics raises experiment winner prediction from $45$–$50\%$ (accuracy-only) to $85\%$ and Pearson $r$ from ${\le}0.65$ to $0.84$ on held-out experiments. The matching training objective delivers $+1.43\%$ save engagement over an accuracy-only baseline and $+0.62\%$ over a heuristic alignment+accuracy production model in two-week A/B tests.
- INDAttending to the Core: Core-Task Attention for Recommendation
by Jingyan Chen, Chenye Sun, Man Zhou, Yunhe Guo, Siyu Gu and Peng JiangIndustrial recommender systems typically support multiple busi- ness objectives through the integration of many specialized models. However, each individual model is usually optimized for a single primary target, such as conversions or purchases. For example, in advertising systems, models are often trained for OCPX-style objectives that tightly couple target prediction with bidding and revenue optimization. Since these target signals are often extremely sparse, multi-task learning (MTL) is widely adopted to leverage denser auxiliary tasks for additional supervision. However, exist- ing MTL approaches typically pursue balanced joint optimization across tasks, which may introduce task interference and degrade the performance of the core task. To address this limitation, we propose CoreAtt, a core-task-centric method that employs a novel attention mechanism to adaptively aggregate informative represen- tations from auxiliary tasks into the core task’s prediction tower. CoreAtt consists of two complementary attention pathways: (1) intra-sample attention, which models instance-level interactions among auxiliary tasks to produce context-aware fusion signals, and (2) inter-sample attention, which assesses each auxiliary sig- nal’s global reliability by comparing its prediction score against the population distribution. These two pathways are fused through a lightweight gating mechanism to enrich the representation for the core task. Notably, CoreAtt achieves strong performance even when built upon a simple Shared-Bottom architecture (CoreAtt-SB). We evaluate CoreAtt-SB on three public recommendation datasets, where it consistently outperforms strong MTL baselines while pre- serving auxiliary task performance. Moreover, online A/B tests on a leading short-video platform show that CoreAtt significantly enhanced platform revenue and advertiser value by 3.3% and 2.3%, respectively. Our code is available here1.
- INDWHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture
by Renqin Cai, Dawei Sun, Yuanjun Yao, Zhiyong Wang, Velvin Fu, Maggie Zhuang, Yu Shi, Zhongnan Fang, Xuan Cao, Jing Qian and Rui LiAs scalability becomes increasingly important in recommendation model, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and cross features; and sequence features from user behavior histories. Wukong and HSTU have emerged as representative scalable backbones for these paths: Wukong scales high-order non-sequence feature-interaction modeling, while HSTU scales long user-behavior sequence modeling. Despite their complementary strengths, practical architectures that combine these two types of feature modeling remain underexplored. We present WHALE, a scalable unified recommendation architecture that jointly models non-sequence and sequence features on top of Wukong and HSTU. Each WHALE layer contains a Wukong module, an HSTU module, and an attention-based fusion module in which Wukong-derived interaction representations query HSTU-derived behavior representations. This design keeps both backbones active throughout the network and enables progressive Wukong–HSTU exchange, allowing high-order feature crosses to repeatedly retrieve fine-grained evidence from long user histories. To make WHALE practical for industrial deployment, we introduce customized Triton kernels and other model-systems co-design techniques to improve training and inference efficiency. On large-scale industrial recommendation data, WHALE achieves consistent gains in offline experiments. Additionally, it delivers positive online gains with a modest serving-throughput trade-off. Overall, WHALE provides a practical example of how these two sources of information can be scalably unified in an industrial recommendation model.
- RESDo We Really Need LLMs to Augment All? A Selective Augmentation Framework with Lightweight Language Models for Multimodal CTR Prediction
by Ziyun Chen, Yuhan Wang, Honghao Li, Mengzi Tang, Qing Xie and Yongjian LiuIn recent years, multimodal models and large language models (LLMs) have been increasingly applied to click-through rate (CTR) prediction, owing to their ability to extract recommendation-relevant information from raw content and thereby alleviate the long-standing issue of collaborative signal sparsity. Existing paradigms typically either feed continuous embeddings from pretrained multimodal encoders into CTR models, or employ LLMs for unified knowledge augmentation. While these approaches have demonstrated promising performance, when and why their gains emerge remains unclear, and the inference cost of large LLMs introduces severe latency and scalability challenges in large-scale deployment. In this work, we empirically show that the gains of LLM-augmented CTR are highly item-dependent: uniformly applying stronger augmentation to all items is often inefficient, while the main benefits concentrate on a subset of items with more complex interaction patterns. Motivated by this observation, we propose A Selective Augmentation Framework with Lightweight Language Models for Multimodal CTR Prediction, dubbed SALM, which revisits the data augmentation paradigm and replaces indiscriminate full-data augmentation with a difficulty-aware selective strategy targeting necessary items. Extensive experiments on real-world multimodal datasets, across multiple CTR backbones and augmentation baselines, demonstrate that SALM consistently improves AUC and reduces LogLoss, while simultaneously reducing LLM-based augmentation cost.
- INDHeterogeneous Ranking in Industrial-Scale Recommender Systems: A Case Study
by Di Bai, Jintao Liu, Zhenwei Tang, Peifan Wu, Nada Al-Thawr and Luoshu WangHeterogeneous recommendation feeds present complex challenges that extend beyond those found in highly homogeneous environments (e.g., music-only or video-only closed-ecosystem platforms). In Google Discover, a unified feed integrates diverse content sourced from the decentralized open web, including web articles, long-form and short-form videos, user-generated content (UGC), and beyond. Different content types exhibit distinct feature densities and user interaction patterns. Building a unified ranking model that sustains high performance across such heterogeneity, while avoiding negative transfer or majority bias, remains a significant industrial challenge. This paper presents an end-to-end case study on the industrial-scale multi-task ranking of heterogeneous feeds, grounded in real-world deployment. We introduce HA-MoE, a heterogeneity-adaptive multi-gated mixture-of-experts architecture that incorporates explicit heterogeneity context into both gating networks and expert representations. This approach enables effective specialization without significantly increasing operational overhead. To support reliable deployment, we introduce LENS, a lightweight observability framework that provides interpretable diagnostics of expert specialization and tracks this functional heterogeneity across continuous retraining. We evaluate our method using Dual-Level AUC (DL-AUC), a heterogeneity-aware evaluation metric that combines global ranking performance with cross-segment ranking correctness. Offline evaluations on a large-scale industrial dataset demonstrate consistent improvements over baseline models. Furthermore, online A/B testing confirms gains in feed activity and exploration metrics. Together, offline and online results validate the effectiveness of our approach for managing heterogeneity in industrial-scale recommender systems.
- INDIDProxy: CTR Prediction with Multimodal LLMs for Cold-Start Recommendation at Xiaohongshu
by Yubin Zhang, Haiming Xu, Guillaume Salha-Galvan, Ruiyan Han, Feiyang Xiao, Yanhua Huang, Li Lin, Luo Yang and Yao HuContent-driven platforms such as Xiaohongshu often leverage clickthrough rate (CTR) prediction models for recommendation. However, these models depend heavily on item ID embeddings, which perform poorly in item cold-start settings. In this paper, we present IDProxy, a production-scale system developed at Xiaohongshu to address this challenge. IDProxy leverages multimodal large language models (MLLMs) to generate proxy embeddings from rich content signals, enabling CTR prediction for new items in the absence of usage data. Through a lightweight coarse-to-fine mechanism, these proxies are aligned with the ID embedding space and trained endto-end with the ranking model, allowing seamless integration into production-facing pipelines. Extensive offline and online experiments demonstrate the effectiveness of the method, which has been deployed in 2025 in Xiaohongshu’s Content Feed and Display Ads features, reaching hundreds of millions of users daily.
5B: Explainability Methods & Human-Centered Evaluation
Date: Wednesday September 30, 14:00 – 15:30 CDT
Session Chair: Christine Bauer
- RESTRACE: Targeted Ranking-Aware Counterfactual Explanation for Sequential Recommendation
by Ungsik Kim, Sang-Min Choi, Gun-Woo Kim and Suwon LeeRanking-constraint counterfactual explanation for sequential recommendation requires query-limited search to decide where to edit and what to substitute—the bottlenecked for query efficiency lies more in how the search space is structured than in the mutation rate alone. We propose TRACE (Targeted Ranking-Aware Counterfactual Explanation), which decomposes the search into three stages under an embedding-accessible, training-free setting: influence-guided position selection, plausibility-aware candidate retrieval, and actual-margin beam pruning. Across five datasets and three recommender architectures, TRACE outperforms the GA-based baseline GECE in validity, query efficiency, edit cost, and likelihood preservation, with the largest gains on bring-in (raising a target item to top-1) and up to an order-of-magnitude reduction in queries on push-out (displacing the current top-1 beyond top-K). Ablations confirm that the gains arise from structuring the search space before evaluation, rather than from increased mutation frequency. Available code: https://anonymous.4open.science/r/recsys-anon-1D31/
- RESWhen attention is bounded, structure matters: personalizing recommendation explanations under time constraints
by Deo Munduku and Elsa NegreRecommendation explanations are often consulted in situations where reading time is limited—for example, when users browse options on a mobile device while on the move—and can process only part of the available information. Yet, little is still known about how such explanations should be structured in this kind of context in order to continue supporting decision-making. In explainable recommender systems, personalization has so far focused mainly on explanation content or linguistic form, much less on explanation structure. In this paper, we study explanation structure as a personalization variable under time constraints. To this end, we introduce a user interpretive schema, defined as an explicit representation of the relative importance the user assigns to the different pieces of information relevant to their decision, and we use this schema to structure the explanation by ordering this information according to its importance to the user. From this, we derive a structuring principle: when reading time is limited, the information that matters most to the user should appear as early as possible. We evaluate this principle in a controlled user study involving 663 participants in a restaurant recommendation scenario, where explanation variants differ only in the order in which information is presented. The results show that a structure aligned with the user’s interpretive schema improves the explanation’s ability to support decision-making when time constraints are strong, whereas this effect weakens as more time becomes available. These findings suggest that explanation structure is not merely a matter of presentation, but constitutes a genuine personalization.
- RESDo We Care About Personalization and Explainability? An Interview Study with News Recommendation Engineers
by Jasmin Kareem, Siddharth Mehrotra, Martijn Willemsen and Maarten de RijkeResearch on explainability in recommender systems largely centers on end users, overlooking the perspectives of those who build and maintain these systems and their potential use cases such as model debugging. In this study, we examine how news engineers and related technical stakeholders perceive and implement personalization and explainability in practice. We conducted 15 semi-structured interviews across nine news organizations, spanning diverse regions in both public and private sectors, to investigate the challenges and motivations shaping their approaches. Our findings reveal that personalization is not always a straightforward or desirable choice for news organizations, as concerns around user tracking, editorial control, and resource constraints often limit its adoption. Even among organizations implementing personalized news recommender systems in production, explainability is rarely prioritized, with day-to-day operational demands frequently taking precedence over longer-term transparency goals. Definitions of explainability vary widely across organizations, though some demonstrate promising internal practices and visualization tools that facilitate communication between engineering teams and newsrooms. Based on our analysis, we provide actionable and practical guidelines for news engineers and researchers on how to adopt explainability methods within a news personalization pipeline.
- RSChoploy: a Plugin-Based Inference Layer for Path-Based Explainable Recommendation over Knowledge Graphs
by Ludovico Boratto, Gianni Fenu, Mirko Marras, Giacomo Medda and Alessio SechiPath-based reasoning over knowledge graphs is a promising approach for explainable recommendation, as it justifies recommendations through entity-relation paths connecting user preferences to suggested items. However, existing implementations mainly target offline training and evaluation, while interactive deployment requires domain-specific handling of model loading, KG access, decoding, and explanation generation. We present hoploy, an open-source inference and explanation layer for the hopwise ecosystem that exposes pre-trained path-reasoning models through configurable APIs. hoploy supports stateless scenarios where users are unknown at training time and provide preferences only at request time. Its plugin architecture lets developers define request/response schemas, configuration files, and decoding logic that maps KG tokens into human-readable explanations. We release and demonstrate hoploy with POI and food plugins, assess their implementation effort and serving footprint, and provide documentation for extension. Resource: https://github.com/tail-unica/hoploy.
- RESThe Utility of LLMs in Recommender Systems Explanation Evaluation
by Kathrin Wardatzky, Oana Inel, Luca Rossetto and Abraham BernsteinExplanations play a crucial role in creating trustworthy recommender systems (RS), yet choosing a good explanation method comes with challenges. Many explanation methods are available, but little guidance exists on which is best for which setting. Existing explanation generation methods often produce abstract outputs that require further formatting to become user-friendly with a seemingly endless pool of options. Running user-based evaluations of all possible options is usually unfeasible, but existing automated evaluation metrics often either assess only the explainer’s abstract output or require comparison with a ground truth, which is generally unavailable. Recent studies have shown that large language models (LLMs) can serve as judges in explanation evaluations, but their reliability has not yet been thoroughly explored. This paper investigates the utility of LLMs in selecting an effective explanation method for a given application. We first explore their ability to generate explanation prototypes given varying information about the RS and the user in the prompts. Specifically, we generate 18 distinct explanation prototypes, which are subsequently evaluated by 14 LLMs of varying sizes across two temperature settings. We compare these against human ratings derived from a user study. Our results show that while LLMs exhibit human-like rating patterns and achieve moderate rank correlation with human raters, their absolute rating agreement is low and varies substantially by model size and evaluation construct. We derive four practical recommendations: keep explanation-generation prompts concise, prefer larger models for evaluation, pre-test evaluation constructs, and audit explanations for factual accuracy, as neither humans nor LLMs reliably detect non-factual content.
- RESWhen Do Contrastive Explanations Really Matter in Recommender Systems?
by Thi Ngoc Trang Tran, Sebastian Lubos, Alexander Felfernig, Mehrdad Rostami, Viet-Man Le and Damian GarberRecommender systems often provide explanations to help users understand why items are suggested. Beyond such non-contrastive explanations, systems can also offer contrastive explanations, such as proposing alternative options (“Something different?”) or explaining why a specific item was not recommended (“Why not Item Y?”). This paper investigates when contrastive explanations matter across different contexts. We report on a user study examining two contrastive explanation types across item domains (movies, shopping, and housing) and recommendation algorithms (collaborative, content-based, and constraint-based). Our results show that contrastive explanations yield context-dependent differences compared to baseline explanations on traditional explanation goals, with generally small to moderate effects. Users tend to perceive contrastive explanations as more necessary in higher-involvement domains and in constraint-based recommendation scenarios, where reasoning about trade-offs, alternatives, and excluded options becomes more important, while baseline explanations remain effective for efficient decision-making. Overall, these findings suggest that contrastive explanations should be viewed as complementary mechanisms and selectively deployed based on recommendation contexts.
- RESFASC: A Feature Aspect-Level Sentiment Consistency Framework for Explainable Recommendation Evaluation
by Chenfu Yu, Qinglin Huang, Xiaoxuan Shen, Qian Wan, Zhicheng Dai, Jianwen Sun and Ruxia LiangRecent research in explainable recommendation commonly uses natural language explanations to improve transparency and user trust. However, reliably evaluating whether explanations are semantically faithful to users’ multi-dimensional preferences remains challenging. Existing methods mainly rely on text similarity metrics (e.g., BLEU, ROUGE) or shallow feature matching, which are insufficient for assessing whether explanations accurately reflect user preferences, especially in multi-aspect settings. To address these limitations, we propose Feature Aspect-Level Sentiment Consistency (FASC), a framework that quantifies semantic consistency between generated explanations and user-authored reference explanations through aspect coverage and sentiment polarity. FASC uses LLMs as auxiliary tools to extract structured aspect–sentiment units from explanation texts, enabling reproducible metrics for aspect coverage, correctness, and sentiment alignment. To validate the framework, we re-annotated several widely used explainable recommendation datasets to construct benchmarks with fine-grained, aspect-level sentiment labels. We further conducted human studies using pairwise comparisons, showing that FASC aligns more closely with human judgments than traditional metrics. Experimental results indicate that FASC can distinguish subtle differences in semantic faithfulness across models. To support future research, we release all annotated datasets and human evaluation results at https://anonymous.4open.science/r/FASC-F54E
RecSys 2026 (Minneapolis)
- About the Conference
- Registration
- Program at Glance
- Program
- Call for Contributions
- Challenge
- Keynotes
- Accepted Contributions
- Presenter Instructions
- Workshops
- Tutorials
- Committees
- Inclusion
- Student Volunteers
- Women in RecSys
- Visa Information
- Addressing Attendance Issues
- Location / Hotel
- Lasting Impact Award
- 60 Milestones for 20 Years




















