Day 1 Posters: Doctoral Symposium + Industry + R&P Notes + Short Papers
Date: Tuesday September 29
Posters are available from 10 am to 4 pm, and are presented by the authors during the coffee breaks.
Doctoral Symposium Papers
- Spot A1From Click Modeling to Offline and Off-Policy Evaluation in Carousel Recommendation
by Jingwei KangCarousel interfaces are widely used in modern recommendation systems. Unlike traditional interfaces that present a single ranked list, carousels simultaneously present several ranked lists to the user, as horizontally swipeable rows stacked on top of each other. In this design, the rankings are closely tied to the two-dimensional layout. Consequently, user behavior is shaped not only by item preference, but also by row organization, viewport constraints, and item context. This tight coupling between ranking and presentation complicates the interpretation of user feedback, introducing new challenges for recommendation evaluation.
- Spot A2Human Value in Recommendations: Towards Improved Evaluation Measures
by Troy ZadaWhen evaluating top-N recommendations, objectives beyond accuracy are increasingly considered, reflecting a broader shift toward user-centric evaluation. However, existing evaluation measures present several limitations: (1) they rely on assumptions about what users truly value, (2) their usage is inconsistent across studies, with different subsets of measures chosen depending on the task, and (3) they often overlook aspects of the intra-list arrangement of recommended items. These limitations highlight the need for evaluation approaches that better align with user preferences. Therefore, in the domain of movie recommendations, this proposed research investigates user preferences over multiple recommendation objectives in a controlled user study. By systematically varying the contribution of each objective and observing user preferences across these variations, the study infers both individual preferences and the relative importance of each objective. The overarching goal is to develop evaluation measures that better reflect how users actually perceive recommendation quality.
- Spot B1Teaching Recommenders to Listen: Bidirectional Natural-Language Control for Personalized News
by Mahamudul HasanPersonalized recommenders learn mostly from implicit signals such as clicks and a few coarse topic ratings. These signals are a narrow channel: a reader cannot easily ask for “more soccer World Cup coverage this month and less celebrity news,” and today’s systems give little indication of whether such a request was understood or acted on. This research asks how to let users steer a personalized news newsletter using natural language. Our aim is to support collection-level feedback in natural language—feedback about the balance and mix of the whole newsletter rather than a single article—and to make the exchange bidirectional, so the system both acts on a request and reports what it did. This extends natural-language critiquing, which has largely addressed individual recommendations, to an assembled collection, and foregrounds the natural-language interaction in a live newsletter setting. We pursue the goal through three planned studies on POPROX, a live platform that emails daily personalized newsletters, with 15 articles in five sections from the Associated Press, to real subscribers and supports randomized experiments. Study 1 investigates how to elicit actionable collection-level instructions and develops a rubric for scoring them. Study 2 builds the channel: an LLM post-ranker over the deployed neural ranker, NRMS, and evaluates it in a five-week randomized controlled trial (RCT). Study 3 adds a per-newsletter explanation of what the system understood, changed, and could not do, and tests whether it improves understanding and the quality of users’ revised instructions, while examining its effects on engagement. None of the studies has been run; this paper presents the plan, the designs, and a baseline-informed analysis.
- Spot B2A Design-Space Approach to Robust Agentic Conversational Recommendation
by Alessandro Francesco Maria MartinaAgentic Conversational Recommender Systems (ACRSs), in which a large language model orchestrates dialogue, item retrieval, and tool use (such as querying item attributes or ranking candidate items), are advancing quickly, yet two problems prevent us from fully understanding their potential and limitations. First, progress is reported at the level of whole systems, so gains cannot be traced to specific design choices. Second, evaluation relies on simulated users who arrive with clear, well-defined needs, leaving untested the users a conversational system should help most—those who cannot yet put their preferences into words. The goal of my PhD project is to address these problems. I take a design-space approach, in three stages: (1) a characterization of the design space of an ACRS: the individual design choices, and a controlled method that studies each in isolation; (2) an analysis of which of these design choices most determine how well an ACRS elicits and satisfies a user’s preferences, and of which effects are invariant across architecture, domain, and model family versus contingent on the conditions of deployment; and (3) an evaluation that spans simulated users across the full range of user behavior and tests the resulting conclusions against real users, mapping when simulation-based evaluation can be trusted. The first stage is complete; the other two are under way and form the focus of the rest of my PhD. Ultimately, I aim to build conversational recommenders that help users move from an uncertain, ill-formed need to a satisfying choice—adapting across the full spectrum of preference-awareness and initiative, robust to that diversity, and grounded in design principles whose conditions of validity are known.
- Spot C1Aligning Algorithms with Axiology: Operationalizing Journalistic Values in News Recommender Systems
by Aishwarya SatwaniRecommender systems shape how people shop, learn, and read the news. Advances in recommendation algorithms have made personalization increasingly influential. Yet most systems are optimized for accuracy and engagement, without addressing the question of what recommender systems should be. When we fail to ask this question, we risk systems that amplify misinformation, reinforce biases, and create echo chambers that affect both individuals and society at large. My research takes on the normative challenge of designing recommender systems that are aligned with professional values and democratic values. News recommendation is a powerful case in point: while it plays a pivotal role in shaping public opinion, current systems lack mechanisms to include journalists’ voices, values, and perspectives. My work addresses this gap through (1) a co-design methodology for eliciting and encoding journalistic values, (2) a values-driven reranking instantiation built on SCRUF-D, and (3) longitudinal field evaluation via POPROX. This work reimagines recommender systems as not merely engines of engagement, but as accountable infrastructures that uphold democratic values.
- Spot C2What Reasoning Should a Recommender Have? Designing Reasoning for Inference-Time Scaling in Recommendation
by Wenhao DengLarge language models can gain accuracy by spending more computation at inference, most visibly through chain-of-thought, where the model proceeds step by step in language before answering. The mechanism that converts the extra computation into a better answer is what we call reasoning. How reasoning should be designed for recommendation, and in what form, is far less settled. This dissertation studies that question, so that a recommender can turn extra inference-time computation into better predictions. Whether the reasoning trace is latent, textual, or mixed is a design choice, not a prior commitment, and the work is organized around three questions. The first asks how to design the reasoning trace for a conventional sequential recommender (RQ1). My first study, RecRec (currently under peer review), gives the recommender a reasoning trace separate from the interests it predicts from, and outperforms prior reasoning-enhanced methods across four datasets. The second question (RQ2) moves to generative recommendation over semantic IDs, where the trace can be designed in more ways, and asks which design is native to the task. The third (RQ3) turns to the economics of these traces. Latent and textual forms cost very different amounts of computation, so the question is which form yields the most accuracy for a given budget, and how to compare methods at matched compute. The field has no settled recipe for any of this, and the appropriate designs are what I bring to the symposium.
- Spot D1Transparency and Control for Provider-side Recommendation Interfaces
by Elizabeth McKinnieAlthough providers are (sometimes) recognized as key stakeholders in a recommendation eco-system, there remains a lack of research centered on evaluation, transparency, and control for (with) providers. My research focuses on advancing provider-centered studies of recommendation and interfaces in a social media context. I have research in progress studying content providers’ experiences with and perceptions of social media provider-centered interfaces. My planned research includes an algorithmic audit with content providers, building on the SMORES framework to create a provider-centric dashboard, and co-designing control mechanisms with content providers.
- Spot D2Discover Weakly No More: User-Centric Design Criteria for Discovery in Music Recommendation
by Omar AhmedRecommendation systems are embedded in the discovery process. They have become a widespread feature and receive a huge amount of user engagement; as a result, they have become influential mediators of exploratory paths through catalogues. This is especially the case with music recommendation where recommenders are heavily used and discovery is of high importance. Despite this, discovery as a complex phenomenon is still largely misunderstood in this research community. It is often operationalised in a variety of ways- such as serendipity and ability to reach into the long-tail- which has resulted in discovery being targeted through partial constructs. This PhD advocates for a unified, user-centric understanding of discovery that can connect these constructs to shared design criteria and evaluation protocols. We do this by conducting user studies to find how positive discovery experiences are described; this allows for a sophisticated definition of discovery (in the context of music recommender systems) to be developed. This, combined with an investigation into how current optimisation goals fall short of affording discovery, lays the groundwork for the creation of clear design criteria and evaluation protocols for future system design. Lastly, the development of a discovery-oriented music recommender system that obeys these criteria sets an example for future system developers. With a clear definition of discovery and a standardisation of design, we lay the groundwork for recommender systems that can fulfil user desires for music discovery.
- Spot E1Toward Human-Aligned Recommender Systems via Cognitive and Behavioral User Modeling
by Mallika MainaliRecommender systems research has extensively explored personalization through user modeling, often relying on static representations of human attributes such as personality traits (e.g., the OCEAN framework), under the assumption that individuals with similar traits behave consistently across contexts. However, cognitive science shows that human decision-making is dynamic, shaped by cognitive biases, contextual trade-offs, uncertainty, and situation-dependent reasoning that are not captured by these representations. As a result, existing recommendation pipelines and evaluation frameworks often fail to reflect how individuals actually make decisions, particularly where preferences are fluid and there is no single optimal choice. This dissertation investigates how recommender systems can be better aligned with individual decision-making through cognitive and behavioral user modeling grounded in cognitive science theory and empirical behavioral signals. This work pursues four research directions: (1) identifying and defining key decision-making attributes that characterize differences in user behavior, (2) investigating how these attributes can be computationally represented using both classical AI approaches and LLM-based reasoning frameworks, (3) exploring whether cognitive attributes can be inferred from behavioral interaction data using the RetailRocket e-commerce dataset, and (4) evaluating whether cognitively informed user representations improve perceived alignment, usefulness, and satisfaction through human-centered user studies. Overall, this research bridges cognitive science and recommender systems by introducing cognitive attributes as a complementary layer of user representation, moving beyond static preference assumptions toward more behaviorally grounded and human-centered recommendation systems.
- Spot E2Unbiased Recommender Systems with Implicit Feedback
by Md Aminul IslamRecommender systems typically rely on implicit feedback (e.g., clicks) to infer user preferences. However, such data is inherently prone to various biases, including position bias and popularity bias. Position bias occurs when higher-ranked items receive more interactions regardless of true relevance. Popularity bias reinforces frequent exposure of popular items while under-recommending relevant, yet less popular ones. Directly learning from such data fails to capture true user preferences, leading to suboptimal recommendations. This research focuses on mitigating position bias and popularity bias in recommender systems. Specifically, I address position bias in learning-to-rank (LTR) systems and popularity bias in collaborative filtering (CF) models and social recommender systems based on graph neural networks. My work develops methods that overcome the limitations of existing approaches to mitigating position bias and popularity bias, enabling more relevant and personalized recommendations that align with users’ preferences.
Industry
- Spot F1MDGR: Masked Diffusion Generative Recommendation
by Lingyu Mu, Hao Deng, Haibo Xing, Jinxin Hu, Yu Zhang, Xiaoyi Zeng and Jing ZhangGenerative recommendation (GR) typically quantizes item embeddings into multi-level semantic IDs (SIDs) and generates the next item via autoregressive decoding. Despite competitive performance, this paradigm inherits three key limitations from language models: (1) autoregressive decoding struggles to capture global dependencies among multi-dimensional features at different SID positions; (2) a fixed decoding order implicitly assumes all users attend to item attributes identically; (3) sequential decoding is inefficient and struggles to meet real-time requirements. To tackle these challenges, we propose MDGR, a Masked Diffusion Generative Recommendation framework that reshapes the GR pipeline from three perspectives: codebook, offline training, and online inference. (1) We adopt a parallel codebook to provide a structural foundation for diffusion-based GR. (2) During training, we adaptively construct masking supervision signals along both the temporal and sample dimensions. (3) During inference, we develop a warm-up–based two-stage parallel decoding strategy for efficient generation of SIDs. Extensive experiments on multiple public and industrial-scale datasets show that MDGR outperforms 13 state-of-the-art baselines. Furthermore, by deploying MDGR on Alibaba Group’s online advertising platform, we achieve a 1.20% increase in revenue, demonstrating its practical value. The code will be released upon acceptance.
- Spot F2LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing
by Juncheng Dong, Ding Tong, Ishan Gupta and Yuyan WangRecommendation systems fundamentally thrive on personalization, operating in a domain where “correctness” is rarely a binary truth but rather a matter of subjective human preference and sociocultural alignment. As Large Language Models (LLMs) are increasingly deployed as autonomous verifiers to evaluate complex safety and quality guidelines in these systems, they face a unique challenge: context-aware preference alignment. Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly enhanced LLM reasoning capabilities, yet these gains are predominantly indexed on objective, mathematical tasks. In this paper, we investigate the generalization of explicit reasoning to subjective, human-centric industry rubrics using production datasets. We expose a fundamental vulnerability: forcing models to apply rigid, math-centric reasoning traces to subjective tasks actively degrades verification performance. Furthermore, applying standard RLVR to rectify this triggers a phenomenon we term reasoning collapse, where the reinforcement policy prematurely truncates reasoning trajectories in favor of rapid, heuristic guessing. To resolve this, we introduce a conditional length-penalized post-training algorithm. By mathematically intertwining strict verification accuracy with bounded reasoning length, we successfully stabilize the policy, halt reasoning collapse, and achieve state-of-the-art performance on production-scale verification tasks. Finally, we establish that the efficacy of a reasoning trace is deeply coupled with its socio-linguistic framing. We present preliminary synthesis results demonstrating massive performance variance across simulated demographic personas, and propose a novel mid-training architecture that dynamically routes reasoning through contextually aligned personas. Ultimately, this work provides a scalable algorithmic patch and a long-term architectural blueprint for aligning reasoning models with the friction of real-world subjective constraints.
- Spot G1Hybrid Retrieval and Multimodal Reranking for Item-to-Item Recommendation in a C2C Marketplace
by Yusuke Shido, Ryo Watanabe, Yuta Ueno, Yuki Yada and Shinya YaginumaItem-to-item recommendations on item detail pages drive a substantial share of session conversions on e-commerce services, including consumer-to-consumer (C2C) marketplaces, where listings are user-generated, short-lived, and diverse in both images and text. Visual vector-search recommendation, in which items are encoded by a pretrained vision model and retrieved by approximate nearest-neighbor (ANN) search using a query vector, has become a common building block for such surfaces. Mercari, Japan’s largest C2C marketplace, deployed a SigLIP-based instance of this approach in production with material online gains. However, visual similarity alone cannot capture buyer intent hinged on textual attributes (e.g., brand, model number, and condition), leaving items with weak user-uploaded images, common in a C2C marketplace, poorly covered. This study extends the visual-only setup into a multimodal, two-stage similar-item recommendation system. The system has two production components: (i) a hybrid retriever that ANN-searches over both image and Japanese text embeddings, and (ii) a multimodal neural reranker that consumes visual and textual item embeddings together with other item metadata. Both components were evaluated via large-scale online A/B tests. The hybrid retriever increased purchases from the similar-item recommendation by 20.4\%, and the reranker further improved them by 12.7\%. We share architecture, feature design, and lessons learned from deploying this multimodal similar-item recommendation system.
- Spot G2Multi-Probe Zero Collision Hash (MPZCH): Mitigating Embedding Collisions and Enhancing Model Freshness in Large-Scale Recommenders
by Ziliang Zhao, Bi Xue, Emma Lin, Tianqi Lu, Mengjiao Zhou, Kaustubh Vartak, Shakhzod Ali-Zade, Tao Li, Bin Kuang, Rui Jian, Bin Wen, Dennis Van Der Staay, Yixin Bao, Xiujin Li, Chao Deng, Henry Wei, Songbin Liu, Qifan Wang and Kai RenEmbedding tables are critical components of large-scale recommendation systems, facilitating the efficient mapping of high-cardinality categorical features into dense vector representations. However, as the volume of unique IDs expands, traditional hash-based indexing methods suffer from collisions that degrade model performance and personalization quality. We present Multi-Probe Zero Colli- sion Hash (MPZCH), a novel indexing mechanism based on linear probing that effectively mitigates embedding collisions. With reasonable table sizing, it often eliminates these collisions entirely while maintaining production-scale efficiency. MPZCH utilizes auxiliary tensors and high-performance CUDA kernels to implement configurable probing and active eviction policies. By retiring obsolete IDs and resetting reassigned slots, MPZCH prevents the stale embedding inheritance typical of hash-based methods, ensuring new features learn effectively from scratch. Despite its collision-mitigation overhead, the system maintains training QPS and inference latency comparable to existing methods. Rigorous online experiments demonstrate that MPZCH achieves zero collisions for user embeddings and significantly improves item embedding freshness and quality. The solution has been released within the open-source TorchRec library for the broader community.
- Spot H1TubiFM: Unified Item, Carousel, and Search Ranking for Streaming Discovery
by Alexandre Salle, Chenglei Niu, Suchismit Mahapatra, Xiaoxiao Chen, Suvash Sedhain, Yaqi Wang, Shervin Shahryari, Saurabh Agrawal, Qiang Chen and Michael TamirPersonalized discovery systems often train separate models for item ranking, carousel ranking, and search, even though these tasks expose complementary signals from the same viewer journey: watches shape carousel and item ranking, search queries reveal intent even when they do not lead to a catalog match, and watch history helps interpret search as rewatching, continuation, or new discovery. We introduce the user story, a serialized representation that turns a user’s cross-surface history—attributes, sessions, watch events with surface and carousel context, and search events—into a single token sequence. By interleaving pretrained language tokens with domain-specific event tokens, user stories let heterogeneous recommendation and search tasks be expressed as prompted next-token prediction over a shared grammar. TubiFM is one instantiation of this approach: a Llama~3.2 1B-based model trained on user stories and prompted to rank items, carousels, or search results without task-specific architectures. In offline evaluation, this single model outperforms specialist baselines across item, carousel, and search ranking. In online A/B tests, TubiFM significantly improves search total viewing time (TVT) by +3.9% and carousel TVT by +0.30%. Item ranking is statistically neutral on TVT (+0.14%), but matches a mature production stack; across all three tasks, TubiFM serves on L40S GPUs and reduces p99 ranking latency from 500ms to 200ms. These results show that shared user stories can improve discovery while simplifying ranking systems.
- Spot H2PinDCO: Whole-Page Aware Dynamic Creative Optimization at Scale
by Yu Hao, Yuchun Li, Peimeng Sui, Meilin Liu, Tianyuan Cui, Hao Li, Zicong Zhou and Akanksha BaidRecent advances in generative AI have substantially accelerated the creation of high-quality ad creatives, dramatically expanding the number of candidate variants per campaign. This shift increases the need for scalable dynamic creative optimization (DCO) systems that can match creatives to the most relevant audiences under stringent latency and cost constraints. We present PinDCO, a production DCO system for ad creative retrieval and selection on Pinterest, a billion-scale visual discovery platform. PinDCO is built around a Creative Component Fusion Network (CCFN) that performs dynamic creative scoring by modeling each creative component (e.g., image, title, layout) with a dedicated tower, using component-specific hyperparameters to account for differing modeling complexity. The component representations are fused to predict a creative-level score conditioned on the ad-level pre- diction, and we improve training data quality via an exploration- exploitation strategy. To account for Pinterest’s waterfall grid layout, where a creative’s rendered size affects nearby content and session-level engagement, we introduce a Pixel-aware Adjustment Module(PAM) that adjusts scores based on creative size to encourage efficient screen real- estate utilization and better whole-page outcomes. To support the large volume of creative candidates, we further employ a light- weight pre-selection model for early pruning, and optimize serving efficiency through caching and dynamic batching. Extensive offline analyses and online A/B experiments demonstrate the effectiveness of PinDCO, yielding a +3.09% lift in ad Click-Through Rate(CTR) with positive whole-page metrics. With the strong performance, we launched PinDCO in the Pinterest Ads platform.
- Spot O1Advancing Relevance Measurement with Vision–Language Models for Web-Scale Search
by Han Wang, Alex Whitworth, Pak Ming Cheung, Zhenjie Zhang, Krishna Kamath, Xi Chen, Roberto Konow and Kurchi Subhra HazraRelevance evaluation plays a crucial role in personalized search systems to ensure that search results align with a user’s queries and intent. While human annotation is the traditional method for relevance evaluation, its high cost and long turnaround time limit its scalability. In this work, we present our approach at Pinterest Search to automate relevance evaluation for online A/B experiments using fine-tuned Vision-Language Models (VLMs). We rigorously validate the alignment between VLM-generated judgments and human annotations, demonstrating that VLMs can provide reliable relevance measurement for experiments while greatly improving the evaluation efficiency. Leveraging VLM-based labeling further unlocks the opportunities to expand the query set, optimize sampling design, and efficiently assess a wider range of search experiences at scale. This approach leads to higher-quality relevance metrics and significantly reduces the Minimum Detectable Effects (MDEs) in online experiment measurements.
- Spot O2Multilingual Semantic Retrieval for Apple Music Search
by Vishalaksh Aggarwal, Kevin Sebastian, Vivek Kanojiya, Leo Le, Nick Tucey and Santosh ShankarApple Music serves listeners across 150+ storefronts in dozens of languages, with a catalog that grows by 100,000+ new tracks daily. At this scale, search recall on misspelled, transliterated, and cross-lingual queries becomes a dominant driver of session quality, particularly for tail queries that account for the majority of unique queries. We present a multilingual semantic retrieval system built on a 305M-parameter Siamese bi-encoder fine-tuned from GTE-multilingual-base with curriculum-scheduled multi-objective training. The model is integrated into the search stack via a hybrid retrieval architecture that blends dense nearest-neighbor results with the existing token-based index using quantile distribution matching, enabling deployment without retraining downstream rankers. Offline, the model achieves a 69% relative improvement in Hit@10 over GTE-multilingual-base. In a worldwide online A/B test, the system delivers a 2.28% relative conversion-rate (CR) lift overall, an 86% reduction in the no-result rate, and gains across every storefront with no observed regressions. The improvement is concentrated where it is needed most: tail queries see a 7.93% relative CR lift, compared with 0.89% for mid-frequency queries and 0.14% for head queries—evidence that semantic retrieval improves recall on hard queries without disturbing well-served popular ones. To our knowledge, this is one of the largest search-quality improvements deployed on the platform.
- Spot P1From M Passes to One: A MapReduce-Style Bootstrap for Pointwise Off-Policy Evaluation in Production Recommenders
by Ghulam Ahmed Ansari and Cheng DingWe present Bootstrap IIPS@K, a pointwise off-policy estimator for production recommender rankers built around a MapReduce-style bootstrap factorization. The factorization constructs all bootstrap pseudo-samples through per-tile binary masks in a single read of the logged data, collapsing -pass variance estimation to one pass at billion-row scale. Bootstrap IIPS@K runs directly on production logs and removes the dedicated random-session logging branch that prior industrial estimators required. The factorization com- poses with any bounded per-session statistic, which lets structural sequence-dependence assumptions be compared empirically from the same single read; on production logs the pointwise prior stays bounded through slate depth = 60 while the cascade prior di- verges at = 10. A non-parametric quantile calibration of model scores keeps Bootstrap IIPS@K stable across multi-objective re- weights between ranker generations. Deployed on two LinkedIn Feed surfaces, Bootstrap IIPS@K reaches 70% offline–online par- ity (16/23 promotions) and Kendall’s = +0.29 (= 0.04) on the broader 26-pair rank-correlation pool; the Li et al. (WSDM 2011) baseline on the same workload shows near-random rank correlation with online lifts ( =−0.13).
- Spot P2DART – Drift Aware Ranking with forecasTing and uncertainty
by Deep Nayak, Shreyas S and Sivaramakrishnan KaveriRanking systems in online platforms across domains must handle temporal drift where artifact distributions and contextual relevance evolve over time. Existing approaches address temporal drift using recency weighting, sliding windows, or forecasting techniques, however, they do not focus on integrating drift signals with ranking systems or handle the data variability introduced by the drift. We propose DART, a two-stage architecture combining a ranking model with a forecasting model that predicts artifact performance for a look-ahead window. An uncertainty-aware attention mechanism dynamically weights forecasted information based on prediction confidence, enabling robust adaptation without online updates to the ranker at serving time. Prior work uses uncertainty either to detect drift or to debias ranking; DART instead uses it as a control signal that modulates information flow from a forecaster to the ranker. Experiments on a large-scale e-commerce promotion ranking dataset show DART’s deployable offline configuration achieves +11.11% relative Mean Reciprocal Rank (MRR) during promotional events and +7.37% pre-event. On the public Coveo SIGIR dataset the same architecture delivers +22% MRR in both high and low-drift periods, showing dual-regime robustness across abrupt and gradual drift; the mechanism is task-agnostic and also improves a regression benchmark. Ablations demonstrate that attention-based integration of forecasted information substantially outperforms direct feature inclusion, reinforcing the importance of uncertainty-aware temporal adaptation in ranking systems.
- Spot Q1LO-FAR: A Cost-Aware Local Filter for Sparse Feature Ranking in Industrial Ad Recommendation
by Egemen Erbayat, Luis Duque, Sohini Roychowdhury, Mohammad Amin and Srihari ReddyIndustrial ad recommendation models rely heavily on sparse, high-cardinality ID-list features that encode user histories and contextual identifiers. Each is backed by a dedicated embedding table, so these features dominate storage, training, and serving cost and must be revisited as traffic and downstream models evolve. Therefore, sparse feature ranking is not just an offline modeling problem but also a recurrent systems decision limited by compute budgets and iteration cadence. We present Localized Feature Ranking (LO-FAR), a CPU-only, model-agnostic workflow that ranks each candidate feature from its stand-alone held-out predictive signal using lightweight local estimators rather than the GPU-bound retraining loops of permutation- and stochastic-gate-based methods. On a production-grade dataset of more than one million logged interactions and 475 sparse ID-list features, LO-FAR completes ranking in approximately two CPU-hours and preserves downstream Normalized Entropy gains on CTR and CVR tasks that match or exceed shuffle-based importance, Binary Stochastic Neurons, and a coverage-based heuristic across budgets of 100–400 retained features. These budgets correspond to deterministic 40–75\% reductions in sparse storage. The contribution is a deployable workflow that shows how, under realistic cost and turnaround constraints, a simple local filter can be a stronger production choice than heavier interaction-aware alternatives.
- Spot Q2Building a User Foundation Model for the Open Web
by Ivan Can Arisoy, Solal Vernier, Merwan Barlier and Blaž ŠkrljUser foundation models have demonstrated strong results in e-commerce and social recommendation, but most industrial deployments assume environments where user identity is stable and persistent. Open-web real-time bidding (RTB) operates on a structurally different data distribution: user identity is fragmented and non-persistent across browsing sessions, and the availability of browsing history depends on user privacy choices. Consequently, a significant portion of traffic carries no historical data, and available records often consist of relatively short, disjointed sessions. As a result, historical signals in this domain are typically represented as aggregated counters and recency buckets, leaving the sequential structure unexploited. To address this limitation, we present a user foundation model that applies self-supervised learning on user browsing histories and show that the learned representation improves multiple downstream production tasks, demonstrating the viability of this approach on the open web. We pre-train a Transformer encoder with masked language modeling and a sequence-level contrastive objective, then fine-tune it on the CTR prediction task. We optimize the encoder’s pre-training pipeline with an LLM-in-the-loop search over a curated catalog of reviewable, code-level edits (lifters), instantiating the LLM-as-optimizer paradigm in an industrial setting. The same encoder representation yields +1.197\% RIG on the production bid win-rate model and +1.354\% RIG on the production CTR ranker; a 12-day live A/B test confirms +1.29\% CTR, -0.89\% eCPC (80\% CI excluding zero on both metrics).
- Spot V1EGR: Embedding-Native Generative Retrieval with a Shared LLM
by Xiaodong Liu, Congfei Zhang, Hsiang-Wei Chao, Siman Wang, Tong Zhao, Xiao Bai, Vincent Zhang, Jingxiao Ma, Zhe Liu, Wenfeng Zhuo, Zichu Li, Jitin Krishnan, Yunzhi Zhou, Yajun Wang, Jinchao Li and Yu ZhangGenerative retrieval is increasingly popular in large-scale recommendation and advertising systems, yet current methods introduce practical complications. Semantic-ID methods rely on quantization, mutable identifier vocabularies, and token-to-item grounding; embedding-based pipelines train the item encoder separately from the query generator, which limits user-item alignment. We propose EGR, an Embedding-native Generative Retrieval framework for recommendation and advertising. EGR uses a single shared LLM to learn item representations from item metadata and user representations from interaction histories in one embedding space. Items are indexed directly as dense vectors, and user histories are encoded as dense retrieval queries. Joint contrastive training groups related items and aligns queries with their target items. We evaluate EGR on public benchmarks, industrial data, and live deployment. Besides outperforming published baselines on Amazon Reviews, on Snap DPA, EGR scales with data, handles cold-start items, and benefits from multimodal input. In production, EGR delivers a +3.12% conversion-rate lift, simplifying system design while improving retrieval quality and ad performance.
- Spot V2Long-term User Engagement Optimization through Model-agnostic Downstream Rewards Learning
by Derek Wang, Filip Ryzner, Kelly He, Armando Ordorica, David Woo, Aditya Mantha, Liyao Lu, Usha Amrutha Nookala, Haoran Guo, Jiacong He, Olafur Gudmundsson, Matt Chun, Krystal Benitez, Dhruvil Deven Badani and Yijie Dylan WangAs recommender systems mature in the past few years, their optimization objectives have evolved from a primary focusing on short-term behavioral signals to a broader emphasis on long-term user engagement and retention. However, directly optimizing retention is difficult because return signals are sparse, delayed, and only partially attributable to earlier recommendations. Prior work has addressed this challenge with sequential modeling and reinforcement learning, but these approaches typically require task specific reward engineering, substantial computational overhead, and surface specific implementations that are difficult to generalize. In this paper, we present a unified, model-agnostic downstream reward framework for optimizing long-term user value in large-scale recommendation systems. First, we formulate the downstream reward learning problem and develop an offline screening framework to identify session level behaviors that are both observable early and predictive of future retention. We then propose several model-agnostic downstream rewards signals derived from observed user action patterns across multiple sources. We further discuss the engineering effort to productionize the proposed rewards derivations and challenges we faced when adding them to our ranking models. Online A/B experiments demonstrate consistent improvements in engagement and retention-related metrics, and the framework has been deployed across multiple Pinterest surfaces, including Homefeed, Related Pins, Search, and Notifications.
- Spot W1HIPA-Net: Heterogeneous Intent Projection Alignment for Feedback Calibration in CVR Prediction
by Mengqing Ye, Jianwei Zhai, Bobo Cheng, Fei Pan and Peng JiangLead advertising is a contact-oriented conversion scenario in industrial recommender systems, where the target action is to initiate follow-up communication rather than complete an immediate transaction. A key challenge in lead CVR prediction is that implicit drop-off behaviors are frequent and informative, but their feedback polarity is not directly observed. Treating such drop-off evidence as fixed negative feedback or ordinary noisy implicit feedback can distort user contact intent. We formulate lead CVR prediction as a semantic feedback calibration problem for unfinished contact behaviors. To address it, we propose HIPA-Net, a Heterogeneous Intent Projection Alignment Network. HIPA-Net learns candidate-aware representations for fine-grained positive, negative, and drop-off feedback behaviors, constructs explicit positive and negative intent anchors from reliable feedback, and calibrates implicit drop-off evidence through anchor-specific projection alignment. HIPA-Net has been deployed in a large-scale industrial lead advertising system. Offline experiments show that HIPA-Net outperforms representative behavior modeling, multi-behavior, feedback-aware, and denoising baselines. Online A/B tests further show +2.184% CTR, +5.426% RPM, and -11.521% user-side negative feedback rate, demonstrating the effectiveness of anchor-guided semantic calibration for contact-oriented recommendation.
- Spot W2Towards Generalizable and Efficient Large-Scale Generative Recommenders
by Qiuling Xu, Ko-Jen Hsiao and Moumita BhattacharyaGenerative recommendation models can model user behavior as sequences of events and provide a shared backbone for multiple recommendation tasks. In production, however, pre-training gains do not automatically translate into downstream application improvements: task headroom, repeated-training cost, serving latency, and item freshness all affect transfer. We describe our experience scaling a generative recommender from 2M to 1B backbone parameters, excluding embedding and decoding layers, in a production-scale title recommendation setting. Across multiple downstream tasks, we observe task-dependent scaling behavior: some tasks approach an empirical ceiling within the observed scale range, while others continue to benefit from additional capacity. This motivates using offset scaling-law fits as a diagnostic for where additional model scale may be more or less useful. We then study production constraints that arise when applying the model in practice. Frequent retraining over trillions of behavior tokens makes training and decoding efficiency important; cached serving can make the immediate next-token target stale; and newly launched titles may need to be scored from semantic metadata before collaborative ID embeddings are reliable. We address these issues with multi-token prediction for serving-latency alignment, sampled softmax and a projected decoding head for efficient repeated training, and semantic item towers with collaborative-embedding masking for cold-start adaptation. In a one-week production-shadow evaluation over 1M users, the 1B-backbone model achieves higher MRR than the 2M-backbone baseline across all reported tasks, including a 22.5\% relative gain on the lower-predictability Task A. Overall, the results support treating model scale as one component of a production transfer problem, alongside task headroom, decoding cost, serving-latency alignment, and item generalization.
R&P Notes
- Spot I1Predicting Custom-Feed Returns for New Bluesky Posts: A Prospective Study
by Yipeng Wang and Mohit SinghalThe conventional approach to cold-start recommendation addresses new users or newly introduced items. Bluesky custom feeds create a different setting: independently operated feeds filter content from a shared public stream. In this setting, newly published posts are the cold-start objects, while the feeds serve as candidates. We propose a cold-start routing task in which a newly ingested public post is the query and all rankable feeds in the monitored panel are ranked according to whether each will subsequently return it. We build a still-evolving collect-first, label-later benchmark dataset. The collected dataset covers a fixed panel of 5,000 monitored feeds and contains 17.804 million public posts, 1.865 million observable post–feed return records, and 625,083 valid feed polls. The labels record whether a post is observed among a feed’s AppView Top-50 results in at least one poll during the 24 hours after publication. The current experiments use two disjoint 24-hour test folds, each paired with a 24-hour training window and separated by a 24-hour outcome-availability gap. Evaluation is conditional on the 602,186 test posts that have at least one positive observed label and satisfy the metric eligibility criteria; these posts account for 9.04% of all 6,661,658 test posts. Across the two folds, LambdaRank achieves the best equal-fold mean values among the evaluated models: 0.7361 for capped Recall@10, 0.6127 for NDCG@10, and 0.7749 for Hit@10.
- Spot I2Whose Posts Get Ranked: Identical-Text Exposure Gaps in Bluesky Custom Feeds
by Yipeng Wang and Mohit SinghalBluesky lets users deploy custom feeds, independently operated recommendation algorithms that the platform serves alongside thousands of others. This paper investigates how evenly these feeds treat posts with the same text. To measure this, we take repeated snapshots of the Top-50 lists that 1,366 public feeds return, and we group posts with identical text, from different authors, that were created before the same list response and closely matched in age. Exposure diverges widely inside these matched sets, which span 250 feeds: in 33% of sets, one copy appears on the list while another does not. Fixed-effects regressions show that this divergence is associated with the author’s history on the specific feed. Authors new to a feed receive less exposure for the same text (−0.061 in reciprocal-rank weight), while authors whose posts the feed has returned before receive more. A new author with more followers than the competing author still loses 74% of head-to-head comparisons. Media and post-type features show no detectable association after multiple-comparison correction. These results are early evidence that access to many independent feeds is not enough to give identical texts equal exposure.
- Spot J1Why Do CF and Review Models Disagree? User-Specific Rating–Review Scale Gaps
by Min-Gyu Jang, Wooseung Kang, Minje Kim, Gun-Woo Kim and Sang-Min ChoiRatings provide a concise summary of user preference, whereas reviews express more selective opinions, attitudes, and product aspects. Given these differences in the feedback they capture, rating-based collaborative-filtering and review-based models may produce different predictions for the same user-item interaction. Although such disagreement can make the choice of model consequential, aggregate prediction accuracy does not explain why the models disagree or which user characteristics are associated with the disagreement. We examine whether these differences can be explained by user-specific discrepancies between rating tendencies and review tone. To represent this discrepancy, we introduce the rating–review scale gap (RRS gap), which compares a user’s relative rating leniency with their review-tone bias. Our results show that large disagreement alone does not indicate which model is more accurate. However, the RRS gap is consistently associated with whether the review-based model predicts higher or lower than the rating-based model, although the strength of this relationship varies across domains. These findings suggest that systematic differences between how users assign ratings and express opinions in reviews can help explain prediction disagreement between rating- and review-based recommender models. Our source code is available at https://anonymous.4open.science/r/recsys2026CDF5/.
- Spot J2Towards Dynamic Relationship Schema Discovery for Complementary News Recommendation
by Kai SugaharaComplementary relationships are commonly used to represent cross-item dependencies in recommendation systems such as e-commerce, where one item functionally complements another. This concept also applies to news recommendation, where complementary relationships capture likely follow-up articles that users read after a given article, such as background explainers or timeline updates. Existing approaches often do not explicitly model document-level complementary relationships and typically rely on fixed schemas, limiting adaptation to evolving news-consumption patterns and potentially reducing recommendation quality and interpretability. This study proposes Dynamic Relationship Schema Discovery (DRSD), an iterative framework that identifies schema-gap article pairs, uses an LLM to propose and annotate new relationship labels, retrains a relationship-conditioned CTR model, and prunes labels that fail to improve validation performance. Preliminary experiments show that DRSD increases Catalog Coverage and filters weak relationship hypotheses through validation-driven pruning while maintaining competitive recommendation quality. These findings suggest that DRSD is a promising direction for refining complementary relationship schemas in news recommendation.
- Spot K1Exploring Revealed Behavior and Stated Profiles in Longitudinal News Recommendation
by Thomas Elmar Kolb, Alain Dominique Starke and Christoph TrattnerNews recommenders increasingly need to balance relevance, diversity, and stated preferences over repeated interactions. We report preliminary evidence from a four-wave user study with 262 participants comparing four personalization trajectories in a content-based news feed. Phase-wise comparisons suggest that on-profile selection varied with the assigned trajectories, while stated topic profiles remained comparatively stable during the two-week study (92.6%–98.0% retention). These preliminary findings suggest that short-term selection behavior partly reflects feed composition rather than stable preference change. Ongoing mixed-effects, exposure-normalized, and attrition-sensitive analyses will test this interpretation.
- Spot K2End-to-End User and Item Embeddings: Specializing Retrieval Representations for Ranking
by Simon Rauch, Vito Bellini, Anton Thielmann, Adrian Gruszczynski and Giuseppe Di BenedettoIndustrial recommender systems typically operate in two stages: retrieving a candidate set from a large catalog, then ranking those candidates using contextual information. The ranking stage relies on features that summarize a user’s prior interactions with the system. These features are often carefully hand-crafted, and designing and maintaining them is time-consuming and computationally expensive. In this paper we propose leveraging the user and item embeddings already produced in the retrieval stage as input features to the downstream ranking model, and fine-tuning them for the ranking objective. We motivate this direction with A/B test results from a large-scale music streaming platform, where retrieval-stage user embeddings improve ranking performance even without fine-tuning. Our results suggest that representations learned for retrieval can reduce the reliance on expensive hand-crafted ranking features, and that fine-tuning them offers a promising path to further gains.
- Spot L1Blended User Embeddings for Steerable Matrix Factorization
by Andrea Pisani, Luca Pagano, Giuseppe Vitello, Maurizio Ferrari Dacrema and Paolo CremonesiWe propose BLUE Steer, a Matrix Factorization-based Recommender System designed for interpretability and steerability. Given user profiles and textual item descriptions, we prompt an inference-only Large Language Model to produce User Summaries. These are embedded and combined to pre-trained User Factors from a Matrix Factorization model. Our preliminary investigation shows that this approach improves ranking accuracy and makes this lightweight model steerable by design, as editing the textual User Summaries predictably influences downstream recommendations. Code is available at https://anonymous.4open.science/r/BLUE-Steer-34E1/.
- Spot L2Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson
by Tanay Chowdhury and Saeideh Shahrokh EsfahaniWe report a practical lesson from building a GPU-free explainable-recommendation serving stack: explanations are pre-generated offline into a per-item candidate pool, and a small CPU-resident model selects one at request time. Every candidate in the pool of size carries an offline BERTScore-F1 label against a reference explanation, so we compare a pairwise learning-to-rank model (LightGBM LambdaRank) against a single-action RL formulation (DPO) and a distilled selector on the XRec Google Local benchmark (2,958 pairs, 5 seeds). LambdaRank outperforms both by 0.019–0.025 F1, a gap more than fifteen times the across-seed standard deviation. The advantage appears structural rather than algorithmic: LambdaRank’s objective consumes the label of every candidate in the pool, whereas DPO’s rollout only observes the label of the sampled side of each preference pair. A separate candidate-source design, selecting from knowledge-graph-grounded paths instead of the cached pool, trades this reference alignment for a Unique-Sentence Ratio (USR) of 1.000. We also encountered two failure modes: further RL fine-tuning on top of an already-distilled policy regresses F1, and an end-to-end RL fine-tune of the generator itself reward-hacks the metric within a few hundred steps. The practical takeaway is that whenever an offline metric can label every candidate in a fixed action set, a pairwise learning-to-rank baseline is worth evaluating before reaching for single-action RL. One caveat: LambdaRank’s training signal differs in form from the evaluation metric we report, since it trains on quintile-binned labels while DPO trains on a blended reward, though both derive from the same underlying BERTScore-F1; a fully decorrelated evaluation remains future work.
- Spot M1How Faithful Is the Reasoning of LLM Recommenders? A Counterfactual Audit
by Arpita ShahLarge language model (LLM)-based recommenders increasingly generate reasoning alongside their recommendations \cite{}, implicitly suggesting that the stated reasoning determines the selected item. We investigate whether these explanations exhibit causal dependence or are post-hoc rationalizations. We introduce a counterfactual audit that edits a model’s own reasoning trace using four operators that (1) flip the stated preference, (2) swap in another user’s reasoning, (3) erase it, or (4) paraphrase it as a control., and then measures whether the top-1 recommendation changes. We define the \emph{faithfulness gap} as the mean top-1 change rate under the three meaning-changing edits (flip, swap, erase) minus the change rate under paraphrasing. We evaluate the same 100 users from the Amazon Video Games dataset using GPT-5 Chat, GPT-5 Reasoning, and DeepSeek-R1. GPT-5 Chat exhibits substantial reason-responsiveness, with a faithfulness gap of $0.57$. In contrast, DeepSeek-R1 and GPT-5 Reasoning yield gaps of $0.02$ and $-0.11$, respectively, providing little evidence that their displayed explanations influence their recommendations beyond sensitivity to the control edit. As a follow-up diagnostic, we apply the same audit to DeepSeek-R1’s \texttt{
} trace and obtain a gap of $0.30$. This result suggests that, for DeepSeek-R1, causal dependence is more strongly associated with the internal reasoning trace than with the explanation presented to the user.
- Spot M2The Illusion of Control: Commensurability, Concentration, and Correlation in Multi-Objective Recommendation
by Patrik Dokoupil and Ladislav PeskaMulti-objective recommender systems (MORS) or recommendation re-ranking approaches often combine diverse goals–such as relevance, novelty, and diversity–via linear scalarization, while objectives’ weights are being exposed as tunable hyperparameters for system optimization or as interactive knobs supporting user control. However, we argue that such an approach creates an \emph{illusion of control}, in which some objectives mechanically dominate the recommendations, and standard normalization techniques do not alleviate the issue. We formalize this through the \emph{3C Bias Framework} (Commensurability, Concentration, and Correlation) and introduce the \emph{Objective Dominance Ratio} (ODR) to diagnose it. We benchmark four widely used normalization schemes on the ML-25M, Steam, and Goodbooks-10k datasets and show that each normalization fails at a specific C, making this a precise diagnostic toolkit. Finally, we propose \emph{Rank-ZCA} normalization, which is–to the best of our knowledge–the only strategy resilient against all 3C biases.
- Spot N1SERA: A Multi-Dimensional Evaluation Framework for Retail Agentic Shopping Assistants
by Sowmya Podila and Bo ShenRetail agentic shopping assistants pose evaluation challenges that existing conversational recommender system (CRS) frameworks do not address: they must maintain constraint context across turns, refuse adversarial requests, protect privacy-sensitive information, and recover when users correct preferences mid-conversation. Recent work in adversarial LLM safety has shown that safety be- haviours validated at the single-turn level degrade under multi-turn pressure, yet no integrated evaluation framework combines safety, compliance, and recommendation quality assessment at the conversation level. We propose SERA (Suite of Evaluation metrics for Retail Agents), a three-dimensional metrics taxonomy, derived from practitioner requirements, of 21 LLM-as-judge metrics spanning Safety, Compliance, and Agent Quality, each with single-turn and multi-turn variants. A risk-proportional aggregation design applies worst-case scoring for safety and compliance metrics, and mean scoring for quality metrics. We describe a persona-driven conversation simulator for offline stress testing and report preliminary reliability analysis on ≈650 seed queries across three categories, simulated into multi-turn conversations with a retail agent under development. Early deployment as a CI/CD gate has produced actionable feedback for agent development; full ablation studies and formal human-annotation campaigns remain future work.
- Spot N2Do SID Alignments Help Generative Recommendation?=
by Shiteng Cao and Zhiheng LiSemantic-ID-based generative recommendation represents each item as a short token sequence and trains an autoregressive model to predict the next item ID. This note studies whether Semantic ID (SID) tokens should be aligned to natural language or specialize into collaborative codes. We report two preliminary studies. First, OpenOneRec representation probes show that SID-to-text explanations are template-sensitive, while SID tokens form clean subspaces separated from ordinary text. Second, on sampled offline logs from a Chinese short-video platform serving hundreds of millions of daily active users, a Xavier-initialized SID model obtains lower validation loss and stronger mid-/large-K HitRate than an alignment-initialized model when the input contains only a system prompt and SID sequence. We further summarize practical failure modes including prompt overfitting, general instruction degradation, SID over-generation, and task trade-offs under model merging. These findings suggest that SID alignment is conditional rather than universally beneficial.
Short Papers
- Spot R1Assessing Sentiment Semantics in KG-Based Explainable Recommender Systems
by Alejandro Ariza-Casabona, Pol Pastells, Ludovico Boratto and Maria SalamoSentiment information can enhance path-based explainable recommendation, yet advances are hindered by the lack of public benchmarks for sentiment-aware knowledge graph reasoning. We address this gap by systematically evaluating how different sentiment representation strategies and reward mechanisms affect utility, beyond-utility, and explanation quality across four real-world datasets under our open-source SAKG reasoning benchmark. Our exploration reveals: (1) sentiment-aware methods improve accuracy and diversity but reduce provider fairness, exposing a consistent trade-off; (2) under extreme sentiment imbalance, adaptive entropy-weighted meta-relations improve robustness in some domains, while naive sentiment sharing across heterogeneous relations causes negative transfer with considerable utility drop; (3) sentiment path consistency rewards enhance recommendation quality and explanation coherence, whereas terminal-only sentiment constraints improve fairness at modest utility cost. Our findings and open benchmark provide actionable guidance for designing sentiment-aware explainable recommenders and a foundation for future research in this underexplored space.
- Spot R2When Fixed-Candidate Offline Evaluation Changes Model Selection in Two-Stage Recommenders
by Teresa ZhangTwo-stage recommenders use a first-stage generator to decide which items can be considered and a second-stage reranker to order the items that survive. Offline evaluation often removes this first-stage variation by scoring all systems on a shared candidate pool. That protocol is appropriate for measuring reranking quality on a fixed action space, but it changes the model-selection problem when compared pipelines differ in their generators, candidate budgets, or retrieval histories. We formalize the distinction between full-pipeline utility on a pipeline’s own candidates and reranking utility on an anchor-defined candidate pool, and we decompose the former into candidate survival and conditional ranking quality. In controlled registries built on ten public benchmark instances, fixed-candidate evaluation selects a different top pipeline from own-candidate evaluation in eight cases under the dataset-specific main anchor. On OTTO, Dressipi, and Diginetica, the top-1 disagreement persists under both tested K=20 anchors: OTTO and Dressipi show sign reversals, while Diginetica shows suppression of the own-candidate advantage large enough to change the selected winner. These results show that fixed-candidate scores are useful reranking diagnostics, but they should not be treated as full-pipeline model-selection scores when first-stage retrieval differs across systems.
- Spot S1Structure-Preserving Projection for Mitigating Modality Bias in LLM-Based Sequential Recommendation
by Tzu-Wei Chiu, Song-Duo Ma, Hsin-Yu Lin and Pu-Jen ChengRecent LLM-based recommenders integrate textual and collaborative signals by projecting collaborative embeddings into the embedding space of the LLM. However, this projection can introduce modality bias that distorts the underlying collaborative structure and limits the usefulness of projected embeddings. To address this issue, we propose a novel structure-preserving projection approach that maintains the relational geometry of collaborative embeddings through dedicated structure-preserving losses. Comprehensive experiments demonstrate that our approach consistently improves recommendation performance, providing a more reliable path for LLM-based recommendation.
- Spot S2Distribution-Level Contrastive Supervision for Generative Recommendation
by Ziqi Xue, Dingxian Wang, Yimeng Bai, Shuai Zhu, Jialei Li, Xiaoyan Zhao, Frank Yang, Andrew Rabinovich, Yang Zhang and Pablo N. MendesRecent generative recommenders improve scalability by retrieving items through token generation instead of traditional ranking over large candidate sets. Yet their training signals are still dominated by discrete code prediction, which overlooks the soft assignment information naturally produced by the tokenizer. This mismatch limits semantic transfer from the tokenizer to the recommender and may hurt overall optimization. We tackle this limitation by introducing a distribution-based supervision scheme for generative recommendation, where multi-level codebook probabilities are treated as soft semantic targets. On top of this design, we develop SODA, a plug-and-play alignment framework that adopts a BPR-style contrastive objective to align recommender representations with target-side distributional representations against negative ones. The proposed method enriches training with finer semantic cues while leaving the decoding stage unchanged. Experimental studies on multiple real-world benchmarks demonstrate that SODA consistently strengthens diverse generative recommendation architectures. Code will be available after acceptance.
- Spot T1Surface Matching in Skill Recommenders: Incomplete Evaluation Hides Simplicity Bias
by Warre Veys, Matthias De Lange, Jens-Joris Decorte, Chris Develder and Thomas DemeesterSkill-extraction encoders are typically evaluated with ranking metrics scored against human-annotated positives. At the scale of ESCO (13,939 skills), annotation is inherently incomplete, and unannotated labels are treated as uniformly wrong irrespective of semantic proximity. We show that the nowadays popular contrastive bi-encoder setup is prone to exhibit simplicity bias, with a tendency for exploiting surface-level lexical overlap between input and skill label as a shortcut. It appears that annotated positives and such shortcuts often coincide, such that ranking metrics remain high even when other top predictions are potentially semantically nonsensical. To analyze this phenomenon at scale, we use an LLM-as-a-judge to flag nonsensical unannotated predictions, validated against human labels. We then study two training-time mitigations: label augmentation with low-overlap positives, and hard negatives targeting the shortcut which reduce judged nonsense while preserving ranking performance. Our findings show how incomplete annotations can silently degrade skill recommender quality, and give practical diagnostic and training-time tools for this setting, enabling us to release a more robust fine-tuned model to the community.
- Spot T2Scaling LLM-Enhanced Linear Autoencoders to Industrial-Size Catalogs
by Maxim Skurikhin, Kiryl Liakhnovich and Oleg LashininLLM-enhanced linear autoencoders (L3AE) incorporate semantic item representations from large language models into collaborative filtering and demonstrate significant gains on long-tail items. However, L3AE requires dense n × n matrices, which makes it impractical for large catalogs – exactly where semantic enrichment would help most. We propose scalable formulations that factor the semantic similarity matrix into a compact low-rank form and compute all required precision-matrix products without ever building any dense n × n matrix. Our method preserves the algebraic structure of L3AE while replacing its dense collaborative precision with SANSA’s sparse approximate inverse, enabling efficient scaling to industrial-size catalogs. Experiments on several public datasets show consistent gains over multiple baselines, and these gains tend to grow in colder, sparser catalogs – a pattern consistent with semantic enrichment helping most where collaborative signals are weakest. Code and configs are available at https://anonymous.4open.science/r/scal3ae
- Spot U1GradSup: Gradient Superposition for Personalised and Scalable LLM Recommendation
by Kanishka Dandeniya, Chirath Dasanayaka, Daswin De Silva, Sam Saltis, Shalinka Jayatilleke, Nishan Mills and Harsha MoraliyageLarge language models (LLMs) have demonstrated strong capabilities in recommendation tasks such as item, sequence, conversational recommendation, and explanation generation. However, LLM weights are typically shared across all users. Adapting these models to individual users remains a fundamental challenge that requires millions of trainable parameters, even when using finetuning methods such as Low-Rank Adaptation (LoRA). Building such personalised adapters would also require large volumes of storage and high computational overhead. To address this challenge of personalised and scalable recommendation, we propose the Gradient Superposition (GradSup) method. Built on the TinyLoRA architecture, GradSup is a closed-form method that computes per-user adapters in a single pass, achieving 15x speedup over iterative fine-tuning at matched parameter budget and 400x compression over full LoRA. Experiments conducted on movie, book and electronics benchmark datasets demonstrate that GradSup outperforms iterative fine-tuning to provide scalable and personalised recommendation that is consistently sustained above the frozen LLM backbone. Code is available at: https://anonymous.4open.science/r/GradSup-9860
- Spot U2Dual Conditional Diffusion for Generative Cross-Domain Recommendation via Disentangled Knowledge Transfer
by Abhradeep Datta, Arin Gupta, Varun Tej Kasula and Ashok Singh SairamCross-domain recommendation (CDR) mitigates data sparsity by transferring knowledge across domains, and dual-target CDR further enables bidirectional transfer to improve both domains simultaneously. However, existing methods largely rely on linear augmentation or reconstruction, which may generate semantically inconsistent representations and noisy supervision. To address these limitations, we propose DCD-XRec, a Dual Conditional Diffusion framework that models cross-domain knowledge transfer as a structured and non-linear generative process. The model first learns expressive user embeddings with graph-based encoders and then decomposes them into shared and domain-specific components using learnable projection heads. On this basis, a bidirectional conditional diffusion module synthesizes target-domain embeddings conditioned on source-domain representations, enabling semantically coherent transfer. In addition, we introduce a dual diffusion objective with two complementary pathways: a real-to-real pathway for stable alignment and a real-to-augmented pathway for richer generative supervision. Experiments on multiple public benchmark datasets show that DCD-XRec consistently outperforms strong baselines, demonstrating the effectiveness of diffusion-based generative modeling for dual-target cross-domain recommendation.
RecSys 2026 (Minneapolis)
- About the Conference
- Registration
- Program at Glance
- Program
- Call for Contributions
- Challenge
- Keynotes
- Accepted Contributions
- Presenter Instructions
- Workshops
- Tutorials
- Committees
- Inclusion
- Student Volunteers
- Women in RecSys
- Visa Information
- Addressing Attendance Issues
- Location / Hotel
- Lasting Impact Award
- 60 Milestones for 20 Years




















