Day 3 Posters: Industry + Reproducibility + R&P Notes + Short Papers + TORS

Date: Thursday October 1
Posters are available from 10 am to 4 pm, and are presented by the authors during the coffee breaks.

Industry

  • Spot A1Personalizing Incremental Video Search with Hybrid Text and ID Embeddings
    by Vivek Kanojiya, Vishalaksh Aggarwal, Daeho Baek, Lyndon Kennedy and Xuetao Yin

    Incremental video search requires high-quality ranking after each keystroke, where intent is often underspecified (e.g., 1–3 character prefixes). We present a personalization system for Apple TV search that combines complementary semantic and collaborative signals at ranking time. Our approach learns two item embedding spaces: (i) a text-based multilingual encoder (TextEmb) fine-tuned on co-engagement triplets via contrastive learning, and (ii) an ID-based collaborative embedding model (IdEmb) trained on interaction-derived positives. At serving time, we construct user representations from recent watch history and inject text- and ID-based user–item cosine similarities into a pairwise XGBoost ranker. We evaluate the system with temporally held-out offline datasets and a three-week online controlled experiment. Offline, for sessions with user history, the personalized ranker improves NDCG@10 by 2.99% and MRR by 3.30% over the non-personalized baseline. Crucially, slice analyses show that personalization is most needed in incremental search, where intent is still forming: on ambiguous prefix queries (1–3 characters), NDCG@10 lift is +8.63%, versus only +1.46% on longer, more fully specified queries. Users with longer watch histories benefit more from personalization than newer users: NDCG lift rises from +2.13% for users with 1–5 history items to +4.37% for users with 51–100. This larger lift occurs even though baseline relevance is lower for long-history cohorts (NDCG@10 drops from 0.733 to 0.680), indicating that personalization adds the most value where default ranking underperforms. Online, treatment yields statistically significant gains of +1.14% tap-through rate and +1.23% conversion rate, with a 2.91% improvement in converted-item rank position. We further analyze coverage–precision trade-offs between semantic and collaborative embeddings through ablations isolating each signal, and evaluate embedding quality on a held-out corpus with LLM-judged similarity labels to reduce click/exposure bias.

  • Spot A2Learned Cross-Task Relationships in Multi-Task Models
    by Victor Zhang, Yiping Yuan, Florian Raudies, Bosun Adeoti, Brian Leung, Sanjay Surendranath Girija and Naijing Zhang

    We propose a framework that learns cross-task relationships in multi-task models by approximating the joint distribution of task labels through targeted pairwise relationships. This approach improves performance via transfer learning and enhances information extraction without the intractable complexity of modeling the full joint space. Although our framework applies to any multi-task system, we demonstrate its efficacy within YouTube’s production recommendation systems. Experiments across the Notifications, Homepage, and Watch Next surfaces show improvements in both accuracy and user satisfaction metrics. Finally, we propose a workflow template to facilitate broader future implementation.

  • Spot B1Hypothesis-Driven Shelf Generation for Personalised Recommendation
    by Aleksandr V. Petrov, Tarun Chillara, Matthew D. Moellman, Lucas de Haas, Yabai Song, Alina Susoykina, Melissa Crawford, Gabriel Negash, Erik Franco, Tasnim Rahman, Binal Jhaveri, Shubham Bansal, Hugues Bouchard, Roberto Mirizzi, Mounia Lalmas and Alois Gruson

    Modern recommendation interfaces organise content into shelves: themed rows such as “More of What You Like” or “New Releases for You.” In production systems, these shelves are typically defined through hand-crafted templates coupled with dedicated retrieval logic. While effective for broad recommendation intents, this approach does not scale to the long tail of individual taste. We present a content-hypothesis-driven shelf generation system for Spotify Home that replaces fixed templates with natural language hypotheses describing what a personalised shelf should contain. The system has three phases: generating shelf hypotheses from user profiles, retrieving catalogue items that fulfil them, and aligning the final shelf by selecting coherent items and revising its title and subtitle. This separation decouples shelf planning from catalogue fulfilment, enabling independent optimisation of both stages, constrained generative retrieval over catalogue entities, and distillation of frontier LLM behaviour into compact models. Our production pipeline combines hypothesis generation, generative retrieval, candidate selection and shelf alignment, offline LLM-as-a-judge evaluation, and precomputed serving. We describe the end-to-end architecture and evaluate it through offline ablations and a large-scale deployment on Spotify Home. Results show that hypothesis-driven shelves substantially expand personalised recommendation supply while achieving engagement comparable to the platform’s strongest curated shelves.

  • Spot B2Interest Sequence for User Modeling in Industrial Short-Form Video Recommendation
    by Yuanzhen Lin, Diego Uribe, Yuan Shao, Ruixiao Sun, Yongle Cao, Zhimeng Jiang, Yuening Li, Yang Gu, Chuan He, Liang Liu and Sourabh Bansod

    Ultra-long user history modeling has been a highly effective approach in modern industrial recommendation systems, with most works heavily utilizing search-based methods and summarization based methods to map massive user interaction logs into latent user representation with algorithm designs to address the scaling challenge. Despite their success in user interest prediction, previous works are presented with significant scaling challenges due to the increasing computation cost in regards to sequence length and inherent information noise of item-based sequence. Furthermore, the lack of semantic interpretability and explicit interest segmentation often results in the loss of long-tail interest exploration. To address these limitations, we propose a novel Interest Sequence Modeling module and detail the real-world deployment within a production-scale ranking model serving billions of users. We introduce the User Interest Tree Profile (UITP), an explicit representation strategy that complements existing implicit chronological sequences by aggregating lifelong engagement metrics across a hierarchical taxonomy to extract explicit historical interest tokens of users. Through sequential modeling, this streaming asynchronous profile is processed into real-time interest-based sequence for model consumption. Offline evaluations demonstrate statistically significant improvements across all prediction targets and extensive live A/B experiments reveal core engagements gains, as well as highly beneficial ecosystem impacts: a $+2.43\%$ increase in exploratory views, a $+4.74\%$ boost in global inter-category diversity, and robust gains in fresh content engagement while maintaining low-latency serving constraints without large training overhead.

  • Spot C1M2R: End-to-End Implicit Scenario Discovery for Multi-Scenario and Multi-Crowd Ranking
    by Yuanhang Zhou, Chaoqun Hou, Wenzheng Fang, Jiaqi Zheng, Cheng Guo, Tong Liu and Bo Zheng

    E-commerce platforms typically host dozens of differentiated shopping scenarios, making multi-scenario modeling a de facto paradigm in industrial recommendation. However, classic methods rely on business rules to partition scenarios, which forcibly isolates similar samples across scenarios while failing to capture heterogeneous distributions within a scenario. Recent works either explore inter-scenario aggregation without addressing intra-scenario heterogeneity, or apply unsupervised clustering that frequently suffers from blurred cluster boundaries and prototype collapse due to the lack of supervision. In this paper, we present M2R, an implicit multi-scenario multi-crowd ranking framework: (1) it adaptively partitions samples into implicit scenarios driven by data distribution rather than business rules, addressing both cross-scenario collaboration and intra-scenario heterogeneity, and serves as a plug-and-play router that integrates seamlessly into mainstream MoE-based multi-scenario backbones; (2) it jointly optimizes the clustering router with the recommendation signal and two clustering-aware auxiliary losses end-to-end, mitigating the negative transfer caused by ambiguous cluster boundaries; (3) it introduces parameter-efficient LoRA experts that scale the expert pool to hundreds while preserving predictive accuracy and substantially reducing parameter counts. Online A/B tests on Taobao show that M2R lifts overall orders across 21 long-tail marketing scenarios by +6.3%, orders in the core Taobaomiaosha channel by +2.1%, and gross merchandise volume in Baiyibutie by +8.8%, demonstrating its practical value for industrial multi-scenario recommendation.

  • Spot C2TSSR-Beta: Enhancing Billion-Scale E-Commerce Semantic Retrieval via Representation-Level Interaction
    by Guohao Tan, Jiahui Wan, Tao Wen, Dongshuai Li, Xingxian Liu, Daoning Jiang, Xiaoyao Qiu, Yuliang Yan, Dan Ou, Haihong Tang and Bo Zheng

    Semantic retrieval in e-commerce search aims to identify a compact candidate set from billion-scale product catalogs with both high recall and low latency. Dual-Encoders dominate this stage due to their efficient dot-product similarity, but this formulation limits model expressiveness and fails to capture fine-grained relationships between queries and items. While prior work has explored interaction-based similarity, its additional cost often prevents deployment at industrial scale. We present TSSR-Beta (Taobao Search Semantic Retrieval Model – Beta), which improves the expressiveness of our production Dual-Encoder TSSR through a plug-in similarity module, termed the Hybrid Interaction Head. This module introduces fine-grained matching in the representation space through two complementary pathways: InteractMLP, which captures explicit matching patterns with residual MLP blocks, and InteractTrans, which models implicit cross-dimensional interactions with a Transformer layer. Their outputs are combined by a Fusion Head to produce the final similarity score. TSSR-Beta introduces only a small parameter overhead to the Dual-Encoder without changing its architecture, enabling it to be (1) pluggable, readily adapting to diverse Dual-Encoder backbones; (2) efficiently trainable, supporting large-batch contrastive learning with massive negative sampling; and (3) industrially deployable, preserving offline item pre-encoding and supporting low-latency online retrieval with Neighborhood-Aware Approximate Nearest Neighbor (NANN). Offline experiments on the Taobao Search show a +3.90pp Hitrate@500 improvement over TSSR. On public Natural Questions and WebQA, our module further improves Recall@1 by +4.60pp and +0.90pp over public Dual-Encoder, respectively. Deployed in Taobao Search, TSSR-Beta delivers low-latency billion-scale retrieval and achieves +0.63% transaction count and +2.69% GMV gains in online A/B tests.

  • Spot E1CAPTS: Channel-Aware, Preference-Aligned Trigger Selection for Multi-Channel Item-to-Item Retrieval
    by Xiaoyou Zhou, Yuqi Liu, Zhao Liu, Xiao Lv, Bo Chen, Ruiming Tang, Guorui Zhou, Han Li and Kun Gai

    Large-scale industrial recommender systems adopt multi-channel retrieval for candidate generation, combining direct user-to-item (U2I) retrieval with two-hop user-to-item-to-item (U2I2I) pipelines. In U2I2I, the system selects a small set of historical interactions as triggers to seed item-to-item (I2I) retrieval across multiple channels. In production, triggers are often selected using rule-based policies or learned scorers and tuned channel by channel. However, these practices face two challenges: biased value attribution, which values triggers by on-trigger feedback rather than downstream retrieval utility, and uncoordinated routing, where channels independently select triggers under a shared quota, increasing cross-channel overlap. To address these challenges, we propose Channel-Aware, Preference-Aligned Trigger Selection (CAPTS), a framework that treats multi-channel trigger selection as a learnable routing problem. CAPTS introduces a Value Attribution Module (VAM) that credits each trigger with subsequent engagement from items retrieved through each I2I channel, and a Channel-Adaptive Trigger Routing (CATR) module that coordinates trigger-to-channel assignment. Offline experiments and large-scale online A/B tests on Kwai, Kuaishou’s international short-video platform, show that CAPTS consistently improves multi-channel recall offline and delivers +0.713% total app time spent and +0.586% average app time spent per device online.

  • Spot E2RecEvolve: A Knowledge-Driven Autonomous Agent System for Recommender Systems
    by Weidi Pan, He Ma, Shuhao Ye, Palaksh Rungta, David McPeek, Junyi Jiao, Arnab Bhadury, Mingyan Gao and Onkar Dalal

    The rise of agentic AI has catalyzed a shift toward self-iterating systems, opening new frontiers for the autonomous optimization of production recommender models. This paper presents the empirical validation of a knowledge-driven autonomous agent system, deployed directly on a production large-scale Two-Tower retrieval model. By delegating the entire research lifecycle, spanning idea generation, code implementation, offline training, and metric evaluation, to a continuous closed-loop autonomous framework, the agent system executed over 40 completed autonomous training runs from scratch. Executing these runs under rigorous production-scale evaluations, the system systematically navigated hidden architectural bottlenecks on the latest production model to achieve a breakthrough $\sim$20\% relative improvement in NDCG, a gain that translated directly to a +3.77\% increase in user satisfaction in live production traffic. Furthermore, the deployment exposed critical vulnerabilities in standard evaluation protocols, as the agent system autonomously discovered reward-hacking shortcuts. These findings prove that an autonomous pipeline can dramatically accelerate the pace of machine learning research and stress-test the rigorousness of underlying experimental infrastructure, while also exposing novel challenges such as reward hacking and redundant exploration of failed hypotheses.

  • Spot F1LLM-Based Re-Ranking for Real Estate Search
    by Nkateko Ntimane, Rafael Guedes, Tiago Cunha and Pedro Nogueira

    QuintoAndar Group operates the leading housing marketplace in Latin America for both rentals and sales. The platform replaces traditionally paper-heavy workflows with a fully digital experience, making housing transactions faster and more accessible to tenants, buyers, and landlords in the region. Finding the ideal home in such a vast catalog is inherently difficult. At the same time, the wide- spread adoption of conversational assistants is reshaping user ex- pectations: people increasingly want to express their needs through open, multi-turn dialog rather than rigid filter menus and faceted search. This shift is particularly pronounced in housing, where in- tent is multi-dimensional, context-dependent, and rarely reducible to a small set of structured constraints. To meet these expectations, we propose a Large Language Model (LLM) based re-ranker that augments a conversational recommendation system by reorder- ing retrieved candidates according to the nuanced, context-rich intent expressed across the user’s conversation. We additionally construct a large-scale offline evaluation dataset for conversational real-estate search, containing 960,000 query-item pairs constructed from both synthetic and production queries and annotated using an LLM-as-a-Judge framework with human validation. We validate our approach both offline, on this proprietary dataset, and online, through a production A/B test. Both evaluations show consistent improvements in ranking quality, including a statistically signif- icant increase in production of +5.3% in click-through rate and +4.8% in scheduled visits, demonstrating the value of integrating conversational context into housing recommendations.

  • Spot F2FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Billion-Scale Livestreaming Recommendation
    by Xinhang Yuan, Zexi Huang, Anjia Cao, Xudong Lu, Zikai Wang, Penghao Zhou, Chang Liu, Wentao Guo and Qinglei Wang

    Modern recommender systems rely heavily on ID-based collaborative filtering, where each item is represented by a unique ID embedding that accumulates collaborative signals from user interactions. Livestreaming recommendation, however, faces a unique challenge within this paradigm. A live room enters the candidate pool only while broadcasting, typically for tens of minutes, so its ID embedding never converges and ID-centric rankers fail to generalize. To address this, we present FLUID, the first framework to fully retire the item ID from a production livestreaming ranker. FLUID introduces a cross-domain multimodal encoder that is jointly trained on livestreams and short videos to produce discrete semantic codes (called LUCID) for content-based item characterization. Next, FLUID applies a staged warmup training scheme to adapt the ranker to LUCID. It first leverages the ranker backbone for a late fusion of cold, slice-level LUCID embedding and ID embedding and then replaces the ID embedding with warm, room-level LUCID embedding before the final online incremental training. Deployed on our online livestreaming ranker with a cross-platform combined user base of over one billion globally, FLUID delivers significant A/B gains of +0.55% Quality Watch Duration, +2.05% Cold-Start Room Views, and +0.05% Active Hours, demonstrating the superiority of its ID-free design at industrial scale.

  • Spot O1Uncertainty-Aware Reward Modeling: A Large-Scale Case Study in Video Recommendation
    by Nitu Sharaff, Brian Y. C. Leung, Shawn Andrews and James Harrison

    Historically, recommendation systems have focused on maximizing precision by treating user preference as a static, predictable target. However, this approach ignores both the inherent randomness of human behavior and the model’s own varying levels of confidence. To compensate, many industrial systems rely on “reserved slots” for exploring user interests- a heuristic that typically utilizes uniform selection. This paper presents a large-scale study on the YouTube video recommendation platform, where we integrated principled exploration directly into the ranking scores of the Homepage Ranking model. To model uncertainty, we evaluate two distinct architectures: a Variational Bayesian Last Layer (VBLL) designed to capture model’s parameter uncertainty, and Quantile Regression (QR) utilized to model the variance within the target reward distribution. By strategically targeting the right exploratory candidates, both paradigms successfully break the traditional explore-exploit trade-off, by driving improvements in both user engagement and content discovery metrics. Large-scale online A/B testing reveals that both paradigms are highly viable for production, successfully balancing exploration and exploitation while driving significant improvements in overall user engagement and content discovery. By analyzing the distinct behaviors of the VBLL and QR deployments, we compare both paradigms and hypothesize how the distinct uncertainties being modeled affect the final reward distributions. Finally, we discuss our ongoing efforts to merge these two paradigms and share preliminary notes on their integration.

  • Spot O2Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation
    by Ziyun Xu, Bosen Ding, Ji Qi, Qingyuan Song, Jizhou Huang, Liwei Wang, Yue Zhang, Qichao Que, Yue Weng, Zhenheng Yang, Junfeng Pan, Jeffrey Santelli and Linhong Zhu

    Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., clicks, dwell time) rather than explicit, natural language inputs. As a result, users experience a persistent discrepancy between their explicit interests and what passive behavioral algorithms deliver, limiting their ability to express nuanced preferences or steer their feed in real time. To address this growing gap between how recommendations are optimized and how users wish to articulate their interests, we present Shape Your Feed (SYF), an LLM-based agentic recommendation framework that enables real-time, multimodal co-curation of content. SYF employs a three-tier architecture: (i) a Perception Flow that captures fine-grained user intent from text prompts, voice commands, and UI interactions; (ii) a Serving Flow that performs real-time agentic re-ranking and pruning of candidate items, grounded in a persistent Semantic Profile encoding evolving user preferences; and (iii) a Self-Evolution Flow that aligns system behavior with human judgments via Direct Preference Optimization (DPO) and an LLM-as-a-Judge ensemble. Offline evaluations show that SYF’s alignment scoring module achieves 98.85\% accuracy, substantially improving over strong few-shot baselines. Large-scale online A/B experiments on production traffic further demonstrate that SYF improves feed relevance and user sentiment, indicating a practical and scalable path toward interactive, user-steerable recommendation in industrial settings.

  • Spot P1Soft Curriculum Learning for Optimizing Fresh and Generalized Recommendations
    by Arnab Bhadury, Siyan Zheng, Anlan Yu, Palaksh Rungta, Jiawei Li, Changping Meng, Dapeng Hong, Chuan He and Onkar Dalal

    Large-scale recommender systems, particularly short-form video platforms, are often bottlenecked by massive popularity feedback loops. In such environments, as models recommend popular items, they generate an overwhelming amount of skewed training data for “head” items. This creates a self-reinforcing cycle where retrieval and ranking models memorize “head” item patterns at the expense of generalizing across the vast “tail” of the catalogue. While Curriculum Learning (CL) offers a powerful mechanism to break this feedback loop by systematically exposing models to progressively more difficult and less frequent examples, its adoption in industrial recommendation has been hampered by hardware utilization inefficiencies or the needs for complicated pre-processing techniques because dynamic data rejection algorithms tend to starve hardware accelearators (TPUs/GPUs) by becoming largely CPU-bound. In this work, we introduce a scalable Soft Curriculum Learning framework designed specifically for continuous training setups within industry-scale retrieval and ranking models. By utilizing loss annealing and in-graph weight adjustments rather than rigid data filtering, we break the popularity feedback and enable dynamic curriculum pacing without sacrificing system throughput. We demonstrate empirical evidence through applications across sequence-based retrieval models (such as SASRec), two-tower retrieval models, and large-scale continuous ranking models. Online A/B tests on our short-video platform demonstrate substantial lifts in both overall user satisfaction and fresh content consumption, all without degrading model throughput.

  • Spot P2Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads
    by Zhimeng Pan, Jiayin Jin, Yang Tang, Jiarui Feng, Kungang Li, Chongyuan Xiang, Jiacheng Li, Runze Su, Chuizheng Meng, Siping Ji, Han Sun, Litian Tao, Ling Leng and Jamieson Kerns

    Offsite conversion data provide valuable signals for ads ranking because they reveal users’ commercial intent beyond their onsite activities. However, incorporating offsite data into large-scale ads models is challenging: conversion records are supplied by different advertisers and partner vendors, often contain sparse or inconsistent metadata, and exhibit behavioral distributions that differ from onsite ad engagements. In this work, we present a production-scale representation-learning framework for leveraging opt-in offsite conversion data in Pinterest Ads ranking models. We construct a heterogeneous onsite-offsite graph that connects various kinds of entities, such as users, ads, items, links, and advertisers, through onsite engagements and opt-in offsite conversions. We then apply TransRA, a revised variant of TransR, to learn scalable ID-based entity representations from this graph. TransRA designates one entity space as an anchor and maps non-anchor entity types into the anchor space, preserving relation-specific transformations while making the learned embeddings easier for different downstream ranking models to consume. We further study how to incorporate pretrained Knowledge Graph Embedding (KGE) representations into ads ranking models. Directly using pretrained KGE embeddings provides limited gains, suggesting a mismatch between graph pretraining objectives and ranking-model training. To address this challenge, we employ the Large ID Embedding Table technique and develop an attention-based KGE fine-tuning method that allows ranking models to adapt graph-derived entity embeddings jointly with the ranking objective while satisfying online serving constraints. We evaluate this framework in both the Ads Engagement Model for Click-Through Rate (CTR) prediction and the Ads Conversion Model for Checkout Conversion Rate (CVR) prediction. Offline experiments show consistent improvements in both models, and online A/B testing demonstrates gains in user engagement, conversion efficiency, advertiser value, CPA and CPC metrics. The framework has been deployed to full-scale traffic in both CTR and CVR models, showing that onsite-offsite graph representation learning can be effectively integrated into production ads ranking systems under industrial-scale constraints.

  • Spot V1Token Factory: Efficiently Integrating Diverse Signals into Large Recommendation Models
    by Xilun Chen, Shao-Chuan Wang, Baykal Cakici, Lukasz Heldt, Lichan Hong, Raghu Keshavan, Aniruddh Nath, Li Wei and Xinyang Yi

    Large Recommendation Models (LRMs) have demonstrated promising capabilities in industry-scale recommendation tasks. However, holistically integrating traditional signals into these transformer based architectures effectively and efficiently remains a major challenge. Conventional approaches that “textualize” these signals directly or create discrete item representations often lead to excessively long prompts, substantial memory footprints, and high computational overhead. To overcome these limitations, we propose “Token Factory”, a framework designed to transform traditional signals into “soft tokens” that can be directly processed by LRMs. This approach enables efficient integration and compression of heterogeneous input features, preventing prompt length explosion while enhancing model performance. We detail the architecture of Token Factory and present experimental results validating its effectiveness in a production-scale recommendation environment.

  • Spot V2Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training
    by Oseong Choi, Hoeinn Kim, Jihoon Lee, Byungsoo Kang and Taeyeong Jang

    Foundation model(FM) for recommendation has shown strong ability to model long-horizon sequential user behavior. In practice, a single pretrained foundation model is often adapted to diverse downstream serving surfaces through Supervised Fine-Tuning(SFT). However, optimizing task-specific objectives such as clicks or likes does not necessarily align the serving policy with the business metrics that determine recommendation quality. We propose a three-phase progressive post-training framework that explicitly separates downstream adaptation from business-metric alignment. The adaptation stage is decomposed into Linear Probing(LP) and Full Fine-Tuning(FFT): LP first stabilizes randomly initialized downstream heads within a frozen pretrained representation space, and FFT then jointly specializes the full model for the target task. On top of this stabilized policy, Reinforcement Fine-Tuning(RFT) aligns the model with practical business objectives using a learned reward model. Rather than directly optimizing the serving policy on sparse business targets, we train the policy on dense implicit feedback and use business-metric supervision only for reward modeling. Offline experiments show that the progressive LP-FFT-RFT framework outperforms single-phase alternatives, and that reward-based alignment yields a stronger serving policy than directly using the reward model itself for ranking. Large-scale online A/B tests further show that the proposed framework improves production recommendation quality over a conventional non-foundation baseline. A reference implementation is available at https://github.com/webtoon/rec-fm-progressive-alignment.

  • Spot W2TransAct V2: Production System for Lifelong User Sequence Modeling at Scale
    by Xue Xia, Saurabh Joshi, Kousik Rajesh, Kangnan Li, Yangyi Lu, Nikil Pancha, Dhruvil Badani, Jiajing Xu and Pong Eksombatchai

    Modeling lifelong user action sequences in CTR prediction faces three production challenges: \textbf{computational complexity} with quadratic transformer costs as sequences extend from hundreds to tens of thousands of actions, \textbf{infrastructure overhead} from storage and network costs that scale with both sequence length and number of candidates per request, and \textbf{weak training supervision} as sequence encoders sit far from CTR prediction heads. We present TransAct V2, a production system deployed at Pinterest serving over 630 million users, that addresses these challenges through: (1) a dual-path training-serving architecture with request-level deduplication and int8 quantization achieving 1\% logging cost, (2) fused Triton kernels with pinned memory management delivering 250$\times$ p99 latency improvement, and (3) a Next Action Loss providing direct supervision for sequence modeling within CTR frameworks. Deployed in production, TransAct V2 demonstrates significant improvements in both engagement and recommendation diversity. To support industry adoption, we open-source our serving optimizations with comprehensive ablation studies. Our work provides a reproducible blueprint for practitioners building large-scale sequential recommenders.

Reproducibility

  • Spot G1WorkRB: A Community-Driven Evaluation Framework for AI in the Work Domain
    by Matthias De Lange, Warre Veys, Federico Retyk, Daniel Deniz, Warren Jouanneau, Mike Zhang, Aleksander Bielinski, Emma Jouffroy, Nicole Clobes, Nina Baranowska, David Graus, Marc Palyart, Rabih Zbib, Dimitra Gkatzia, Thomas Demeester, Tijl De Bie, Toine Bogers, Jens-Joris Decorte and Jeroen Van Hautte

    Today’s evolving labor markets rely increasingly on recommender systems for hiring, talent management, and workforce analytics, with natural language processing (NLP) capabilities at the core. Yet, research in this area remains highly fragmented. Studies employ divergent ontologies (ESCO, O*NET, national taxonomies), heterogeneous task formulations, and diverse model families, making cross-study comparison and reproducibility exceedingly difficult. General-purpose benchmarks lack coverage of work-specific tasks, and the inherent sensitivity of employment data further limits open evaluation. We present \textbf{WorkRB} (Work Research Benchmark), the first open-source, community-driven benchmark tailored to work-domain AI. WorkRB organizes 13~diverse tasks from 7~task groups as unified recommendation and NLP tasks, including job, skill recommendation, candidate recommendation, similar item recommendation, and skill extraction and normalization. WorkRB enables both monolingual and cross-lingual evaluation settings through dynamic loading of multilingual ontologies. Developed within a multi-stakeholder ecosystem of academia, industry, and public institutions, WorkRB has a modular design for seamless contributions and enables integration of proprietary tasks without disclosing sensitive data. WorkRB is available under the Apache 2.0 license at https://github.com/techwolf-ai/WorkRB.

  • Spot G2Attacking and Defending Multi-Agent Collaborative Filtering Systems Through Connectivity
    by Anjun Hu, Hanting Xie, Saranya Govindan, Jas Kandola and Kurt Cutajar

    Multi-agent collaborative filtering systems coordinate autonomous LLM-powered user and item agents through natural-language interaction to refine preferences and generate recommendations. These systems inherit vulnerabilities from both their data-driven nature and multi-agent interactions, which manifest in distinct ways. Understanding how connectivity modulates vulnerability in these systems could facilitate the development of more robust recommendation pipelines. In this work, we adapt attacks and defenses from the general multi-agent systems (MAS) literature to the agent-based CF setting, evaluating them under systematically varied connectivity in the AgentCF framework, where CF connectivity is characterized along two axes: (i) candidate count (the number of item candidates per turn per user, measuring user-side interaction density) and (ii) catalog concentration (the degree of item catalog overlap across users). Our contributions include: (1) Adaptation: we reproduce MAS-inspired attacks and defenses in the agentic CF domain, confirming partial transferability of original observations. (2) Characterization: we characterize how the two aspects of connectivity shape attack and defense outcomes, revealing role asymmetries between user and item agents, non-monotonic temporal dynamics in attack efficacy, and divergent patterns across dissemination and extraction attack goals. (3) Prediction: motivated by empirical findings, we assess the applicability of epidemic-inspired static metrics in ranking CF configurations by expected attack outcome without running full adversarial simulations, enabling cost-efficient robustness assessment. Implementation is available at https://anonymous.4open.science/r/connacf

  • Spot H1JTH: A Dataset for Evaluating Cold-Start and Temporal Dynamics in Job Recommendation
    by Julien Romero, Yann Millet and Éric Behar

    Online job‐matching platforms face a fundamentally different landscape from e-commerce or media recommendation. A job posting is typically live for only a few weeks, while an active candidate disappears as soon as they accept an offer. Besides, both sides may enter and leave the market multiple times over a career. These fleeting lifetimes (median 25 days for jobs, 10 days for candidates in our data) yield three compounding hurdles: (i) pervasive cold start for every newly posted vacancy and every first-time applicant, (ii) extreme interaction sparsity, and (iii) complex temporal overlap, because the window during which two entities can actually meet is narrow and highly variable (the median application process is of 8 days in our data).

  • Spot H2τ-Rec: A Verifiable Benchmark for Agentic Recommender Systems
    by Bharath Narasimhan and Karthik Narasimhan

    As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. Current benchmarks often rely on “LLM-as-a-judge” evaluations, which introduce subjectivity, high costs and inconsistency. We present ????-Rec, a benchmark for agentic recommender systems that replaces subjective evaluation with verifiable rewards and a reveal-tagged elicitation (RTE) mechanism that controls how task constraints surface during dialogue. By testing agents against structured catalog predicates and employing a pass^k reliability metric, τ-Rec provides a systematic test for consistent reasoning. Our evaluation of nine configurations across five model families — GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Flash, DeepSeek V4 Flash, Qwen3 32B and GPT-5 mini — reveals a steep reliability cliff, where even the best model achieves only ∼57% at pass^1 and ∼35% at pass^4, highlighting a critical gap in current conversational agent deployment. All code and data are publicly available at https://github.com/nbharaths/tau-rec.

R&P Notes

  • Spot I1Repetition Is Not Enough: Evaluating Repeat Recommendations Through Utility and Timing
    by Abhishek Srivastava

    Repeat consumption is a common pattern in e-commerce, and recommender systems in such settings are typically evaluated using accuracy-oriented measures of predictive performance. While these measures are effective for assessing correctness, they do not fully capture whether recommendations are useful or delivered at the right time. In repeat-heavy environments, predicting previously consumed items does not always correspond to meaningful or well-timed recommendations. In this work, we revisit evaluation for repeat recommendation and propose a utility and timing aware framework comprising three metrics: Repeat Utility (RU), Temporal Repeat Utility (TRU), and Temporal User Satisfaction (TUS). RU measures repeat recommendation usefulness via subsequent re-engagement. TRU rewards timely repeats. TUS combines timing and re-engagement strength. Initial experiments indicate the potential of these metrics over traditional metrics, to capture associations with user re-engagement, alignment with revisit patterns, and timely responses to recommendations.

  • Spot I2Deadweight Loss in Recommendation Systems
    by Minje Kim, Hae-Yeong Cho, Wooseung Kang, Hye-Jin Jeong, Gun-Woo Kim and Sang-Min Choi

    Content recommender systems can be viewed as platform-mediated allocation mechanisms among users, creators, and the platform. We model platform-side intervention as a wedge that can leave mutually beneficial user–creator interactions unrealized. To operationalize this effect, we construct an empirical Pareto frontier over user-side and creator-side utility and compare the resulting allocations with those obtained when one top-$K$ position is reserved for platform-beneficial exposure. Experiments show that slot reservation can reduce user-side ranking utility and creator-side exposure, while displacing relevant matches retained in the reference allocation. We operationally define deadweight loss (DWL) as the loss incurred when relevant user–creator matches are displaced without preserving either user-side utility or creator-side exposure. Our code is available at \url{https://anonymous.4open.science/r/dwl-in-rec}.

  • Spot J1Your 5-Core Train Split May Not Be 5-Core: Core Filtering and Temporal Splitting Do Not Commute
    by Danil Gusak and Evgeny Frolov

    Temporal recommender benchmarks are often specified as “5-core plus temporal split,” but this phrase does not define a unique evaluation protocol.We show that iterative -core filtering and temporal splitting do not commute: when the full log is filtered before the split, users and items can survive because of validation or test interactions unavailable during model fitting. The resulting fitted train split may therefore violate the same p-core property claimed for the benchmark. Across six timestamped datasets, fitted-train violations at p = 5 reach 38.2% of users and 22.3% of items. These protocol choices also reshape the evaluated population, targets, and candidate universe, and can change conclusions: across eight recommenders, the top-ranked model changes in 8 of 14 matched p-core settings with p > 0, and in 5 of 10 settings after a test-size guardrail. Thus p-core temporal benchmarks require an explicit filter/split order, not just a pruning threshold. We recommend reporting fitted-train core violation and publishing split manifests as lightweight checks for reproducible temporal evaluation.

  • Spot J2Should I Be Polite to My LLM Relevance Judge? Tone as a Severity Operating-Point Shift
    by Tian Zhang and Meng Li

    Large language models are widely used as relevance judges in information retrieval, yet their labels shift with surface-level features of the prompt. We study one such feature—tone, or how politely the prompt addresses the model—which prior work has examined only on generation tasks, with contradictory conclusions and no mechanism. Across eight judge models, five classifier-calibrated politeness levels, and paraphrase controls, we find that the effect is strongly model-dependent—pronounced in one judge, small in most—but that, wherever tone does move agreement, the change is explained not by better judgment but by a shift in the judge’s severity operating point: the leniency with which it assigns absolute scores. Agreement rises or falls only as this shift moves the judge toward or away from the human annotators’ strictness—tone turns the judge’s strictness knob, not its intelligence knob. We formalize this as a falsifiable, pre-specified prediction and confirm it across models (Spearman = −0.62, = 0.0005). The account unifies prior contradictory findings and identifies tone as a predictable validity threat for calibration-based metrics such as Cohen’s , but not for ranking-based metrics such as NDCG.

  • Spot K1More Test Users, More Overconfidence: Seed Variance in Sequential Recommender Tests
    by Danil Gusak, Anna Volodkevich and Evgeny Frolov

    More test users should make offline evaluation more precise. For stochastic recommenders, however, paired tests over users can become overconfident. They compare two models, not whether one training procedure performs better on average than the other. Drawing conclusions about those procedures requires repeated runs, not more users. A variance model separating run-to-run and user-to-user variation gives the inflation factor $1+U\sigma_a^2/\sigma_{\iid}^2$: more users reduce uncertainty about the two models but not across seeds. With 15 SASRec runs on fixed sequential benchmarks, this inflation predicts how often user-level tests falsely detect a difference between runs of the same procedure: rejection rates reach 80% on the 138K-user MovieLens-20M, while tests across runs remain near 5%. For near-ties in the commonly reported 1-2% range, the winner changes with the seed. We propose a lightweight seed-aware protocol using paired differences across runs and a variance-inflation diagnostic.

  • Spot K2Personalized Fashion Discovery via Session Trajectories and Soft Negatives
    by Gwendolyn Rippberger and Julia Neidhardt

    Fashion e-commerce routes every shopper through the same fixed category tree and reads a dislike as a permanent verdict. Both assumptions fail during exploratory browsing, where shoppers have no clear target and a “No” often means not this, right now. This work-in-progress note proposes navigation built per user, per session: the shopper is modeled as a single point in the item-embedding space, pulled toward liked items and pushed away from disliked ones, with every influence decaying over the session. The decay turns negative feedback into a soft negative: a rejected direction fades and can return, in contrast to the clean negative of Rocchio-style relevance feedback. A prototype on the public H&M catalog (≈105K articles) implements the full loop, from a cluster-based seed screen to retrieval. We outline a two-stage evaluation (a simulation with drifting preferences, followed by a within-subjects user study) addressing three research questions on interaction efficiency, reading negative feedback, and perceived user experience.

  • Spot L1CoSID: Concept-Conditioned Semantic-ID Decoding for Efficient Generative Recommendation
    by Danil Gusak, Anna Volodkevich and Evgeny Frolov

    Generative recommenders retrieve items by autoregressively decoding semantic IDs (SIDs). The standard autoregressive SID interface (SID-AR) represents each history item with K code tokens, expanding a T-item history to TK tokens, while trie-constrained beam search repeatedly re-enters the full backbone during generation. Composing each item’s K code embeddings into a single input token restores item-level history length, but every decoding step still runs through the full backbone. We introduce CoSID, a concept-conditioned SID decoder that encodes the history once into a compact next-item concept and delegates the entire beam search to a lightweight KV-cached decoder. Across four datasets under a global temporal split, CoSID matches or surpasses the baselines in accuracy, maintains comparable or broader catalog coverage, and delivers up to $6.2\times$ higher throughput at beam width 100 and $7.3\times$ at beam width 1000. Because autoregressive SID decoding no longer re-enters the backbone, throughput remains nearly independent of backbone depth.

  • Spot L2Cross-Spectrum Candidate Set Analysis for Popularity-Aware Recommendation
    by Hye-Jin Jeong, Ye-Jin Lee, Gun-Woo Kim and Sang-Min Choi

    Popularity-biased recommender systems can produce unequal recommendation utility across users with different popularity preferences. In particular, users who prefer niche items can experience lower recommendation accuracy even when the system’s overall performance improves. To examine this disparity, we propose a multi-task framework with a shared backbone and three heads specialized for Niche, Diverse, and Blockbuster-focused user groups. By analyzing intersections among their candidate sets, we identify group-specific and cross-group consensus items. Experiments show that the heads generate different candidate structures and popularity patterns across models and datasets, suggesting that their set intersections may support future popularity-aware reranking.

  • Spot M1Music Discovery Quality and the Value of Familiarity
    by Gustavo Escobedo, Geoffray Bonnin, Markus Schedl and Bruno Sguerra

    Arguably, driving music discovery is one of the most important functions of recommender systems; however, defining what characterizes discovery and determining what constitutes a successful one are both nontrivial problems. Here, we argue that repetition offers a valuable lens: when users discover a track they like, they typically listen to it repeatedly several times. In this work, we use this pattern as an indicator to characterise the quality of successful discoveries. We also argue that familiarity is a fundamental part of the discovery process. We therefore use different familiarity measures to generate user profiles from long term consumption logs and use them to predict discovery success. Our experiments, performed on listening data from an online music streaming platform, indicate that familiarity measures are a good proxy for discovery prediction.

  • Spot M2What price fairness? Evaluating Energy – Fairness – Accuracy Trade-off in Recommender Systems
    by Abhirup Mitra, Oleg Lesota and Antonela Tommasel

    Fairness-aware recommender systems aim to mitigate systematic imbalances in recommendation outcomes, including how visibility, relevance, and opportunities are distributed among users, items, and providers. However, these systems are usually evaluated in terms of accuracy and fairness alone, while their computational and environmental costs remain largely invisible. This omission matters because fairness interventions may affect the cost of recommendation in different ways. Training-time methods modify model optimization, post-processing methods add computation at inference time, and both may depend on the model, dataset, hardware, and deployment setting. We examine whether provider-side fairness in recommendation comes with a measurable green cost. We compare in-processing, graph-level reweighting and post-processing interventions across multiple models, two datasets, and two hardware settings. We measure recommendation quality, provider-side exposure, and energy consumption separately across training and inference stages. Our results show that the green cost of fairness is not uniform, post-processing shifts cost to repeated serving, while in-processing and graph-level methods avoid re-ranking overhead but vary substantially across models, datasets, and hardware. Findings call for evaluating fairness-aware recommendation as a three-way trade-off between accuracy, fairness, and computational cost.

  • Spot N1Training seeds and model-selection stability in recommender-system evaluation
    by Juan Manuel Rodriguez, Oleg Lesota and Antonela Tommasel

    Recommender-system experiments often rely on a single random training seed, assuming that run-to-run stochasticity has limited impact on evaluation conclusions. This assumption is risky, as a training seed may influence several algorithm-dependent mechanisms, including parameter initialization, mini-batch ordering, dropout, masking, latent sampling, and training-time negative sampling. We examine this assumption by fixing the data partition and varying the training seed across hyperparameter configurations. We analyze seed effects at three levels: user-level metric sensitivity, validation-based model selection and recommendation-list agreement. Results show that seed variation is often detectable. Its impact depends on whether configurations are clearly separated, whether validation results transfer to test, and whether similar scores lead to similar top-lists. Findings suggest that reporting single-seed results can overstate the stability of recommender system evaluation, and that training seeds should be treated as part of the evaluation protocol rather than as incidental implementation noise.

  • Spot N2Revisiting the Assumptions of Shared-Account Sequential Recommendation
    by Ye-Jin Lee, Hye-Jin Jeong, Wooseung Kang, Gun-Woo Kim and Sang-Min Choi

    Shared-account sequential recommendation (SSR) has been proposed to address the limitation of sequential recommendation (SR) models under shared-account environments. In this paper, we conduct a dataset-centric study to examine how shared-account environments affect SR models and which characteristics of shared-account data drive recommendation performance. By comparing representative SR models on matched single-user and shared-account datasets, we show that performance degradation depends on model architecture and is mitigated by bidirectional contextual modeling. We further establish an analysis framework covering five key shared-account characteristics—group size, sequence length, user similarity, sparsity, and item popularity—and analyze their influence on recommendation performance. Our findings provide support for the motivation of SSR.

Short Papers

  • Spot Q1Towards Efficient Hyperbolic Representation Learning for Recommender Systems
    by Tendai Mukande and Noel E. O Connor

    Real-world relational data in recommender systems (RSs), particularly in domains such as e-commerce, often exhibit hierarchical structures, such as linking User → Product Category → Sub-Category → Item. Message-passing models such as GNNs, which propagate information across nodes, have been applied in most existing RS models. However, repeated averaging in message-passing for complex hierarchical data often leads to oversmoothing and oversquashing, where node embeddings collapse to similar vectors, reducing their discriminative power and degrading model performance. To address these issues, we propose HyperRec, a hyperbolic representation model for heterogeneous recommendation, which preserves the hierarchical structure in the embedding space and mitigates feature collapse. Experimental results on three real-world datasets show that HyperRec achieves superior performance with competitive efficiency. The implementation code will be open-sourced.

  • Spot Q2PriCoRec: A Privacy-Aware Cloud–Device Collaborative Framework for Ad Recommendation under Feature Constraints
    by Dairui Liu, Zhongyi Lu, Jitao Lu, Aghiles Salah, Mete Sertkan, Roger Zhe Li, Changhong Jin, Barry Smyth, Xingsheng Guo and Ruihai Dong

    Privacy regulations increasingly restrict cloud processing of sensitive user data (e.g., age, gender), hindering traditional cloud-only recommendation models. To mitigate this challenge, we propose a Privacy-aware Collaborative cloud-device ads Recommendation framework (PriCoRec) that allows ad recommendations to be personalized while keeping sensitive data on-device. While separating recommendation into cloud-based and on-device stages enables privacy-aware deployment, naive splitting suffers from degraded shortlist quality and inefficient on-device inference due to limited private features. We therefore design a collaborative framework that explicitly compensates for this feature gap. Our approach comprises a cloud-based pre-ranking stage using cloud-accessible features, and an on-device ranking stage that locally incorporates highly personalized features. We introduce a diversity regularizer to pre-ranking to improve candidate quality. Moreover, to ensure controlled on-device power consumption and computational cost, we incorporate a cloud-guided training mechanism that enhances the performance of the device model while keeping the model lightweight. Empirical results demonstrate that the proposed framework maintains strong recommendation performance while keeping sensitive features on-device.

  • Spot R1Understanding ID-Text Complementarity in Sequential Recommendation
    by Liam Collins, Bhuvesh Kumar, Clark Mingxuan Ju, Tong Zhao, Donald Loveland, Leonardo Neves and Neil Shah

    Dense retrieval-based Sequential Recommendation (SR) systems increasingly leverage text features to represent items, and generally do so in one of two ways: (i) by completely replacing ID embeddings with text embeddings, or (ii) by carefully combining ID and text embeddings through complex fusion mechanisms. While these lines of work have produced strong recommendation performances, they offer conflicting perspectives on ID-text complementarity, or the extent to which ID and text embeddings specialize in modeling different signals: the former suggests a lack of complementarity, and the latter argues it exists but must be harnessed carefully. Moreover, neither workstream conducts an in-depth study of ID-text complementarity. We aim to clarify this picture by developing a rigorous understanding of the complementarity of ID- and text-based SR models in this work. Our study reveals that these models do learn complementary signals, meaning that either should provide performance gain when used properly alongside the other. Motivated by this, we introduce and evaluate a new, simple SR baseline that preserves ID-text complementarity through independent model training, then harnesses it via ensembling. Despite this method’s simplicity, we show it outperforms several competitive SR baselines, implying a third perspective on ID-text complementarity: both features are necessary to achieve state-of-the-art SR performance, but complex fusion strategies are not.

  • Spot R2Personalized Recommendation Tool Learning via Autonomous Language Agents
    by Mingdai Yang, Zhiwei Liu, Weizhi Zhang, Yibo Wang, Hao Peng and Philip Yu

    Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and extensive world knowledge, previous LLM-based agents suffer from hallucination and context-length limitations, and thus are not suitable for full-ranking recommendation tasks. To overcome these drawbacks, we propose an agent-based recommendation framework, Personalized Recommendation Tool learning via autonomous language Agents (PRTA), in which an LLM acts as a central planner interacting with multiple recommendation models as tools. The LLM-based agent is responsible for high-level reasoning and personalized tool selection, while traditional recommendation models perform full-ranking scoring, leveraging their scalability in modeling behavioral patterns. To support personalized tool selection, we design reflection mechanisms that enable the agent to evaluate and compare tools for each user based on user profiles and candidate ranked lists. Extensive experiments across three public datasets demonstrate the superiority of PRTA over traditional recommendation and LLM-based baselines in improving full-ranking recommendation performance. Our code implementation is available online.

  • Spot S1LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation
    by Seungheon Doh, Bruno Sguerra, Sergio Oramas, Elena V. Epure and Juhan Nam

    Conversational Recommendation Systems (CRS) aim to achieve two primary objectives: recommending relevant items and generating natural language recommendation responses, i.e., system utterances that present and justify suggested items to the user. While recommendation accuracy is effectively measured by established ranking metrics, the evaluation of response generation poses a more fundamental challenge. Although human evaluation remains the gold standard, its cost and scalability constraints have motivated the adoption of LLM-as-a-judge as a promising proxy, whose alignment with human judgment in the context of CRS remains an open question. In this paper, we present the first user study to empirically assess the reliability of LLM-as-a-judge for evaluating CRS responses. We sample 20 multi-turn music recommendation sessions and generate candidate system responses using four instruction-tuned LLMs (1B to 4B parameters), inducing variance in response quality across model scales. We collect 400 ratings from 20 domain-expert annotators, who evaluate each response across two orthogonal dimensions: Personalization Quality and Explanation Quality. Through bootstrapped correlation analysis with human preference scores, we demonstrate that LLM-based judges outperform all reference-based baselines. Furthermore, we analyze how judge performance varies according to model scale and conditioning information, providing practical guidance for deploying LLM-as-a-judge as a cost-effective proxy for human evaluation.

  • Spot S2Hierarchical GTV Estimation: Bridging the Gap between Container-Level Ranking and Item-Level Conversion
    by Zerong Lan, Fan Zhang, Chuang Chen, Tianchun Huang, Teng Zhang and Xingxing Wang

    Emerging e-commerce platforms, such as live-streaming and Online-to-Offline (O2O) services, inherently operate under a “Container-based” display paradigm. In this setting, a critical granularity mismatch exists: the platform performs ranking and estimation at the container level (e.g., livestream rooms or stores), whereas the actual transactions occur at the item level within these containers. This structural discrepancy complicates Gross Transaction Value (GTV) estimation, introducing unique challenges including high label variance, severe label sparsity, and dynamic item heterogeneity. To address these challenges, we propose a novel framework named Hierarchical GTV Estimation (HGE). HGE employs a Hybrid Set-Aware Encoder (HSAE) to model the dynamic composition of items and feature interactions. To mitigate label variance, we introduce a Proxy-Label Learning (PLL) strategy that decomposes post-click GTV into purchase probability, quantity, and unit price. Furthermore, a Hybrid Boosting & Bagging Strategy (HBBS) is designed to handle label sparsity by effectively combining item-level aggregation with container-level inference. Extensive experiments on large-scale industrial datasets demonstrate that HGE significantly outperforms state-of-the-art baselines, achieving a +0.63% lift in XAUC offline and a +1.42% increase in Revenue Per Search (RPS) in online A/B tests. HGE framework has been fully deployed into our main traffic.

  • Spot T1Interaction Modality and Trust: Investigating System-Driven Conversational Recommendation and Faceted Search
    by Laura Modre, Ahmadou Wagne, Thomas Elmar Kolb and Julia Neidhardt

    This work adopts an interdisciplinary approach to investigate how interactions with conversational recommender systems (CRSs) can foster user trust, compared to faceted search interfaces commonly used in e-commerce. We present an empirical comparison between a system-driven CRS and a functionally equivalent faceted search system, isolating the effect of interaction modality through a between-participants user study (N = 144). While perceived explainability and anthropomorphism were positively associated with trust across both conditions, no significant differences emerged in trust or its antecedents between conditions. In contrast, CRS users reported significantly lower perceived control. This difference reflects a design trade-off: the system-driven structure constrained user actions in order to ensure safe dialogue. Our results suggest that designers must balance the benefits of structured guidance against the agency cost of delegating control to the system.

  • Spot T2BookForYou: Leveraging BART-Generated Narrative Tropes for Content-Based Book Recommendations
    by Ritu Rajesh Kanchi, Ruhul Amin Hazarika and T. Gopalakrishnan Thirumoorthy

    This paper presents BookForYou, a hybrid content-based book recommendation system that integrates deep generative modeling with semantic retrieval. The system utilizes a fine-tuned BART model (139M parameters) to extract structured narrative trope tags from book descriptions, while Sentence-BERT (SBERT) performs initial retrieval from a catalog of 42,350 titles. Candidates are re-ranked using a late-fusion approach that combines semantic similarity with an Asymmetric Trope Overlap (ATO) metric. Evaluation on a held-out test set (n=89) demonstrates that the system achieves an NDCG@10 of 0.705, significantly outperforming SBERT-only (0.392, p < 0.001), BM25 (0.364), and popularity-based baselines (0.136). Ablation analysis identifies ATO as the dominant performance driver: trope-based re-ranking alone (0.691) improves upon SBERT by 76.3%, while the optimal hybrid configuration (alpha=0.55) provides an additional 2.1% gain (p=0.018). Furthermore, the system achieves a 92.7% trope match rate and a serendipity score of 0.679. By utilizing matched tropes as natural-language explanations, BookForYou offers a transparent and highly effective approach to narrative-driven recommendation.

  • Spot U2Candidate Retrieval for Provider Fairness under Severe Recommendation Imbalance
    by Patrik Dokoupil, Ludovico Boratto and Ladislav Peska

    Provider fairness in recommender systems is commonly addressed through reranking, which assumes that the candidate set already contains sufficient items from under-exposed provider groups. We study a setting in which this assumption breaks down: catalog representation does not translate into recommendation exposure, and minority items are largely missing from the candidates passed to the reranker. In this regime, methods such as FA*IR, calibration-based reranking, or minority boosting can only partially improve fairness because the items needed to rebalance the ranking are absent upstream. We address this bottleneck with FairSAE, an activation-steering method that shifts user representations toward minority items through a single contrastive direction in a latent space induced by a sparse autoencoder (SAE), before candidate retrieval. This pre-retrieval intervention surfaces minority items from the full catalog and can be combined with existing fairness-aware rerankers. Across the MovieLens-25M and Steam datasets, FairSAE-steered retrieval substantially improves the fairness-accuracy trade-off compared to standalone reranking. These results show that provider fairness under severe recommendation imbalance depends not only on reranking, but also on which items are retrieved in the first place. See https://bit.ly/fairsae for codes and data.

TORS

Back to program