Session 7:
7A: Online Experimentation & Exploration
Date: Thursday October 01, 08:30 – 10:00 CDT
Session Chair: Morten Arngren
- RESProbabilistic Residual Learning for Online Recommendations
by Wenyuan Wang, Yusong Zhao, Zihao Xu, Hengyi Wang, Qi Xu, Zhigang Hua, Yan Xie, Yi Wang, Zihao Zhao, Bo Long, Chengzhi Mao, Shuang Yang, Hengguan Huang and Hao WangModern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and items. As a result, these systems often suffer from the black-box nature and computational complexity of the underlying models, making it difficult to systematically enhance their recommendation capabilities. To address this problem, we propose Probabilistic Residual Learning (PRL), a causal Bayesian recommendation model that models the residual between ground-truth and base predictions, enabling targeted refinement of existing systems. Specifically, PRL (1) divides users into clusters in an unsupervised manner and identifying causal confounders that influence latent variables, (2) learns sub-models for each confounder given the observable variables, and (3) generates recommendations by aggregating the rating residuals under each confounder using do-calculus. Experiments demonstrate that our plug-and-play PRL is compatible with various base DL recommender systems, improving their performance while automatically discovering meaningful user clusters. Auxiliary materials (including the Appendix) are at https://anonymous.4open.science/r/PRL_Appendix-CDD7/PRL_RecSys_Appendix.pdf.
- RESAdaptive Retraining of Recommender Systems via Reinforcement Learning
by Diego Russo, Valerio La Gatta, Claudio Spasiano and Vincenzo MoscatoModern recommender systems operate in dynamic environments where user preferences drift, new items arrive, and interaction patterns evolve, causing deployed models to become progressively stale. Retraining is essential to maintain recommendation quality, yet prior work has largely treated the how and when of retraining separately: adaptation strategies are evaluated under fixed schedules, while scheduling policies assume predefined updates. We formalize retraining as a sequential, resource-constrained decision problem that jointly determines when and how to update a recommender. Rather than introducing new retraining algorithms, our approach leverages existing strategies, including full retraining, fine-tuning, and sample-based updates, selecting the most effective action at each timestep. We introduce RTagent, an agent instantiated via reinforcement learning, which learns a meta-policy optimizing long-term cumulative performance under a global constraint on the number of retraining operations. Evaluation on MovieLens 1M and Yelp across three recommender architectures (SVD, CAFE, and NeuMF) shows that RTagent consistently outperforms static schedules, closely approaches full-retraining effectiveness while operating under the same budget constraint, and exhibits interpretable, architecture-specific retraining rhythms, demonstrating the benefits of sequential, strategy-aware retraining decisions.
- RESTSMOO: Solving Multi-Objective Experimentation with Constrained Thompson Sampling
by Krishna Chaitanya Kalagarla, Yi Liu, Lin Chai and Wenyang LiuTraditional online A/B experimentation limits the number of treatments that can be evaluated concurrently. Bandit-based adaptive experimentation algorithms address this by dynamically reallocating traffic across an order of magnitude more treatments. Yet existing methods involve a fundamental trade-off: single-metric Thompson Sampling-based methods are robust to novelty effects but cannot accommodate multiple launch criteria, while elimination-based multi-objective methods support multiple constraints but risk prematurely removing promising treatments. We introduce TSMOO (Thompson Sampling with Multi-Objective Optimization), a method that bridges this gap by combining multi-metric optimization with continuous learning. Grounded in stochastically constrained best-arm identification, TSMOO extends single-metric Thompson Sampling to multi-objective batch traffic allocation by estimating multi-constraint feasibility at each allocation step, while preserving all treatments throughout exploration. It further incorporates uplift modeling that mitigates temporal effects shared across treatments and control and ensures that the Gaussian distribution assumption holds. In simulations replaying historical real-world experimentation patterns, TSMOO achieves 93-94% success rates in multi-winner settings and 63-66% under novelty effects, outperforming both single-metric and elimination-based baselines.
- RSCWatchLens: A Configurable Platform for Online Video Recommendation Experiments
by Deogyong Kim and Dongha LeeStudying how video recommender systems shape user behavior requires online experiments that link playback behavior with the recommendation conditions that produced it. Existing user-study infrastructure provides one or the other, but not both within a single experimentation workflow. We present WatchLens, an open-source platform that fills this gap. WatchLens adopts a modular architecture in which user interfaces, content sources, and recommendation policies are independently configurable, with policies assignable separately to the feed and the watch page, while a standardized logging layer attaches the recommendation policy and ranking position to every event at the time of recording. This design enables researchers to analyze how recommendation policies and ranking positions shape downstream playback behavior, session continuation, and navigation between the feed and the watch page, with the linkage between policy and outcome available within each event rather than reconstructed afterwards. We demonstrate WatchLens through a short-form video case study that holds the interface and content pool constant while varying only the watch-page policy, showing how the platform supports session-level comparison of recommendation effects on real viewing behavior. WatchLens is released as a publicly available, single-server deployable system for reproducible online video recommendation research.
- INDREEF: Real-time end-to-end Explore-Exploit framework for e-commerce Feeds
by Devashish Gupta, Divay Jindal, Venkata Velugoti, Karthik Sundar, Vaishnav Chandak, Bhavuk Singhal, Vinit Rongata, Ravindra Yadav and Debdoot MukherjeeIn modern e-commerce, recommendation feeds must balance exploitation (capitalizing on known user intent) with exploration (introducing novel catalogs to prevent “filter bubbles”). However, standard ranking architectures often suffer from negative transfer, where optimizing for discovery degrades conversion performance. We propose REEF (Real-time end-to-end Explore-Exploit Framework), an intent-aware, decoupled architecture that orchestrates traffic between specialized Explore and Exploit rankers. The two rankers are bound at the data-label level by the HandShake: an emergent synchronization mechanism, arising from conditional labeling strategy that operationalizes serendipitous discoveries into conversion pathways in real-time without added inference latency. Offline experiments demonstrate that REEF improves both diversity and relevance. Currently deployed at Meesho\footnote{www.meesho.com} for ~300 million monthly active users, REEF has significantly improved structural feed diversity and long-term business metrics.
- INDTowards Reliable Social A/B Testing: Spillover-Contained Clustering with Robust Post-Experiment Analysis
by Xu Min, Zhaoxu Yang, Kaixuan Tan, Juan Yan, Xunbin Xiong, Zihao Zhu, Kaiyu Zhu, Fenglin Cui, Yang Yang, Sihua Yang and Jianhui BuA/B testing is the foundation of decision-making in online platforms, yet social products often suffer from network interference: user interactions cause treatment effects to spill over into the control group. Such spillovers bias causal estimates and undermine experimental conclusions. Existing approaches face key limitations: user-level randomization ignores network structure, while cluster-based methods often rely on general-purpose clustering that is not tailored for spillover containment and has difficulty balancing unbiasedness and statistical power at scale. We propose a spillover-contained experimentation framework with two stages. In the pre-experiment stage, we build social interaction graphs and introduce a \emph{Balanced Louvain} algorithm that produces stable, size-balanced clusters while minimizing cross-cluster edges, enabling reliable cluster-based randomization. In the post-experiment stage, we develop a tailored CUPAC estimator that leverages pre-experiment behavioral covariates to reduce the variance induced by cluster-level assignment, thereby improving statistical power. Together, these components provide both structural spillover containment and robust statistical inference. We validate our approach through large-scale social sharing experiments on Kuaishou, a platform serving hundreds of millions of users. Results show that our method substantially reduces spillover and yields more accurate assessments of social strategies than traditional user-level designs, establishing a reliable and scalable framework for networked A/B testing.
7B: Health, Society & Mobility Applications
Date: Thursday October 01, 08:30 – 10:00 CDT
Session Chair: Ding Tong
- RESImproving Rare Medication Recommendation with Counterfactual Data Augmentation and Large Language Models
by Shinhwan Kang, Soo Yong Lee, Jaewon Kim, Kijung Shin and Buru ChangAI-based medication recommendation systems have attracted substantial attention due to their potential to enhance patient safety and therapeutic outcomes. Despite the clinical importance of accurately recommending rarely prescribed medications (rare-meds), we observe that most existing methods show significantly lower predictive performance for rare-meds. We attribute this issue to two intrinsic limitations: (a) the inherent scarcity of data for rare-meds and (b) limited consideration of co-recommended medications. To address these limitations, we propose GenRxR, a novel framework based on large language models (LLMs). GenRxR leverages the medical knowledge and clinical reasoning capability of LLMs to generate counterfactual medical data, mitigating the data scarcity issue for rare-meds. It also integrates an LLM into the medication recommendation process to model relationships among co-recommended medications. To further enhance the clinical reasoning, we introduce an instruction tuning step that aligns the LLM’s capability with the recommendation task, enabling better handling of clinical context, including rare-meds cases. In our experiments, we show that GenRxR outperforms 14 (including 5 LLM-based) baselines in most cases. Specifically, it achieves up to 30.9% higher predictive performance for rare-meds than the strongest baseline.
- RESSafety-Aware Next-POI Recommendation with Large Language Models
by Rami Zaboura, Ludovico Boratto and Adir SolomonPoint of Interest (POI) recommendation has become a core task in location-based services, with modern systems increasingly driven by deep learning models that achieve strong predictive accuracy. Yet, despite these advances, most approaches optimize primarily for relevance, giving limited attention to an equally important real-world factor: user safety. In this study, we propose a safety-aware next-POI recommendation method that leverages a Large Language Model (LLM) to generate predictions informed by both mobility patterns and crime-derived safety signals. By integrating crime statistics with POI data and encoding safety information directly into trajectory prompts, our approach produces recommendations that better reflect real-world risk. Through tailored prompt engineering, we finetune an LLM to incorporate safety considerations, yielding predictions that align with user preferences while prioritizing personal security. Experimental results show that our method substantially improves the safety profile of recommended POIs and surpasses state-of-the-art baselines in overall accuracy.
- INDGuess Where You Go: Generative Next Point-of-Interest Recommendation in Amap
by Penglong Zhai, Bowen Zheng, Jie Li, Yifang Yuan, Yue Liu, Sicong Wang, Mingyang Yin, Tingting Hu, Shuaijun Guo, Fanyi Di and Xin LiGenerative retrieval enables recommender systems to retrieve items by generating compact item identifiers, but scaling it to industrial scenarios remains challenging due to redundant or colliding token assignments and insufficient integration of heterogeneous item signals. These challenges are particularly critical for next Point-of-Interest (POI) recommendation, where models must represent structured spatial entities, capture sequential mobility patterns, and produce predictions consistent with real user behavior. We propose Gwhere, an end-to-end industrial framework that integrates semantic identifier (SID) generation with LLM-based generative next POI recommendation. Gwhere first learns discriminative POI SIDs through a contrastive residual-quantization tokenizer that aligns textual, visual, spatial, and collaborative signals. Based on these SIDs, Gwhere adapts LLMs to mobility scenarios via continued pretraining on enriched spatio-temporal corpora, supervised fine-tuning, and Exposure-Aware Kahneman-Tversky Optimization (EAKTO), a reinforcement learning objective for behavioral preference alignment. Experiments on public datasets and Amap’s large-scale industrial dataset demonstrate the effectiveness of Gwhere. The system has been deployed in Amap’s homepage service under high-concurrency and low-latency constraints. Long-term online A/B tests show improvements of 5.83% in P-CTR and 6.20% in U-CTR over the production baseline. The implementation is publicly available at https://github.com/alibaba/SimCIT.
- RESTowards welfare-oriented recommendations in activity-travel behavior
by Ekin Ugurel and Takahiro YabeWhile mainstream recommender systems (RS) rely on diverse heuristics to rank alternatives, they generally lack a principled account of user welfare (i.e., whether accepting the recommendation will leave the user better off than other alternatives). The problem is particularly acute in activity-based travel behavior, where users incur costs they cannot recoup (i.e., energy, time) regardless of eventual satisfaction. As a result, existing systems may recommend options based on popularity or collaborative filtering, but may still leave users worse off than nearby or self-selected alternatives. We address this gap by introducing a welfare-oriented framework for activity recommendation that evaluates suggestions in terms of net utility, defined as experienced benefit minus travel costs. Specifically, we formalize two operational decision criteria: Positive Utility Probability (PUP) recommends only when the probability of non-negative net utility exceeds a threshold, while Regret Minimization (RM) recommends only when expected regret relative to the user’s best organic alternative falls below a tolerance level. To evaluate these criteria, we develop an agent-based simulation in which heterogeneous synthetic travelers interact with multiple RS over time in a spatial environment with realistic travel costs, congestion, and behavioral feedback loops. This framework enables controlled counterfactual evaluations, and offers a practical foundation for designing RS that treat user welfare as a primary objective rather than an incidental byproduct.
- PPFSome People Are Worth Recommending For: Twenty Years of Progress, Gaps, and Opportunities For and With Children at RecSys
by Robin Ungruh and Maria Soledad PeraRecommender systems play a pivotal role in shaping children’s digital ecosystems. Despite substantial progress in recommender systems research over the past two decades, work that considers children as the main stakeholder, or even acknowledges them as a user group, remains in its infancy. Taking the 20th anniversary of the ACM RecSys conference as a milestone for reflection, we present a comprehensive review of 20 years of research published at RecSys and its co-located workshops, mapping the evolution of research related to children. Through this historical lens, we take inventory of the community’s contributions and identify unique challenges that emerge when designing, evaluating, and deploying recommenders for young people. Drawing on these insights, we outline critical gaps in the current landscape and propose a research agenda for advancing child-aware recommender system research.
- RESReducing Perceived Polarization through Affect-Balanced News Reframing
by Jia-Hua Jeng, Alain D. Starke, David Elsweiler and Christoph TrattnerEmotionally charged news headlines can amplify negative reactions and increase perceptions of societal polarization. In this paper, we investigate whether large language models can be used to reframe news headlines in a more affect-balanced manner, combining fear and hope, without undermining engagement or perceived fairness. Across three studies, we analyze emotional framing in news headlines using large-scale interaction data, validate that LLMs can reliably generate fear-hope reframings, and evaluate their effects in a controlled user study (N = 80). Our results show that fear-hope reframed headlines significantly reduce perceived polarization compared to original headlines (roughly a 20-25 decrease relative to the original condition’s mean, substantial given the brief exposure), while reducing negative emotional states such as anger and hostility. At the same time, we find no significant decrease in engagement intentions, perceived fairness, or willingness to pay for news. Emotional effects are not uniform, indicating that LLM-based reframing changes specific emotional responses rather than broadly reducing emotionality. Together, these findings suggest that carefully designed LLM-driven headline reframing can reduce perceived polarization and some negative emotional responses while preserving key engagement-related outcomes.
- INDTowards Full Candidate Interaction: A Comprehensive Comparison Network for Better Route Recommendation
by Hanyu Guo, Chao Chen, Longfei Xu, Chengzhang Wang, Kaikui Liu and Xiangxiang ChuWe argue that the decision-making essence of route recommendation is comparative judgment: users choose a route because it is better than alternatives in specific aspects. The decision-critical information resides in segment-level spatial differences of non-overlapping parts between routes, which is irreversibly lost through item-level feature aggregation. Existing methods, whether attention-based or pairwise ranking approaches, follow an item-first paradigm that can only infer pairwise relations indirectly from individual route representations. To address this, we propose the Comprehensive Comparison Network (CCN), which inverts the information flow by constructing explicit comparison features from non-overlapping segments between route pairs and reasoning directly in the pairwise space. CCN introduces a Comprehensive Comparison Block that enables context-aware pairwise reasoning, where the comparison between two routes is informed by how both compare against all other candidates. We further develop an interpretable Pair Scoring Network that decomposes pairwise preferences into independent physical fields, providing field-level explanations for route selection. CCN has served as the production ranking model in Amap for over two years, achieving 85.70% offline route-trajectory coverage rate and +1.2% online improvement over the previous production model. We also release a large-scale route recommendation dataset comprising 175 million users, 512 million samples, and 6 billion routes across 370 cities.
RecSys 2026 (Minneapolis)
- About the Conference
- Registration
- Program at Glance
- Program
- Call for Contributions
- Challenge
- Keynotes
- Accepted Contributions
- Presenter Instructions
- Workshops
- Tutorials
- Committees
- Inclusion
- Student Volunteers
- Women in RecSys
- Visa Information
- Addressing Attendance Issues
- Location / Hotel
- Lasting Impact Award
- 60 Milestones for 20 Years




















