Session 7:

7A: Online Experimentation & Exploration

Date: Thursday October 01, 08:30 – 10:00 CDT
Session Chair: Morten Arngren

  • RESProbabilistic Residual Learning for Online Recommendations
    by Wenyuan Wang, Yusong Zhao, Zihao Xu, Hengyi Wang, Qi Xu, Zhigang Hua, Yan Xie, Yi Wang, Zihao Zhao, Bo Long, Chengzhi Mao, Shuang Yang, Hengguan Huang and Hao Wang

    Modern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and items. As a result, these systems often suffer from the black-box nature and computational complexity of the underlying models, making it difficult to systematically enhance their recommendation capabilities. To address this problem, we propose Probabilistic Residual Learning (PRL), a causal Bayesian recommendation model that models the residual between ground-truth and base predictions, enabling targeted refinement of existing systems. Specifically, PRL (1) divides users into clusters in an unsupervised manner and identifying causal confounders that influence latent variables, (2) learns sub-models for each confounder given the observable variables, and (3) generates recommendations by aggregating the rating residuals under each confounder using do-calculus. Experiments demonstrate that our plug-and-play PRL is compatible with various base DL recommender systems, improving their performance while automatically discovering meaningful user clusters. Auxiliary materials (including the Appendix) are at https://anonymous.4open.science/r/PRL_Appendix-CDD7/PRL_RecSys_Appendix.pdf.

  • RESAdaptive Retraining of Recommender Systems via Reinforcement Learning
    by Diego Russo, Valerio La Gatta, Claudio Spasiano and Vincenzo Moscato

    Modern recommender systems operate in dynamic environments where user preferences drift, new items arrive, and interaction patterns evolve, causing deployed models to become progressively stale. Retraining is essential to maintain recommendation quality, yet prior work has largely treated the how and when of retraining separately: adaptation strategies are evaluated under fixed schedules, while scheduling policies assume predefined updates. We formalize retraining as a sequential, resource-constrained decision problem that jointly determines when and how to update a recommender. Rather than introducing new retraining algorithms, our approach leverages existing strategies, including full retraining, fine-tuning, and sample-based updates, selecting the most effective action at each timestep. We introduce RTagent, an agent instantiated via reinforcement learning, which learns a meta-policy optimizing long-term cumulative performance under a global constraint on the number of retraining operations. Evaluation on MovieLens 1M and Yelp across three recommender architectures (SVD, CAFE, and NeuMF) shows that RTagent consistently outperforms static schedules, closely approaches full-retraining effectiveness while operating under the same budget constraint, and exhibits interpretable, architecture-specific retraining rhythms, demonstrating the benefits of sequential, strategy-aware retraining decisions.

  • RESTSMOO: Solving Multi-Objective Experimentation with Constrained Thompson Sampling
    by Krishna Chaitanya Kalagarla, Yi Liu, Lin Chai and Wenyang Liu

    Traditional online A/B experimentation limits the number of treatments that can be evaluated concurrently. Bandit-based adaptive experimentation algorithms address this by dynamically reallocating traffic across an order of magnitude more treatments. Yet existing methods involve a fundamental trade-off: single-metric Thompson Sampling-based methods are robust to novelty effects but cannot accommodate multiple launch criteria, while elimination-based multi-objective methods support multiple constraints but risk prematurely removing promising treatments. We introduce TSMOO (Thompson Sampling with Multi-Objective Optimization), a method that bridges this gap by combining multi-metric optimization with continuous learning. Grounded in stochastically constrained best-arm identification, TSMOO extends single-metric Thompson Sampling to multi-objective batch traffic allocation by estimating multi-constraint feasibility at each allocation step, while preserving all treatments throughout exploration. It further incorporates uplift modeling that mitigates temporal effects shared across treatments and control and ensures that the Gaussian distribution assumption holds. In simulations replaying historical real-world experimentation patterns, TSMOO achieves 93-94% success rates in multi-winner settings and 63-66% under novelty effects, outperforming both single-metric and elimination-based baselines.

  • RSCWatchLens: A Configurable Platform for Online Video Recommendation Experiments
    by Deogyong Kim and Dongha Lee

    Studying how video recommender systems shape user behavior requires online experiments that link playback behavior with the recommendation conditions that produced it. Existing user-study infrastructure provides one or the other, but not both within a single experimentation workflow. We present WatchLens, an open-source platform that fills this gap. WatchLens adopts a modular architecture in which user interfaces, content sources, and recommendation policies are independently configurable, with policies assignable separately to the feed and the watch page, while a standardized logging layer attaches the recommendation policy and ranking position to every event at the time of recording. This design enables researchers to analyze how recommendation policies and ranking positions shape downstream playback behavior, session continuation, and navigation between the feed and the watch page, with the linkage between policy and outcome available within each event rather than reconstructed afterwards. We demonstrate WatchLens through a short-form video case study that holds the interface and content pool constant while varying only the watch-page policy, showing how the platform supports session-level comparison of recommendation effects on real viewing behavior. WatchLens is released as a publicly available, single-server deployable system for reproducible online video recommendation research.

  • INDREEF: Real-time end-to-end Explore-Exploit framework for e-commerce Feeds
    by Devashish Gupta, Divay Jindal, Venkata Velugoti, Karthik Sundar, Vaishnav Chandak, Bhavuk Singhal, Vinit Rongata, Ravindra Yadav and Debdoot Mukherjee

    In modern e-commerce, recommendation feeds must balance exploitation (capitalizing on known user intent) with exploration (introducing novel catalogs to prevent “filter bubbles”). However, standard ranking architectures often suffer from negative transfer, where optimizing for discovery degrades conversion performance. We propose REEF (Real-time end-to-end Explore-Exploit Framework), an intent-aware, decoupled architecture that orchestrates traffic between specialized Explore and Exploit rankers. The two rankers are bound at the data-label level by the HandShake: an emergent synchronization mechanism, arising from conditional labeling strategy that operationalizes serendipitous discoveries into conversion pathways in real-time without added inference latency. Offline experiments demonstrate that REEF improves both diversity and relevance. Currently deployed at Meesho\footnote{www.meesho.com} for ~300 million monthly active users, REEF has significantly improved structural feed diversity and long-term business metrics.

  • INDTowards Reliable Social A/B Testing: Spillover-Contained Clustering with Robust Post-Experiment Analysis
    by Xu Min, Zhaoxu Yang, Kaixuan Tan, Juan Yan, Xunbin Xiong, Zihao Zhu, Kaiyu Zhu, Fenglin Cui, Yang Yang, Sihua Yang and Jianhui Bu

    A/B testing is the foundation of decision-making in online platforms, yet social products often suffer from network interference: user interactions cause treatment effects to spill over into the control group. Such spillovers bias causal estimates and undermine experimental conclusions. Existing approaches face key limitations: user-level randomization ignores network structure, while cluster-based methods often rely on general-purpose clustering that is not tailored for spillover containment and has difficulty balancing unbiasedness and statistical power at scale. We propose a spillover-contained experimentation framework with two stages. In the pre-experiment stage, we build social interaction graphs and introduce a \emph{Balanced Louvain} algorithm that produces stable, size-balanced clusters while minimizing cross-cluster edges, enabling reliable cluster-based randomization. In the post-experiment stage, we develop a tailored CUPAC estimator that leverages pre-experiment behavioral covariates to reduce the variance induced by cluster-level assignment, thereby improving statistical power. Together, these components provide both structural spillover containment and robust statistical inference. We validate our approach through large-scale social sharing experiments on Kuaishou, a platform serving hundreds of millions of users. Results show that our method substantially reduces spillover and yields more accurate assessments of social strategies than traditional user-level designs, establishing a reliable and scalable framework for networked A/B testing.

Back to program