Session 4:
4A: Agentic Recommendation & Autonomous Optimization
Date: Wednesday September 30, 10:45 – 12:15 CDT
Session Chair: Joeran Beel
- PPFA Position Paper on Recommender Systems in the Era of Autonomous Agents
by Aixin SunFor decades, recommender systems (RecSys) have been optimized to serve human users, leveraging behavioral data to predict preferences. However, the rapid deployment of autonomous agents powered by large language models (LLMs) introduces a paradigm shift: recommendation consumers are predicted to be increasingly a mixture of humans and authorized agents acting on their behalf, as a client-side proxy. This position paper reviews insights from prior human-centric RecSys research and outlines the transition to this more complex setting. We characterize the resulting tripartite interactions among humans, agents, and platforms, highlighting the dynamics that can arise across these relationships. This shift presents new opportunities for RecSys research, while introducing unique challenges for evaluation, alignment, and system design.
- RESBiLPR: Bidirectional Teacher-Student Agent Interaction for Context-Aware Learning Path Recommendation
by Zejun Chen, Weiwei Chen, Suojuan Zhang, Zhi Zheng, Dawei Jin, Ziwei Zhao, Tong Xu, Jing Cui, Jiaqi Long and Enhong ChenLearning path recommendation is a critical component of intelligent education systems, aiming to plan a personalized sequence of learning resources for each student based on their cognitive state. Existing methods predominantly rely on unidirectional modeling for recommendations, failing to adequately capture the bidirectional interaction between teachers and students. This leads to a lack of feedback-driven adaptation and difficulty in forming an effective instructional closed loop. Furthermore, current learning path recommendations are often limited to static student-exercise matching. They cannot perceive and respond to dynamic learning contexts, which results in insufficient adaptability. This limitation stems from an inadequate consideration of key contextual factors, including real-time cognitive states, interaction history, exercise semantics, and knowledge structures. To address these issues, this paper proposes a Bidirectional Teacher-Student Agent Interaction for Context-Aware Learning Path Recommendation (BiLPR), which implements a bidirectional, dynamic, and synergistic process. Specifically, the Teacher Agent integrates domain knowledge graphs with semantic reasoning to thoroughly mine features of the learning context. This enables dynamic exercise adaptation and recommendation strategies underpinned by knowledge transfer. The Student Agent simulates the evolution of dynamic cognitive states and behaviors during authentic learning processes, providing feedback on its performance. This interaction establishes a novel iterative closed loop of recommendation, feedback, and reflection. Evaluated on two real-world educational datasets, Junyi and ASSIST2009, the proposed method significantly outperforms baseline models in recommendation effectiveness. The code is available at https://anonymous.4open.science/r/BiLPR-D16D.
- RESRecRec: Latent Interests Recursive Reasoning for Sequential Recommendation
by Wenhao Deng, Junchen Fu, Hanwen Du, Alexandros Karatzoglou, Ioannis Arapakis, Hangjun Guo, Kaiwen Zheng, Yongxin Ni and Joemon JoseSequential recommender systems rely on a single forward pass to encode user interaction histories and predict the next item. Increasing inference-time computation through latent reasoning, with the model proceeding step by step before the final prediction, has been recently explored in sequential recommendation with promising results. However, how to structure the reasoning process for sequential recommendation remains an open question. Existing approaches couple reasoning and prediction in a single d-dimensional state, limiting reasoning depth and often relying on multi-stage pipelines with reinforcement learning. We propose RecRec (Recursive Reasoning for Recommendation), an RL-free framework that decouples reasoning from prediction, overcoming the fixed d-dimensional state bottleneck of prior methods. RecRec consists of a Context Compressor and a Recursive Reasoner, trained in two simple supervised stages. The Context Compressor distills the backbone’s hidden states into a small set of latent interests, with an Interest Diversity Regularizer encouraging each interest to capture a distinct aspect of user behavior. The Recursive Reasoner then refines these interests by reasoning in a separate intermediate latent space. Deep supervision lets the reasoning depth be freely adjusted at inference without retraining. On four real-world datasets, RecRec outperforms state-of-the-art reasoning-enhanced methods, and on three of four datasets, gains extend past the training-time depth. Our findings point to a decoupled, multi-vector recipe that unleashes latent reasoning from the single-state bottleneck of prior methods, suggesting reasoning-state structure as a design axis to explore further in sequential recommendation.
- INDWhich LLM to Fine-Tune? Agent-Driven Model Selection at Scale
by Chen Luo, Yulin Liu, Yi Liu, Xuejing Lei, Yuchen Yan, Xin Zhang, Huimin Zeng, Hongda Mao and Monica ChengOpen-source model hubs now host over two million public AI models, yet teams building customer-facing AI systems must still determine which model to fine-tune for production deployment—a decision that shapes the quality, latency, and cost experienced by hundreds of millions of users. At Amazon, we spent over a year iterating on this process across multiple production use cases, where model selection remained manual, slow, and heavily biased toward a small set of familiar model families despite the rapidly expanding open-source ecosystem. We show that model selection is a recommendation problem, and introduce \textsc{AgentRec}, a multi-stage retrieval-and-ranking framework that progressively narrows hundreds of candidate models using increasingly expensive but more faithful evaluation signals. Across public benchmarks and Amazon production systems, AgentRec reduces model selection from multi scientist-weeks to couple unattended GPU-hours while matching or exceeding the quality of exhaustive manual exploration. Our results suggest that, for industrial teams deploying fine-tuned LLMs at scale, model selection can evolve from an ad-hoc bottleneck into a repeatable and continuously automated system for discovering high-quality models under real-world deployment constraints.
- INDSelf-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents
by Haochen Wang, Yi Wu, Daryl Chang, Li Wei and Lukasz HeldtOptimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyperparameter search space and, more critically, designing sophisticated optimizers, architectures, and reward functions to capture nuanced user behaviors. Achieving substantial improvements in these areas is a non-trivial task, traditionally relying on extensive manual iterations to test new hypotheses. We propose a self-evolving system that leverages Large Language Models (LLMs), specifically those from Google’s Gemini family, to autonomously generate, train, and deploy high-performing, complex model changes within an end-to-end automated workflow. The self-evolving system consists of an Offline Agent (Fast Loop) that performs high-throughput hypothesis generation to optimize for proxy metrics, and an Online Agent (Slow Loop) that validates candidates against delayed north star business metrics in live production. Our agents act as specialized Machine Learning Engineers (MLEs): they exhibit deep reasoning capabilities, discovering novel improvements in optimization algorithms and model architecture, and formulating innovative reward functions that target long-term user engagement. The effectiveness of this approach is demonstrated through several successful production launches at YouTube, confirming that autonomous, LLM-driven evolution can surpass traditional engineering workflows in both development velocity and model performance.
- INDGRIP: Generation and Reasoning for User Profile Completion
by Riwei Lai, Yu Xia, Li Chen, Beibei Kong, Lei Cheng, Chengxiang Zhuo, Zang Li and Chenyun YuUser profiles, such as age and interest tags, form the backbone of modern recommender systems. However, in real-world scenarios, user profiles frequently encounter the problem of incomplete profile data, restricting the effectiveness of downstream recommendation tasks. Although large language models (LLMs) have shown remarkable potential in understanding user profiles, existing methods mainly focus on extracting explicit user information from external data and generating profile summaries. Such shallow reasoning patterns often fail to accurately infer missing user attributes under sparse data conditions. In this paper, we present GRIP (Generation and Reasoning from Incomplete Profiles), the first foundation model tailored for comprehensive profile inference under data sparsity for recommendation. GRIP introduces a unified three-stage training framework: (1) continual pre-training on large-scale structured user data for robust correlation modeling; (2) decomposed chain-of-thought reasoning with iterative self-distillation to facilitate in-depth profile inference; and (3) reinforcement learning with multi-dimensional rewards to jointly optimize factuality and reasoning coherence. Comprehensive experiments on realistic benchmarks along with online business deployments have demonstrated that GRIP significantly outperforms LLM-based methods in completing missing attributes and delivers substantial business gains in downstream recommendation tasks.
- INDA Self-Triggered Agentic Push Recommendation System
by Zhao-Yu Zhang, Qingying Chen, Chunyuan Zheng, Jing Zhou, Jian Sun, Siqi Chen, Leiying Chen, Chuan Zhou, Huiyou Jiang, Xin Tao, Haoxuan Li and Zhouchen LinPush notification is a critical recommendation scenario on large-scale platforms, allowing the system to proactively reach users outside the application to improve long-term re-engagement. However, designing an optimal push system requires handling a complex action space for the “whether and when” delivery problem under strict system resource constraints, i.e., systems cannot naively evaluate every user at every second to find the optimal delivery moment. Existing solutions typically fall into two paradigms: the first is a two-stage approach that relies on offline user-level uplift modeling combined with integer programming solvers to allocate optimal delivery times and frequencies. The second periodically activates the push system to decide whether to send a push. However, the former cannot effectively utilize real-time information, and the latter results in resource inefficiency. In addition, such a multi-stage solution suffers from local optima. In this paper, we propose a proactive, self-triggered end-to-end agentic push recommendation system, which is already fully deployed at Douyin with over 1 billion users, allowing the system to generate push time and deternmine whether to send in a closed loop with both real-time effectiveness and efficiency. Specifically, the agentic system consists of two decision transformer-based agents: a planning agent that determines when to schedule the next push time with a gated ordinal regression method, and an action-execution agent that decides whether to send a push based on a trajectory reward. In addition, we further introduce a lightweight filtering agent to both control the computational overhead and act as a crucial safeguard against unreasonable planning behaviors. Extensive experiments demonstrate that this approach can both increase user active days by 0.3% and reduce push permission disablement by 1.9%. In addition, adding the filtering agent can reduce the computational overhead by 74%.
4B: Fairness, Privacy & Safety
Date: Wednesday September 30, 10:45 – 12:15 CDT
Session Chair: Robin Burke
- RESGTP: Mitigating Popularity Bias in PLM-based Sequential Recommendation via Group-Aware Token Pruning
by Ruilin Yuan, Dugang Liu, Hao Chen and Zhong MingIncorporating item textual information with pretrained language models (PLMs) still incurs popularity bias at the semantic representation level. Although existing studies have explored debiasing at the token level, most of them treat tokens in a uniform manner based on item inputs, without finely distinguishing the distributional differences between unique and shared tokens in popular and tail items. Consequently, it remains difficult to characterize the heterogeneous impact of different tokens on popularity bias. To address this issue, we first conduct a validation study to investigate token distribution patterns in popular and tail items. Our findings reveal that unique tokens are more likely to introduce bias than shared tokens, primarily driven by the distributional discrepancies between popular-unique and tail-unique tokens. Based on this observation, we find that strategically pruning biased tokens can enhance item exposure fairness without sacrificing recommendation accuracy. Accordingly, we propose GTP, a group-aware token pruning framework with an adaptive selection network. GTP explicitly categorizes tokens into popular-unique, tail-unique, and shared groups, and learns differentiated retention strategies for different types of unique tokens, thereby effectively mitigating popularity bias induced by unique tokens. Extensive experiments on three real-world datasets demonstrate that GTP consistently alleviates popularity bias while maintaining recommendation performance.
- RESFrom Constraint to Control: Modeling Expected Fairness in Ranking Systems
by Tristan Cladière, Antoine Gourru, Bissan Audeh and Christine LargeronModern learning to rank systems have achieved remarkable performance across a wide range of applications. However, they may also exhibit disparities in exposure, raising concerns about fairness, especially in sensitive domains such as healthcare, judicial decision-making, and recruitment, where biased rankings may have critical societal consequences. A common approach to mitigate such issues is to incorporate fairness constraints or regularization terms into the training objective. Yet, this provides limited insight into how these constraints influence the final level of bias, and consequently require costly hyperparameter tuning to reach a desired fairness outcome. In this work, we study fairness in ranking by leveraging threshold-based constraints on disparate exposure, which, under a distributional approximation, induce a predictable transformation of the exposure distribution. We derive a closed-form expression for the expected disparate exposure as a function of the threshold, and introduce an anchored formulation that accounts for practical optimization limits. This formulation enables practitioners to directly select a threshold that achieves a desired fairness target, eliminating the need for extensive hyperparameter tuning. An experimental evaluation on standard learning to rank benchmarks confirms that the proposed model closely matches empirical behavior. These results demonstrate that fairness can be explicitly modeled, predicted and controlled, providing novel and sound approach to tuning fairness in ranking systems. Finally, we disclose our source code for full reproducibility.
- RESA Redundancy Reduction Approach for Controllable Sequential Recommendations
by Veronika Ivanova, Marina Munkhoeva, Ivan Razvorotnev and Evgeny FrolovSequential recommendation must operate under long-tailed item distributions and popularity-driven concentration, often forcing practitioners to trade short-list accuracy against long-tail exposure. In this work, we study feature decorrelation as a mechanism for shaping representation geometry in dot-product sequential recommenders, and analyze how this, in turn, affects popularity-driven concentration. We propose a decorrelation-regularized training framework that augments next-item prediction with an auxiliary redundancy-reduction term, and instantiate it with BT-SR, which uses the Barlow Twins objective. To form label-consistent positive pairs without synthetic corruptions, we pair user histories that share the same next-item target. Beyond accuracy, we provide a geometric analysis showing how decorrelation suppresses shared low-rank directions in the user representation space that can give popular items a global scoring advantage, and we introduce a bucket-based alignment concentration metric to quantify this effect. Experiments on five public benchmarks show that BT-SR consistently improves next-item ranking quality, while the decorrelation strength acts as a simple control knob that reallocates accuracy across head and tail items, enabling accuracy–exposure trade-offs. Our analysis also reveals that the impact on head-vs-tail exposure differs across datasets, reflecting interactions between decorrelation and data temporal structure.
- RESGive the Long-tail More SPACE: Promoting Provider Fairness in Next POI Recommendation
by Anran Zhang, Jiaqi Jiang, Jiahui Jin and Yuhan ZhaoNext point-of-interest (POI) recommendation predicts users’ future destinations from historical mobility sequences and has become a key component of location-based services. However, mainstream models often concentrate exposure on a small set of popular POIs, leaving long-tail merchants systematically under-exposed. While provider fairness has recently attracted increasing attention, directly applying existing provider-fairness techniques to POI recommendation is problematic: (i) users face execution constraints; and (ii) POIs face resource supply constraints. These coupled constraints render provider fairness in POI recommendation a fundamentally different—and more challenging—problem than in purely digital settings. To address this, we propose SPACE (Supply- and Physics-Aware Conditional Embedding generation), a model-agnostic framework that improves long-tail POI exposure via virtual user generation under explicit feasibility and supply control. SPACE consists of three stages: (1) community inference to capture heterogeneous user execution constraints; (2) unbalanced optimal-transport allocation to decide how many virtual users each tail POI should receive from which communities under POI-specific supply budgets; and (3) constraint-guided latent diffusion to generate POI-conditional, community-consistent virtual user embeddings. The generated user–POI pairs can be seamlessly used to train existing recommenders without modifying their architectures. Extensive experiments on three real-world datasets demonstrate that SPACE substantially improves provider fairness while maintaining—and often improving—recommendation accuracy across multiple backbone models. Our code is publicly available at https://anonymous.4open.science/r/anonym046A/.
- RESMembership Inference Attacks on In-Context Learning Recommendation
by Jiajie He, Min-Chun Chen, Xintong Chen, Xinyang Fang, Yuechun Gu and Keke ChenLarge language models (LLMs) based recommender systems (RecSys) can adapt flexibly across different domains. It uses in-context learning (ICL), i.e., prompts, including sensitive historical user-specific item interactions, to customize the recommendation functions. However, no study has examined whether such private information may be exposed by novel privacy attacks. We design two membership inference attacks (MIAs): ItemMem, and RecInertia, aiming to identify whether system prompts contain the victim’s information. We have carefully evaluated them on the latest open-source LLMs and three well-known RecSys datasets. The results confirm that the MIA threat to LLM RecSys is realistic and can be more sophisticated than prompt extraction. They utilize the unique prompt structures in ICL RecSys and cannot be easily mitigated with existing defense methods on prompt extraction.
- RESDecoupled Learning and Selection in Slate Recommendation for Privacy and Stability Under Noisy Scores
by Sam Urmian, Qinyi Liu and Mohammad KhalilMany recommender systems do not show users the raw list produced by a learned model. They first score possible items, then apply a repeatable rule layer that removes restricted items, adds variety, enforces constraints, and decides how many items to show. We study what can be said when these two steps are separated explicitly. The aim is not to propose a new recommender, but to understand which privacy, auditability, and stability guarantees follow from this common design pattern. Our main result is a certificate for when the final recommendation list stays unchanged. If the recorded gap between each chosen item and the closest alternative is large enough, then small changes in learned scores cannot change the selected list. This gives a practical way to audit whether a recommendation was robust to noisy scores. The result also explains why mixing a changing model with a fixed reference model can reduce top-list changes under score noise. This stability claim is separate from privacy: if the learned model is trained with a formal privacy guarantee and the rule layer uses only public, fixed, or separately privacy-accounted inputs, then the full system and its audit log inherit that guarantee. If the fixed reference model is trained without privacy protection, the system may still be more stable, but it is not private end-to-end. We test these claims in the settings where they apply. Controlled top-list change tests match the predicted stability pattern, and ranking-change tests on OULAD, MovieLens-25M, and Amazon Musical Instruments show the same stable-reference/noisy-model effect. OULAD and EdNet margin diagnostics find certified cases with no slate changes. Simulated repeated recommendation runs on OULAD and EdNet show that selector rules can bound target drift and make decision changes replayable, while final user-facing utility effects remain mixed. Overall, the paper characterizes what follows from separating learning from repeatable rule-based selection: certifiable stability and scoped privacy claims, not a universally best recommender.
- INDConAlign: Conditional Alignment Framework for Balancing Biased and Unbiased Recommendation
by Jingcheng Zhang, Yihan Wang, Qi Song and Liyin HongIndustry recommender systems trained on observational data suffer from various biases that create filter bubbles, causing user interests to collapse into narrow categories and severely degrading long-term engagement. While utilizing unbiased uniform data for debiasing has shown promise, existing methods remain impractical for industrial deployment due to limitations such as neglect of factual(biased) recommendation performance and the substantial computational overhead. To overcome these limitations, we propose ConAlign (Conditional Alignment Framework), a conditional debiasing approach for industrial deployment. The key innovation of ConAlign lies in a discrete gating-based conditional alignment mechanism that selectively transfers knowledge from the biased tower to the unbiased tower. Following a selective intervention paradigm rather than universal correction, it seamlessly balances factual accuracy and unbiased preference estimation while supporting real-time streaming adaptation. To our knowledge, based on publicly available literature, ConAlign is the first architecture successfully deployed in a large-scale industrial system that utilizes a small fraction of unbiased random traffic for debiasing. Extensive offline experiments on three real-world datasets rigorously validate the effectiveness of our proposed framework, and the code is publicly available at \url{https://github.com/JcZhangzz/ConAlign}. Furthermore, large-scale online A/B testing on Kuaishou demonstrates significant improvements in long-term user engagement and interest diversity, with negligible latency overhead.
RecSys 2026 (Minneapolis)
- About the Conference
- Registration
- Program at Glance
- Program
- Call for Contributions
- Challenge
- Keynotes
- Accepted Contributions
- Presenter Instructions
- Workshops
- Tutorials
- Committees
- Inclusion
- Student Volunteers
- Women in RecSys
- Visa Information
- Addressing Attendance Issues
- Location / Hotel
- Lasting Impact Award
- 60 Milestones for 20 Years




















