Session 4:

4A: Agentic Recommendation & Autonomous Optimization

Date: Wednesday September 30, 10:45 – 12:15 CDT
Session Chair: Joeran Beel

  • PPFA Position Paper on Recommender Systems in the Era of Autonomous Agents
    by Aixin Sun

    For decades, recommender systems (RecSys) have been optimized to serve human users, leveraging behavioral data to predict preferences. However, the rapid deployment of autonomous agents powered by large language models (LLMs) introduces a paradigm shift: recommendation consumers are predicted to be increasingly a mixture of humans and authorized agents acting on their behalf, as a client-side proxy. This position paper reviews insights from prior human-centric RecSys research and outlines the transition to this more complex setting. We characterize the resulting tripartite interactions among humans, agents, and platforms, highlighting the dynamics that can arise across these relationships. This shift presents new opportunities for RecSys research, while introducing unique challenges for evaluation, alignment, and system design.

  • RESBiLPR: Bidirectional Teacher-Student Agent Interaction for Context-Aware Learning Path Recommendation
    by Zejun Chen, Weiwei Chen, Suojuan Zhang, Zhi Zheng, Dawei Jin, Ziwei Zhao, Tong Xu, Jing Cui, Jiaqi Long and Enhong Chen

    Learning path recommendation is a critical component of intelligent education systems, aiming to plan a personalized sequence of learning resources for each student based on their cognitive state. Existing methods predominantly rely on unidirectional modeling for recommendations, failing to adequately capture the bidirectional interaction between teachers and students. This leads to a lack of feedback-driven adaptation and difficulty in forming an effective instructional closed loop. Furthermore, current learning path recommendations are often limited to static student-exercise matching. They cannot perceive and respond to dynamic learning contexts, which results in insufficient adaptability. This limitation stems from an inadequate consideration of key contextual factors, including real-time cognitive states, interaction history, exercise semantics, and knowledge structures. To address these issues, this paper proposes a Bidirectional Teacher-Student Agent Interaction for Context-Aware Learning Path Recommendation (BiLPR), which implements a bidirectional, dynamic, and synergistic process. Specifically, the Teacher Agent integrates domain knowledge graphs with semantic reasoning to thoroughly mine features of the learning context. This enables dynamic exercise adaptation and recommendation strategies underpinned by knowledge transfer. The Student Agent simulates the evolution of dynamic cognitive states and behaviors during authentic learning processes, providing feedback on its performance. This interaction establishes a novel iterative closed loop of recommendation, feedback, and reflection. Evaluated on two real-world educational datasets, Junyi and ASSIST2009, the proposed method significantly outperforms baseline models in recommendation effectiveness. The code is available at https://anonymous.4open.science/r/BiLPR-D16D.

  • RESRecRec: Latent Interests Recursive Reasoning for Sequential Recommendation
    by Wenhao Deng, Junchen Fu, Hanwen Du, Alexandros Karatzoglou, Ioannis Arapakis, Hangjun Guo, Kaiwen Zheng, Yongxin Ni and Joemon Jose

    Sequential recommender systems rely on a single forward pass to encode user interaction histories and predict the next item. Increasing inference-time computation through latent reasoning, with the model proceeding step by step before the final prediction, has been recently explored in sequential recommendation with promising results. However, how to structure the reasoning process for sequential recommendation remains an open question. Existing approaches couple reasoning and prediction in a single d-dimensional state, limiting reasoning depth and often relying on multi-stage pipelines with reinforcement learning. We propose RecRec (Recursive Reasoning for Recommendation), an RL-free framework that decouples reasoning from prediction, overcoming the fixed d-dimensional state bottleneck of prior methods. RecRec consists of a Context Compressor and a Recursive Reasoner, trained in two simple supervised stages. The Context Compressor distills the backbone’s hidden states into a small set of latent interests, with an Interest Diversity Regularizer encouraging each interest to capture a distinct aspect of user behavior. The Recursive Reasoner then refines these interests by reasoning in a separate intermediate latent space. Deep supervision lets the reasoning depth be freely adjusted at inference without retraining. On four real-world datasets, RecRec outperforms state-of-the-art reasoning-enhanced methods, and on three of four datasets, gains extend past the training-time depth. Our findings point to a decoupled, multi-vector recipe that unleashes latent reasoning from the single-state bottleneck of prior methods, suggesting reasoning-state structure as a design axis to explore further in sequential recommendation.

  • INDWhich LLM to Fine-Tune? Agent-Driven Model Selection at Scale
    by Chen Luo, Yulin Liu, Yi Liu, Xuejing Lei, Yuchen Yan, Xin Zhang, Huimin Zeng, Hongda Mao and Monica Cheng

    Open-source model hubs now host over two million public AI models, yet teams building customer-facing AI systems must still determine which model to fine-tune for production deployment—a decision that shapes the quality, latency, and cost experienced by hundreds of millions of users. At Amazon, we spent over a year iterating on this process across multiple production use cases, where model selection remained manual, slow, and heavily biased toward a small set of familiar model families despite the rapidly expanding open-source ecosystem. We show that model selection is a recommendation problem, and introduce \textsc{AgentRec}, a multi-stage retrieval-and-ranking framework that progressively narrows hundreds of candidate models using increasingly expensive but more faithful evaluation signals. Across public benchmarks and Amazon production systems, AgentRec reduces model selection from multi scientist-weeks to couple unattended GPU-hours while matching or exceeding the quality of exhaustive manual exploration. Our results suggest that, for industrial teams deploying fine-tuned LLMs at scale, model selection can evolve from an ad-hoc bottleneck into a repeatable and continuously automated system for discovering high-quality models under real-world deployment constraints.

  • INDSelf-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents
    by Haochen Wang, Yi Wu, Daryl Chang, Li Wei and Lukasz Heldt

    Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyperparameter search space and, more critically, designing sophisticated optimizers, architectures, and reward functions to capture nuanced user behaviors. Achieving substantial improvements in these areas is a non-trivial task, traditionally relying on extensive manual iterations to test new hypotheses. We propose a self-evolving system that leverages Large Language Models (LLMs), specifically those from Google’s Gemini family, to autonomously generate, train, and deploy high-performing, complex model changes within an end-to-end automated workflow. The self-evolving system consists of an Offline Agent (Fast Loop) that performs high-throughput hypothesis generation to optimize for proxy metrics, and an Online Agent (Slow Loop) that validates candidates against delayed north star business metrics in live production. Our agents act as specialized Machine Learning Engineers (MLEs): they exhibit deep reasoning capabilities, discovering novel improvements in optimization algorithms and model architecture, and formulating innovative reward functions that target long-term user engagement. The effectiveness of this approach is demonstrated through several successful production launches at YouTube, confirming that autonomous, LLM-driven evolution can surpass traditional engineering workflows in both development velocity and model performance.

  • INDGRIP: Generation and Reasoning for User Profile Completion
    by Riwei Lai, Yu Xia, Li Chen, Beibei Kong, Lei Cheng, Chengxiang Zhuo, Zang Li and Chenyun Yu

    User profiles, such as age and interest tags, form the backbone of modern recommender systems. However, in real-world scenarios, user profiles frequently encounter the problem of incomplete profile data, restricting the effectiveness of downstream recommendation tasks. Although large language models (LLMs) have shown remarkable potential in understanding user profiles, existing methods mainly focus on extracting explicit user information from external data and generating profile summaries. Such shallow reasoning patterns often fail to accurately infer missing user attributes under sparse data conditions. In this paper, we present GRIP (Generation and Reasoning from Incomplete Profiles), the first foundation model tailored for comprehensive profile inference under data sparsity for recommendation. GRIP introduces a unified three-stage training framework: (1) continual pre-training on large-scale structured user data for robust correlation modeling; (2) decomposed chain-of-thought reasoning with iterative self-distillation to facilitate in-depth profile inference; and (3) reinforcement learning with multi-dimensional rewards to jointly optimize factuality and reasoning coherence. Comprehensive experiments on realistic benchmarks along with online business deployments have demonstrated that GRIP significantly outperforms LLM-based methods in completing missing attributes and delivers substantial business gains in downstream recommendation tasks.

  • INDA Self-Triggered Agentic Push Recommendation System
    by Zhao-Yu Zhang, Qingying Chen, Chunyuan Zheng, Jing Zhou, Jian Sun, Siqi Chen, Leiying Chen, Chuan Zhou, Huiyou Jiang, Xin Tao, Haoxuan Li and Zhouchen Lin

    Push notification is a critical recommendation scenario on large-scale platforms, allowing the system to proactively reach users outside the application to improve long-term re-engagement. However, designing an optimal push system requires handling a complex action space for the “whether and when” delivery problem under strict system resource constraints, i.e., systems cannot naively evaluate every user at every second to find the optimal delivery moment. Existing solutions typically fall into two paradigms: the first is a two-stage approach that relies on offline user-level uplift modeling combined with integer programming solvers to allocate optimal delivery times and frequencies. The second periodically activates the push system to decide whether to send a push. However, the former cannot effectively utilize real-time information, and the latter results in resource inefficiency. In addition, such a multi-stage solution suffers from local optima. In this paper, we propose a proactive, self-triggered end-to-end agentic push recommendation system, which is already fully deployed at Douyin with over 1 billion users, allowing the system to generate push time and deternmine whether to send in a closed loop with both real-time effectiveness and efficiency. Specifically, the agentic system consists of two decision transformer-based agents: a planning agent that determines when to schedule the next push time with a gated ordinal regression method, and an action-execution agent that decides whether to send a push based on a trajectory reward. In addition, we further introduce a lightweight filtering agent to both control the computational overhead and act as a crucial safeguard against unreasonable planning behaviors. Extensive experiments demonstrate that this approach can both increase user active days by 0.3% and reduce push permission disablement by 1.9%. In addition, adding the filtering agent can reduce the computational overhead by 74%.

Back to program