Session 2:
2A: Long-Horizon Sequential Recommendation
Date: Tuesday September 29, 14:00 – 15:30 CDT
Session Chair: Yuyan Wang
- RESCONGA: Continual Neural Gated Architecture for Long-History Sequential Recommendation
by Hoang Vu Le, Tuong Bach Hy Nguyen and Bac LeExisting sequential recommendation models rely on absolute positional encodings and fixed context windows, producing two structural failure modes: out-of-distribution degradation on histories longer than the training window, and hard context limits that discard long-range interactions. We present CONGA (COntinual Neural Gated Architecture), which addresses these limitations through three contributions: (1) Rotary Positional Embeddings (RoPE) with a norm-preserving property (||R_m||_F = sqrt(d) for any sequence length), accelerated by custom CUDA kernels; (2) KromHC multi-stream fusion with exact doubly-stochastic mixing via Kronecker-product parametrization, where ablation confirms the expressivity gain arises from balanced gradient flow rather than additional parameters, together with a data-adaptive stream selection mechanism that prevents overfitting on sparse corpora; and (3) TITANS neural associative memory adapted to discrete-item recommendation—the first such proof-of-concept—via a two-phase training protocol with a structural forgetting-prevention property: the base encoder is frozen, preserving its short-sequence predictions, while Phase 2 only adds a learned memory term. Evaluated under a rigorous full-ranking protocol across four benchmarks (ML-1M, Beauty, Yelp, Steam), CONGA achieves state-of-the-art performance with gains up to +29.7% on ML-1M—the densest, longest-history dataset where all three failure modes are simultaneously active.
- RESResidual Dominance as a Structural Account of Last-Item Reliance in Causal Self-Attention Recommenders
by Keito Kozaki, Keigo Sakurai, Ren Togo, Takahiro Ogawa and Miki HaseyamaTransformer-based sequential recommenders with causal self-attention are known to rely heavily on the most recent interaction at inference time, but the structural origin of this behavior remains unclear. We study this problem by combining prediction-time diagnostics with norm-based analysis of the full attention block. First, we show that SASRec-style models exhibit highly localized last-item reliance. We then find that, although self-attention aggregates contextual information, the residual pathway has a dominant influence on the final representation and sharply reduces the expression of preceding context, yielding what we term residual dominance. To probe this interpretation, we use inference-time residual scaling as a controlled diagnostic intervention. Changing the residual strength induces a monotonic trade-off between structural mixing and last-item reliance, and reveals that some correct predictive signals already exist beyond the final position but are weakly expressed under standard inference. Overall, our results provide structural evidence that extreme last-item reliance in causal self-attention recommenders is closely tied to the dominance of the residual pathway at inference time. The code is available at: https://anonymous.4open.science/r/Residual-1BC6.
- RESThe Cost of Continuous Time: Diagnosing Solver Sensitivity in Recommenders and Mitigating It via Training-Free Routing
by Zixu Li and Sergio Augusto Romaña IbarraContinuous-time recommendation models irregular user behavior naturally, but under heavy-tailed temporal gaps the solver becomes part of the deployment problem. Under an aligned evaluation protocol, we find that explicit integrators (e.g., explicit Euler) exhibit substantial utility degradation on extreme temporal gaps under temporal corruption, a pattern consistent with stiffness-related numerical sensitivity. Implicit solvers mitigate this degradation through fixed-point iterations, but their iterative computational graphs substantially increase tail latency, approaching a two-fold increase in P99 under our fixed software stack. To address this trade-off, we propose Stiffness-Aware Routing (STAR), a training-free, quantile-calibrated inference-time routing policy. Rather than relying on parameterized gating networks, STAR uses offline temporal priors to route only high-risk long-gap requests to an implicit solver while keeping the remaining traffic on the fast explicit path. Experiments on Amazon, Yelp, and ML-1M show that STAR improves the Pareto frontier between robustness and latency under the aligned evaluation protocol. On Amazon, for example, STAR confines utility degradation to 2.52% (compared with 7.80% for Always-Explicit) while capping P99 latency at 18.44ms, thereby avoiding the 24.46ms tail-latency cost of an Always-Implicit baseline. On dense ML-1M, however, Always-Implicit underperforms Always-Explicit, indicating that no single solver is uniformly optimal across temporal regimes. Overall, the results show that deployment-time solver allocation is more effective than uniformly applying a single integration strategy in the ODE-based setting studied here.
- RESInformation-Aware Long Sequence Compression for Sequential Recommendation
by Wooseung Kang, Minje Kim, Suwon Lee, Gun-Woo Kim and Sang-Min ChoiSequential recommendation (SR) aims to predict a user’s next interaction by modeling temporal dependencies in historical behavior sequences. However, modeling long sequences introduces two challenges: longer histories often include noisy interactions irrelevant to a user’s core interests, and increasing sequence length substantially raises computational cost while often degrading prediction accuracy due to noise accumulation. We present RDSR, a Rate-Distortion-based Sequential Recommendation framework grounded in a task-oriented rate-utility view. Instead of directly modeling full-length sequences, RDSR combines fixed-capacity token selection with VIB-based latent compression to retain task-relevant information under explicit rate control. This suppresses irrelevant interactions while preserving essential preference signals, effectively reducing sequence length and computational overhead. Extensive experiments show that RDSR improves the performance–efficiency trade-off across attention-, MLP-, and SSM-based SR backbones, with performance gains depending on backbone inductive bias. Our code and supplementary material are available at https://anonymous.4open.science/r/Recsys_RDSR-1D6C/
- RESDP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation
by Dwipam Katariya, Thomas Caputo, Akshat Shreemali, Juan Manuel Origgi, Pranab Mohanty, Nam Nguyen, Kalanand Mishra, Nikita Seleznev and James MontgomeryTransformers have redefined sequential recommendation by effectively modeling dynamic user behaviors and long-range dependencies. However, they remain inherently inefficient: standard architectures operate at a fixed rate, allocating comparable computation to every item in a user’s history regardless of its information content. This leads to prohibitive computational overhead on long sequences and increased sensitivity to behavioral noise. To address this, practitioners often resort to lossy sequence compression, staged modeling, or truncation. This limits the model’s ability to leverage the full context of long histories during inference. Inspired by the recent success of Byte Latent Transformers, we propose DP-Rec, a dynamic latent patching architecture for recommendation. DP-Rec shifts from item-level modeling to patch-level modeling by segmenting interaction sequences using contrastive entropy surprise to identify informative behavioral boundaries. A lightweight patch encoder compresses these temporally contextualized segments into a reduced set of dynamic latent behavior vectors, which are then processed by a larger latent transformer and decoded for next-item prediction. Extensive experiments show that, under constrained computational budgets, DP-Rec scales effectively to long sequences and achieves a superior efficiency–accuracy trade-off over both non-compressed and fixed-size compression baselines.
- RESHamiltonian Spectral-Temporal Dissipative Dynamics for Sequential Recommendation
by Shuiying Liao and P. Y. MokSequential recommendation requires understanding how user preferences evolve over time, yet most existing models treat such evolution as a first order process where the next state depends solely on the current latent representation. Nevertheless, real user behavior often exhibits richer dynamics, including inertia, periodicity, and sudden shifts that cannot be fully captured by these first order assumptions. Motivated by these behavioral characteristics, we reconceptualize sequential recommendation through the lens of second order dynamical systems and introduce the Hamiltonian Spectral Recommender (HSR), a novel framework that models user interest trajectories in a dissipative Hamiltonian phase space. HSR decomposes user interest evolution into position (long term preference) and momentum (short term tendency) components, and uses a spectral symplectic integrator to propagate these dynamics efficiently in the frequency domain. A learnable dissipation mechanism further captures natural interest decay, while a short impulse refinement module models abrupt behavioral fluctuations commonly observed in sparse interaction logs. This design jointly accounts for global periodic patterns, inertial evolution, and localized shocks — three phenomena that are underrepresented in existing sequential models. Extensive experiments on three benchmark datasets demonstrate that HSR consistently outperforms state-of-the-art Transformer-based and state space model based recommenders. Ablation studies further verify the necessity of each component, including spectral propagation, second order coupling, and impulse refinement. This study highlights the value of incorporating dynamical systems perspectives into sequential recommendation, offering an effective alternative to prevailing first order modeling approaches.
- INDSequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching
by Lin Guan, Jia-Qi Yang, Zhishan Zhao, Jiaqi Huang, Hangyu Wang, Longbin Li, Beichuan Zhang, Haonan Jiang, Jinan Ni, Xiangyu Fan, Xiaowen Li, Ziyao Ren, Yuhang Qi, Xiaolong Zhu, Xuanyuan Luo, Qiwei Chen, Yi Cheng and Lele YuModern short-video recommenders must exploit ultra-long user histories—often on the order of $10^5$ interactions per user—but are constrained by strict latency and training-throughput budgets. As a result, production systems typically truncate histories or rely on two-stage retrieve-then-rank pipelines, sacrificing long-term signals and breaking end-to-end optimization. While Stacked Target-to-History Cross Attention (STCA) enables end-to-end modeling up to the 10K regime, directly scaling it to 100K remains prohibitive for serving: its dominant target-conditioned cross attention is query-dependent and its cost still grows linearly with the raw history length. We present \textbf{SequenceO1}, an end-to-end framework deployed at full traffic on Douyin that scales long-sequence ranking to the \textbf{100K} regime at billion scale. At the model level, we propose \textbf{Sketch Attention (SA)}, which compresses an ultra-long history into a fixed-size, user-only sketch using learnable prototypes and \emph{prototype-wise} normalization (each token distributes mass over prototypes). We then perform target-conditioned reasoning at two time scales: STCA over a recent 10K suffix for recency and STCA over the fixed-size sketch for ultra-long signals, followed by lightweight fusion. At the system level, the user-only sketch enables cache-first reuse: it can be computed once and reused across multiple targets and consecutive requests, removing $n$-dependent computation from the ranking critical path on cache hits. We further improve efficiency with multi-request user-level batching in training and a fused FlashSA kernel for sketching under ragged batching. Together, these model and system optimizations make end-to-end 100K sequence modeling practical in production.
2B: User Modeling, Intent & Cross-Domain Personalization
Date: Tuesday September 29, 14:00 – 15:30 CDT
Session Chair: Ludovico Boratto
- RESDPGR: Dual-Domain Spatiotemporal Generative Retrieval for Intent-Aware Local Life Service Recommendation
by Lei Shao, Fei Xiong, Meng Wang, Shiqi Tian, Zihan Yang, Ran Li, Xing Dong, Ming Liu and Hao GuLocal life service recommendation (LLSR) spans content recommendation and Point-of-Interest (POI) recommendation. On platforms such as Dianping, users browse content (notes, videos, reviews) and interact with POIs (collecting restaurants, planning check-ins) within the same session. The core challenge is that dual domain actions are causally linked: a content click on a food review and a subsequent POI collect are two reflections of the same latent user intent, not merely two separate problems. Yet existing generative retrieval methods treat content and POI signals as independent, missing this shared latent structure. To address these limitations, we propose DPGR (Dual-domain Spatiotemporal Generative Retrieval), a unified dual-domain generative retrieval framework for LLSR. DPGR introduces a Spatiotemporal State-Conditioned Token Modeling mechanism that injects dynamic user context into multiple stages of the encoder, enabling state-dependent reweighting and adaptive preference balancing. Instead of directly generating items, DPGR learns discrete intent codes via quantization of dual-domain behaviors and predicts the top intents for the target session. Each intent code independently retrieves candidates via parallel ANN, achieving diverse coverage with no additional latency. Offline experiments on public and internal datasets show that DPGR outperforms state-of-the-art generative retrieval baselines. Online A/B tests on Dianping demonstrate significant gains in both domains: in the content domain, visit views increase by 1.486% and watch time by 1.015%; in the POI domain, POI clicks increase by 0.527% and POI collects by 5.209% (all < 0.05).
- INDMend the Measurement Gap: Latent User Preference Modeling for Short-Form Video Recommendation
by Shuo Chang, Jiangguo Zhang, Yueqi Wang, Zihuan Diao, Dapeng Hong, Ali Montazer, Joyneel Misra, Tomer Margolin, Sourabh Bansod and Ningren HanRecommender systems rely heavily on heterogeneous behavioral feedback to infer user preference. Although abundant, these signals are imperfect measurements: the same observed behavior can arise from different underlying states, such as genuine enjoyment, passive consumption, or inattention. The challenge is especially acute in short-form video, where watch-based signals are strongly affected by measurement confounders such as video duration — the same watch time can imply different levels of preference for videos of different lengths, while ratio-based metrics can systematically favor short videos. As a result, optimizing raw engagement can amplify measurement artifacts rather than improving user value. We propose a Factorized Latent Value Model (FLVM) for measuring user preference from heterogeneous behavioral feedback. The model treats observed behaviors as noisy measurements of a low-dimensional, factorized latent value state and uses structured output heads to model heterogeneous feedback signals. A restricted baseline path captures predictable variation from measurement-confounding features such as video duration, user propensity, and session context, while a routed latent path estimates preference-relevant value advantage. The resulting latent value score can be integrated into an existing recommender system as a ranking feature or ranking score. On YouTube Shorts, a major short-form video platform, this model improves offline metrics and lifts a primary viewer enjoyment metric by 2.67\% in online A/B tests.
- RESDPGFlow: Decoupled Preference Guided Flow Matching for Cross Domain Sequential Recommendation
by Xiaoxin Ye, Chengkai Huang, Hongtao Huang, Shoujin Wang and Lina YaoCross-Domain Sequential Recommendation (CDSR) aims to improve next-item prediction by leveraging users’ sequential behaviors across multiple domains. Despite recent progress, existing CDSR models suffer from two fundamental limitations: (1) they often entangle transferable, domain-invariant interests with domain-specific preferences, leading to negative transfer across heterogeneous domains; and (2) they are highly sensitive to noisy interactions such as misclicks, and abrupt domain transitions. Generative models have recently emerged as a promising paradigm for modeling complex preference dynamics and mitigating noise. However, diffusion-based approaches rely on Gaussian initialization and stochastic denoising, resulting in unstable inference. Flow Matching (FM) offers a deterministic and efficient alternative by directly learning preference transport trajectories, yet existing FM-based recommenders are restricted to single-domain settings and fail to account for domain-dependent signals critical in CDSR. To bridge these gaps, we propose DPGFlow, the first Flow Matching framework specifically designed for CDSR. DPGFlow explicitly disentangles user preferences into domain-invariant and domain-specific components and injects them as structured guidance into a domain-aware conditional flow field. This design enables stable and efficient few-step inference, suppresses noise propagation, and facilitates effective knowledge transfer under heterogeneous and noisy behaviors. Extensive experiments on multiple real-world CDSR benchmarks demonstrate that DPGFlow consistently outperforms state-of-the-art baselines, while exhibiting strong robustness under noise, cold-start, and domain-transition scenarios. The data and code are available at here.
- RESCalibrating User Preferences for Cross-Domain Recommendation via Target-Guided Representation Mapping
by Guohang Zeng, Jie Lu and Guangquan ZhangAs an interdisciplinary field between transfer learning and recommender systems, cross-domain recommendation (CDR) leverages a data-rich source domain to overcome the data sparsity issue in the target domain. In this paper, we study the problem of user preference calibration in CDR when the source domain contains noisy interactions, an aspect overlooked in previous studies. We demonstrate that noisy interactions in the source domain introduce a calibration gap — a divergence between the user representations learned in the noisy source domain and the user’s true preferences — which leads to inaccurate preference transfer to the target domain. To address this, we propose a robust CDR framework called Denoising Cross-Domain Recommendation (DCDR), which incorporates a target-guided representation mapping mechanism. The intuition behind this component lies in leveraging the cleaner user representations in the target domain to construct an explicit mapping function, thereby deriving calibrated user representations that correct the distorted source-domain preferences toward their true, preference-aligned counterparts. Notably, the proposed DCDR method is agnostic to specific CDR models, making it a general framework applicable to various existing CDR approaches. Experimental results show that our method effectively calibrates user preferences and mitigates the impact of noisy preference transfer, outperforming existing single-domain denoising approaches across multiple real-world recommendation tasks.
- RESThere’s Something About You: Epistemic Recommendation for Latent Interest Discovery
by Daniel Nemirovsky, Priya Khokher, Adarsh Jois, Marco Zagha and Joaquin DelgadoHow well does a recommender system know you? These systems are typically trained on the silhouette of user activity to predict immediate engagement, such as the next click or stream. Yet this narrow focus may paradoxically expose how fragmented and incomplete the system’s knowledge of the user really is. In contrast to recommending from established user knowledge, in this work, we recommend to enrich it – an approach we term epistemic recommendation. To realize this, we propose EGRec (Epistemic Gain Recommender). For each item it could recommend, EGRec constructs two hypothetical futures (the user engages, or does not) and uses Jensen-Shannon divergence to quantify how much either outcome would enrich the model’s understanding of the user across the full catalog. We explore two variants depending on how outcomes are valued: epistemic gain (EG), which considers what the model would learn from any outcome, and expected epistemic gain (EEG), which weights learning by how likely each outcome is. For practicality, we train a lightweight prediction head on frozen model embeddings to approximate EG in a single forward pass. We formalize our intuition by introducing a dual regret framework analyzing both reward regret and coverage regret, and prove that EGRec achieves bounded reward regret O(T_explore) and bounded coverage regret O(1), while greedy relevance incurs linear coverage regret Ω(T). We evaluate EGRec both as a standalone strategy and as a signal augmenting bandit methods on KuaiRec and MovieLens-1M, against baselines spanning UCB, Thompson Sampling, MMR, and greedy relevance. Our results show that combining EG with bandit exploration (UCB+EG) achieves the best balance of recommendation accuracy and interest coverage, improving alpha-NDCG by 9-19% over the strongest exploratory baseline and calibration by up to 20%. Moreover, we show that EG signal quality depends on the expressiveness of the base model’s user representation, a dependency that serves as a practical diagnostic for when the approach will be most effective. Our work points to a broader principle in recommendation, that the value of recommending an item lies not merely in expected engagement but extends to what the system is poised to learn about the user from the outcome.
- RESBridge the Unseen Gap: Enhancing Non-overlapping Cross-domain CTR Prediction via Profile Retrieval
by Jingyang Bin, Xing Tang, Wei Zeng, Jianan Su, Kaixin Shen, Jingtong Wu, Chaohua Yang, Kailiang Hao, Dugang Liu and Xiuqiang HeClick-through rate (CTR) prediction is a fundamental task in industrial recommender systems. Cross-domain CTR prediction, which leverages data from a source domain to improve performance in a target domain, has emerged as a key strategy. However, most existing methods rely on overlapping users or items across domains to enable knowledge transfer, which fails in prevalent real-world scenarios where domains are functionally or geographically isolated (e.g., cross-country services). In this paper, we introduce a novel paradigm shift, from implicit representation alignment to explicit retrieval-based instance transfer. We propose LLM-PRIT, a framework for Large Language Model-generated Profile Retrieval & Instance Transfer. Our framework operates in three cohesive stages. First, it utilizes an LLM as a universal semantic interpreter to generate domain-agnostic, transferable profiles for users and items, encapsulating open-world knowledge. Second, instead of directly using these textual profiles, it employs them as semantic anchors to retrieve the most relevant historical instances from the source domain. This step explicitly establishes cross-domain correlations while avoiding the modality gap. Finally, it transfers knowledge by efficiently fine-tuning the target CTR model on the retrieved instances, preserving the model’s inherent feature-interaction capabilities. We conduct extensive experiments on a public and a real-world industrial dataset. Both online and offline results demonstrate the effectiveness of our LLM-PRIT, bridging the unseen gap with open-world semantic information.
RecSys 2026 (Minneapolis)
- About the Conference
- Registration
- Program at Glance
- Program
- Call for Contributions
- Challenge
- Keynotes
- Accepted Contributions
- Presenter Instructions
- Workshops
- Tutorials
- Committees
- Inclusion
- Student Volunteers
- Women in RecSys
- Visa Information
- Addressing Attendance Issues
- Location / Hotel
- Lasting Impact Award
- 60 Milestones for 20 Years




















