Session 2:

2A: Long-Horizon Sequential Recommendation

Date: Tuesday September 29, 14:00 – 15:30 CDT
Session Chair: Yuyan Wang

  • RESCONGA: Continual Neural Gated Architecture for Long-History Sequential Recommendation
    by Hoang Vu Le, Tuong Bach Hy Nguyen and Bac Le

    Existing sequential recommendation models rely on absolute positional encodings and fixed context windows, producing two structural failure modes: out-of-distribution degradation on histories longer than the training window, and hard context limits that discard long-range interactions. We present CONGA (COntinual Neural Gated Architecture), which addresses these limitations through three contributions: (1) Rotary Positional Embeddings (RoPE) with a norm-preserving property (||R_m||_F = sqrt(d) for any sequence length), accelerated by custom CUDA kernels; (2) KromHC multi-stream fusion with exact doubly-stochastic mixing via Kronecker-product parametrization, where ablation confirms the expressivity gain arises from balanced gradient flow rather than additional parameters, together with a data-adaptive stream selection mechanism that prevents overfitting on sparse corpora; and (3) TITANS neural associative memory adapted to discrete-item recommendation—the first such proof-of-concept—via a two-phase training protocol with a structural forgetting-prevention property: the base encoder is frozen, preserving its short-sequence predictions, while Phase 2 only adds a learned memory term. Evaluated under a rigorous full-ranking protocol across four benchmarks (ML-1M, Beauty, Yelp, Steam), CONGA achieves state-of-the-art performance with gains up to +29.7% on ML-1M—the densest, longest-history dataset where all three failure modes are simultaneously active.

  • RESResidual Dominance as a Structural Account of Last-Item Reliance in Causal Self-Attention Recommenders
    by Keito Kozaki, Keigo Sakurai, Ren Togo, Takahiro Ogawa and Miki Haseyama

    Transformer-based sequential recommenders with causal self-attention are known to rely heavily on the most recent interaction at inference time, but the structural origin of this behavior remains unclear. We study this problem by combining prediction-time diagnostics with norm-based analysis of the full attention block. First, we show that SASRec-style models exhibit highly localized last-item reliance. We then find that, although self-attention aggregates contextual information, the residual pathway has a dominant influence on the final representation and sharply reduces the expression of preceding context, yielding what we term residual dominance. To probe this interpretation, we use inference-time residual scaling as a controlled diagnostic intervention. Changing the residual strength induces a monotonic trade-off between structural mixing and last-item reliance, and reveals that some correct predictive signals already exist beyond the final position but are weakly expressed under standard inference. Overall, our results provide structural evidence that extreme last-item reliance in causal self-attention recommenders is closely tied to the dominance of the residual pathway at inference time. The code is available at: https://anonymous.4open.science/r/Residual-1BC6.

  • RESThe Cost of Continuous Time: Diagnosing Solver Sensitivity in Recommenders and Mitigating It via Training-Free Routing
    by Zixu Li and Sergio Augusto Romaña Ibarra

    Continuous-time recommendation models irregular user behavior naturally, but under heavy-tailed temporal gaps the solver becomes part of the deployment problem. Under an aligned evaluation protocol, we find that explicit integrators (e.g., explicit Euler) exhibit substantial utility degradation on extreme temporal gaps under temporal corruption, a pattern consistent with stiffness-related numerical sensitivity. Implicit solvers mitigate this degradation through fixed-point iterations, but their iterative computational graphs substantially increase tail latency, approaching a two-fold increase in P99 under our fixed software stack. To address this trade-off, we propose Stiffness-Aware Routing (STAR), a training-free, quantile-calibrated inference-time routing policy. Rather than relying on parameterized gating networks, STAR uses offline temporal priors to route only high-risk long-gap requests to an implicit solver while keeping the remaining traffic on the fast explicit path. Experiments on Amazon, Yelp, and ML-1M show that STAR improves the Pareto frontier between robustness and latency under the aligned evaluation protocol. On Amazon, for example, STAR confines utility degradation to 2.52% (compared with 7.80% for Always-Explicit) while capping P99 latency at 18.44ms, thereby avoiding the 24.46ms tail-latency cost of an Always-Implicit baseline. On dense ML-1M, however, Always-Implicit underperforms Always-Explicit, indicating that no single solver is uniformly optimal across temporal regimes. Overall, the results show that deployment-time solver allocation is more effective than uniformly applying a single integration strategy in the ODE-based setting studied here.

  • RESInformation-Aware Long Sequence Compression for Sequential Recommendation
    by Wooseung Kang, Minje Kim, Suwon Lee, Gun-Woo Kim and Sang-Min Choi

    Sequential recommendation (SR) aims to predict a user’s next interaction by modeling temporal dependencies in historical behavior sequences. However, modeling long sequences introduces two challenges: longer histories often include noisy interactions irrelevant to a user’s core interests, and increasing sequence length substantially raises computational cost while often degrading prediction accuracy due to noise accumulation. We present RDSR, a Rate-Distortion-based Sequential Recommendation framework grounded in a task-oriented rate-utility view. Instead of directly modeling full-length sequences, RDSR combines fixed-capacity token selection with VIB-based latent compression to retain task-relevant information under explicit rate control. This suppresses irrelevant interactions while preserving essential preference signals, effectively reducing sequence length and computational overhead. Extensive experiments show that RDSR improves the performance–efficiency trade-off across attention-, MLP-, and SSM-based SR backbones, with performance gains depending on backbone inductive bias. Our code and supplementary material are available at https://anonymous.4open.science/r/Recsys_RDSR-1D6C/

  • RESDP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation
    by Dwipam Katariya, Thomas Caputo, Akshat Shreemali, Juan Manuel Origgi, Pranab Mohanty, Nam Nguyen, Kalanand Mishra, Nikita Seleznev and James Montgomery

    Transformers have redefined sequential recommendation by effectively modeling dynamic user behaviors and long-range dependencies. However, they remain inherently inefficient: standard architectures operate at a fixed rate, allocating comparable computation to every item in a user’s history regardless of its information content. This leads to prohibitive computational overhead on long sequences and increased sensitivity to behavioral noise. To address this, practitioners often resort to lossy sequence compression, staged modeling, or truncation. This limits the model’s ability to leverage the full context of long histories during inference. Inspired by the recent success of Byte Latent Transformers, we propose DP-Rec, a dynamic latent patching architecture for recommendation. DP-Rec shifts from item-level modeling to patch-level modeling by segmenting interaction sequences using contrastive entropy surprise to identify informative behavioral boundaries. A lightweight patch encoder compresses these temporally contextualized segments into a reduced set of dynamic latent behavior vectors, which are then processed by a larger latent transformer and decoded for next-item prediction. Extensive experiments show that, under constrained computational budgets, DP-Rec scales effectively to long sequences and achieves a superior efficiency–accuracy trade-off over both non-compressed and fixed-size compression baselines.

  • RESHamiltonian Spectral-Temporal Dissipative Dynamics for Sequential Recommendation
    by Shuiying Liao and P. Y. Mok

    Sequential recommendation requires understanding how user preferences evolve over time, yet most existing models treat such evolution as a first order process where the next state depends solely on the current latent representation. Nevertheless, real user behavior often exhibits richer dynamics, including inertia, periodicity, and sudden shifts that cannot be fully captured by these first order assumptions. Motivated by these behavioral characteristics, we reconceptualize sequential recommendation through the lens of second order dynamical systems and introduce the Hamiltonian Spectral Recommender (HSR), a novel framework that models user interest trajectories in a dissipative Hamiltonian phase space. HSR decomposes user interest evolution into position (long term preference) and momentum (short term tendency) components, and uses a spectral symplectic integrator to propagate these dynamics efficiently in the frequency domain. A learnable dissipation mechanism further captures natural interest decay, while a short impulse refinement module models abrupt behavioral fluctuations commonly observed in sparse interaction logs. This design jointly accounts for global periodic patterns, inertial evolution, and localized shocks — three phenomena that are underrepresented in existing sequential models. Extensive experiments on three benchmark datasets demonstrate that HSR consistently outperforms state-of-the-art Transformer-based and state space model based recommenders. Ablation studies further verify the necessity of each component, including spectral propagation, second order coupling, and impulse refinement. This study highlights the value of incorporating dynamical systems perspectives into sequential recommendation, offering an effective alternative to prevailing first order modeling approaches.

  • INDSequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching
    by Lin Guan, Jia-Qi Yang, Zhishan Zhao, Jiaqi Huang, Hangyu Wang, Longbin Li, Beichuan Zhang, Haonan Jiang, Jinan Ni, Xiangyu Fan, Xiaowen Li, Ziyao Ren, Yuhang Qi, Xiaolong Zhu, Xuanyuan Luo, Qiwei Chen, Yi Cheng and Lele Yu

    Modern short-video recommenders must exploit ultra-long user histories—often on the order of $10^5$ interactions per user—but are constrained by strict latency and training-throughput budgets. As a result, production systems typically truncate histories or rely on two-stage retrieve-then-rank pipelines, sacrificing long-term signals and breaking end-to-end optimization. While Stacked Target-to-History Cross Attention (STCA) enables end-to-end modeling up to the 10K regime, directly scaling it to 100K remains prohibitive for serving: its dominant target-conditioned cross attention is query-dependent and its cost still grows linearly with the raw history length. We present \textbf{SequenceO1}, an end-to-end framework deployed at full traffic on Douyin that scales long-sequence ranking to the \textbf{100K} regime at billion scale. At the model level, we propose \textbf{Sketch Attention (SA)}, which compresses an ultra-long history into a fixed-size, user-only sketch using learnable prototypes and \emph{prototype-wise} normalization (each token distributes mass over prototypes). We then perform target-conditioned reasoning at two time scales: STCA over a recent 10K suffix for recency and STCA over the fixed-size sketch for ultra-long signals, followed by lightweight fusion. At the system level, the user-only sketch enables cache-first reuse: it can be computed once and reused across multiple targets and consecutive requests, removing $n$-dependent computation from the ranking critical path on cache hits. We further improve efficiency with multi-request user-level batching in training and a fused FlashSA kernel for sketching under ragged batching. Together, these model and system optimizations make end-to-end 100K sequence modeling practical in production.

Back to program