Session 6:
6A: Graph, Knowledge & Content Representations
Date: Wednesday September 30, 16:00 – 17:30 CDT
Session Chair: Giuseppe Spillo
- PPFEachMovie: archaeology of the first latent recommender system
by John DeTrevilleThis retrospective presents a preliminary archaeological reconstruction of EachMovie (1995–97), believed to be the first large latent recommender system; it became broadly available online a full decade before the Netflix Prize competition was announced. Running on two office PCs, EachMovie gave its users highly personalized movie recommendations by computing 20-dimensional latent vectors from its evolving 2.8M-vote dataset (97.63% sparse). EachMovie’s Joint Quartic Polak-Ribière iterative solver converged rapidly, and used a novel Hallucination-Free objective function to overcome the inherent problems of zero- or mean-filled matrices. EachMovie’s Early Voting data-augmentation strategy pre-loaded and stabilized the latent space every week by importing votes from professional reviews. EachMovie worked surprisingly well, cultivating a high degree of user engagement and trust, and noted movie critic Roger Ebert called it “uncannily accurate.”
- RESLearning Sparse Representations of Multimodal Content for Enhanced Cold Item Recommendation
by Gregor Meehan and Johan PauwelsThe scale and rapid growth of item catalogs in modern digital platforms present significant challenges to recommender system (RS) practitioners. Most RSs use embedding similarity to predict user-item preferences, but storage and low-latency retrieval of these embeddings is challenging in industry-scale catalogs. Furthermore, newly added items do not have corresponding embeddings and cannot be recommended effectively; previous works often tackle this item cold-start problem by generating cold item representations from auxiliary content, such as images or descriptive text, so that user preferences can be predicted without historical interactions. In this paper, we argue that sparse embeddings have notable advantages over standard dense vectors in this content-based cold-start paradigm. We describe how existing cold-start training regimes can be adapted for sparse representation learning, and build on insights from linear attention to design a pre-sparsification activation technique that induces sharpness and denoising effects in learned item-item similarities. We show that the resulting sparse embeddings achieve significant improvements in cold-start recommendation accuracy over dense embeddings at considerably lower storage costs, especially for users with multiple interests. Through comprehensive experiments on four multimodal RS datasets, we also demonstrate the interpretability of sparse content embeddings and their robustness in the trade-off between size and accuracy.
- RESGateBoxGCN: Hard-Soft Gated Box Embeddings with Graph Convolution for Recommendation
by Fan Mo, Takashi Wada, Rongqin Chen, Chongxian Chen, Xin Fan, Tianwei Chen and Hayato YamanaThis paper proposes GateBoxGCN, a box embedding framework that relaxes the strict positivity constraint on offsets via a hard-forward/soft-backward gating mechanism, improving recommendation performance. Box embeddings have been explored as a technique to model user preferences via high-dimensional boxes defined by centers and offsets. However, existing box-based methods restrict all offset dimensions to be positive to ensure valid box geometry, limiting the model’s flexibility and expressiveness. To address this limitation, we relax the constraint to allow offsets to be negative values. We then use negative offset dimensions to explicitly model unevaluable dimensions, such as those arising from unobserved user preferences or noise. During inference, we exclude unevaluable dimensions and calculate user-item preference scores by using only the robust ones. To handle unevaluable dimensions, we further introduce a hard-forward/soft-backward gating mechanism, where unevaluable dimensions are filtered out by the hard gate during forward propagation while the soft gate provides gradients to these dimensions during backpropagation, enabling end-to-end learning of the gating mechanism and user/item box representations. Experimental results on real-world datasets confirm the effectiveness of our method over state-of-the-art baselines.
- PPFFrom Side Information and Knowledge Graphs to Large Language Models: Two Decades of Knowledge Integration in Recommender Systems
by Claudio Pomo, Diego Baquero Sanz, Liam Claude Morris, Ludovico Boratto, Fedelucio Narducci and Tommaso Di NoiaOver the past two decades, Recommender Systems (RSs) have undergone successive transformations in how they encode and enable external knowledge: evolving from constraints and hand-crafted rules, to side information and feature matrices, to linked data and knowledge graphs, and, most recently, to Large Language Models (LLMs). This Past/Present/Future retrospective views this trajectory as a sequence of representational translations rather than a series of outright replacements. We analyze how these translations have transformed the functional capabilities of RSs, the recurring design patterns across technological eras, and the new risks that arise as knowledge integration becomes generative and conversational. By analyzing representative contributions from the RecSys literature, we identify several persistent regularities. External knowledge is repeatedly leveraged to mitigate sparsity and cold-start problems; human-interpretable structure remains fundamental for explanation, user control, and system governance; and hybrid architectures systematically reappear whenever no single representation simultaneously satisfies both scalability requirements and demands for traceability and auditability. We argue that LLMs should not be seen as replacements for structured knowledge representations, but rather as a new interaction and mediation layer that requires explicit grounding in verifiable, structured sources. We close by outlining a cautious research agenda for 2026–2030, centered on the development of hybrid, inspectable, and accountable RSs.
- RESGSPRec: On Improving Item Representations in Graph Signal Processing for Collaborative Filtering
by Ahmad Bin Rabiah and Julian McAuleyGraph-based collaborative filtering methods act as low-pass filters in the spectral domain and discard the intermediate-frequency components where community-level user preferences reside. Existing GSP-based methods address this through increasingly sophisticated filter designs, yet derive item representations from the user-item interaction matrix alone. The interaction matrix captures which items each user interacted with, but not which items users interacted with close together in their interaction ordering. We propose GSPRec, a graph spectral collaborative filtering framework that produces richer item spectral representations by incorporating item-item proximity derived from user interaction ordering before spectral filtering. GSPRec derives item-item edges from user interaction ordering via multi-hop diffusion and incorporates them into the graph topology. The resulting Laplacian exposes intermediate-frequency structure that a Gaussian bandpass filter selectively amplifies. A low-pass filter retains broad popularity trends. Extensive experiments on four real-world datasets show that GSPRec outperforms all GSP-based and GCN-based CF baselines, with average improvements of 5.12% in NDCG@10. Ablation studies establish that graph construction and filter design are coupled: incorporating item-item proximity without the bandpass filter falls below all GSP baselines, while bandpass filtering without item-item proximity already surpasses them.
- RESStabilizing Stability and Plasticity in Graph-based Continual Recommender System
by Yixin Chen, Xiangmeng Wang and Qian LiReal-world graph-based recommender systems face increasing challenges as interaction graphs evolve continuously, exposing models to persistent out-of-distribution shifts. Continual learning has emerged as a promising paradigm for incremental updates without retraining from scratch. However, existing methods primarily emphasize preserving historical knowledge (i.e., stability) and fail to address the unique challenges of graph structures. We identify fundamental challenges in graph-based continual recommendation. Through empirical analysis, we reveal three key challenges: (i) over-stabilization induced by message passing limits the absorption of new knowledge, i.e., lack of plasticity; (ii) improving plasticity for new items degrades performance on historical items, exposing a stability–plasticity trade-off; and (iii) evolving graph topology weakens the preservation of learned representations, i.e., limited stability. To address these challenges, we propose SSPRec, a graph prompt-based continual learning framework for OOD recommendation that explicitly preserve plasticity, enhance stability, and effectively balance the stability–plasticity trade-off. SSPRec freezes a pre-trained backbone and adapts to evolving graph slices via lightweight prompts and user-specific control. Specifically, we design contextual and temporal prompts to enhance stability, introduce forward-knowledge-guided contrastive objectives to improve plasticity, and develop a preference-shift-aware mechanism to adaptively balance stability and plasticity at the user level. Extensive experiments demonstrate that SSPRec consistently outperforms state-of-the-art methods under evolving graph settings.
6B: Groups, Markets & Multi-Stakeholder Evaluation
Date: Wednesday September 30, 16:00 – 17:30 CDT
Session Chair: Ladislav Peška
- RESMODE: Mutual Optimality in Direct Effects of Reciprocal Recommendations in Matching Markets
by Yoji TomitaMatching platforms such as job posting services and online dating platforms have become widely used over the past decade. For a matching platform to be successful, it is crucial to design appropriate reciprocal recommendation systems (RRSs) that consider the preferences of both sides of users (job candidates and employers) and prevent opportunities from being concentrated too heavily on a few popular users. However, prioritizing concentration mitigation too much can lead to recommending undesirable results to some individual users, resulting in their dissatisfaction. In this paper, we formulate the concept of optimality of direct effects of the recommendation list for an individual user, given the recommendations to other users. Furthermore, we propose a novel method, MODE, that computes mutually optimal recommendations in direct effects. Experiments with synthetic and real-world data demonstrate that MODE surpasses other existing methods in terms of mutual optimality of direct effects, exhibits faster processing speeds, and enables a higher expected number of matches.
- RESConsensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders
by Cedric Waterschoot, Nava Tintarev and Francesco BarilePrevious work in group recommender systems has demonstrated a sensitivity to the distribution of preferences within a group. Specifically, the selection of the preference aggregation strategy benefits from considering such group configurations. In this paper, we study whether LLMs are able to mimic this sensitivity and to select the ideal aggregation strategy and corresponding recommendation according to nuanced human perceptions of fairness, satisfaction, and consensus. We do this by fine-tuning Large Language Models (LLMs) on human survey data to serve as real-time judgmental models within the recommendation pipeline. Using a reasoning dataset distilled from DeepSeek-V3.1 and human ground truth assessments, we develop Judgmental Llama and Judgmental OLMo to simulate group assessments. Our pipeline successfully generates multiple recommendation candidates based on social choice-based aggregation strategies and dynamically selects the one that maximizes these predicted human-like evaluations. We further validate these suggestions in a user study (n=284) and find that our methodology achieved the highest scores for satisfaction and group consensus. Furthermore, we find that LLM judgments are most aligned with human perceptions of fairness, satisfaction and consensus when we also consider interaction effects between our LLM-based method and group configuration (e.g., minority or coalition). These findings give further support for dynamically adapting aggregation strategies to specific within-group preference distributions, and highlight the advantage of using LLMs for an adaptation that is aligned with subjective human judgments.
- PPFTowards the Human-Centered Study of Recommender System Providers
by Elizabeth McKinnie and Robin BurkeProviders are essential stakeholders in recommender systems, but they have been under-studied in the recommender systems field. While prior work has examined topics of fairness and diversity in ways that represent provider-side concerns, so far the field has resisted conceptualizing providers as users and applying human-centered methods to understand and improve their experience. We identify five areas for future study required to bring providers into the fold as first-class users of recommender systems: impacts, interfaces, explanation, governance, and design. We propose concrete research questions for study, referencing existing work in RecSys and adjacent fields.
- REPRAre We Really Making Progress in Group Recommendation? Unmasking the Tie-Breaking Illusion
by Song-Duo Ma and Pu-Jen ChengRecent group recommendation methods have reported strong improvements on standard benchmarks, but it remains unclear whether these gains always reflect genuine advances in modeling group preferences. In this paper, we show that several recent methods are affected by a systematic evaluation bias caused by the interaction between training-time score compression and evaluation-time deterministic tie-breaking. Specifically, an additional sigmoid transformation before the BPR objective can greatly increase tied top scores, making top-K metrics such as HR@K and NDCG@K highly sensitive to how ties are resolved. We revisit recent representative methods and their baselines onCAMRa2011 and Mafengwo under both group and user recommendation settings, and evaluate them with a tie-aware protocol that computes the exact expectation of HR@K and NDCG@K under uniform random tie-breaking. Our results show that many previously reported improvements shrink substantially under tie-aware evaluation, and the relative ranking of methods can change markedly. We further show that the additional sigmoid may act as implicit margin smoothing during optimization, and that temperature-scaled BPR can retain much of this benefit without inducing severe tie inflation. Overall, our findings highlight the importance of tie-aware evaluation for establishing reliable progress in group recommendation. The anonymized artifact package is available at https://anonymous.4open.science/r/RecSys-2026-Reproducibility-Artifact-A4CB/.
- REPRWe’ve Already Been There: A Study of Data Leakage in Ephemeral Group Recommender Systems
by Tom LupickiWe identify cross-interaction leakage, a data leakage problem in ephemeral group recommendation (EGR) evaluation that substantially inflates performance and obscures the task’s true difficulty. EGR models infer group preferences for ad hoc groups with no prior group-level history by learning from individual member interaction histories. Across three commonly used EGR benchmarks, we find that every group-item interaction is already present in every corresponding group member’s individual history, constituting up to 23.4% of all individual user-item interactions. Unlike previously studied data leakage problems in recommender evaluation, the held-out evaluation labels are directly present in the individual histories of the evaluated group’s own members. The task is effectively reduced to retrieving already-seen shared choices. We reproduce four EGR models on two benchmarks and evaluate them after removing all group-item interactions from individual user histories. Under this protocol, NDCG@20 drops by 18.3-42.9% on Weeplaces and 91.4-98.3% on Yelp across all models. This is not simply an artifact of data sparsity; randomly removing the same number of interactions per user yields drops of 4.7-18.8% and 34.5-63.7%, respectively. Further, all models are outperformed by a parameter-free heuristic under the standard leaky protocol. On a separately constructed dataset, we evaluate a user-inductive protocol to limit prior exposure to held-out group members’ individual histories, advancing toward leakage-free EGR evaluation.
RecSys 2026 (Minneapolis)
- About the Conference
- Registration
- Program at Glance
- Program
- Call for Contributions
- Challenge
- Keynotes
- Accepted Contributions
- Presenter Instructions
- Workshops
- Tutorials
- Committees
- Inclusion
- Student Volunteers
- Women in RecSys
- Visa Information
- Addressing Attendance Issues
- Location / Hotel
- Lasting Impact Award
- 60 Milestones for 20 Years




















