Demos Presentation
Date: Tuesday September 29
Time: 5:30 – 7:30 PM
- Slot ALoom: LLM-Powered Naturally Embedded Recommendation
by Piriyakorn Piriyatamwong and Satvik GargNews publishers usually present recommendations in fixed locations outside the main reading flow, where engagement tends to be low and maintaining high-quality recommendations at scale is expensive. Although research has improved what gets recommended, how recommendations are presented has changed little. This demonstration presents Loom, a modular pipeline that embeds recommendations inline within article text. Given an article and, optionally, a reader profile inferred from previous interactions, Loom first retrieves candidate articles and then uses an LLM to determine where recommendations should be inserted and how they should be phrased in context. Recommendations are rendered as clearly marked, clickable inline spans that preserve the tone and flow of the surrounding text. The system is recommendation-model agnostic through a stable retrieval API and continuously collects interaction signals, including clicks, reading time, and survey responses, creating feedback signals that can be used to improve future recommendation and personalization without additional human curation. Attendees can browse articles with live insertions, toggle personalization, and inspect the generated placements.
- Spot BScikit-Rank: Scikit-learn-Compatible Neural Ranking Models for Tabular Recommender Systems
by Aleksandr Milogradskii, Ilya Veselov, Yaroslav Klyukin, Maryia Kurdun, Oleg Lashinin and Kiryl LiakhnovichDeep \& Cross Networks (DCN) are established deep learning architectures for modeling feature interactions in CTR prediction, ranking, and recommender systems. Although they can achieve metrics comparable to gradient-boosting methods, existing implementations are often organized around framework-specific training loops rather than estimator-style interfaces, such as XGBoost, LightGBM, and CatBoost, thereby limiting their interoperability with tabular workflows for preprocessing, cross-validation, and hyperparameter optimization. To address these limitations, we present Scikit-Rank, an open-source library that exposes DCN models via a scikit-learn-compatible API and is designed to support additional neural ranking architectures. Scikit-Rank provides configurable numerical and categorical feature encoders, as well as more than 20 pointwise, pairwise, listwise, ordinal, and composite learning objectives. Furthermore, in our experiments, encoder and objective choices improved the target evaluation metrics by up to 2\% relative to reproduced DCN evaluation baselines while decreasing the number of trainable parameters without degrading predictive quality.
- Spot CPerseus: A Demo of Modular Personalization over Heterogeneous Event Sequences
by Andrei Babkin, Anna Nikiforova, Kiryl Liakhnovich, Maryia Kurdun, Oleg Lashinin and Marina AnanyevaReal-world personalization systems combine purchases, views, searches, offers, transactions, and context signals with heterogeneous schemas. We present Perseus, a Python framework for modeling heterogeneous user-event sequences through one configurable pipeline: given a reference point (user, timestamp, context, target), Perseus assembles only prior events, encodes them, applies a sequence backbone, and routes the resulting user embedding to a task-specific head. The demo covers retrieval, ranking, classification, and regression on MovieLens-1M – switching task types requires only a YAML configuration change. We validate Perseus against baselines and through experiments with richer features, additional event types, and cross-domain histories, showing that it combines heterogeneous events, declarative configuration, multi-task heads, backbone variety, offline inference, and multi-type feature encoders in one framework.
- Spot SReBait: From Clickbait to Informed Choice in Live News Recommendation Feeds
by Yi-Cheng Chang, Yi-Syuan Tsai, Xin-Ling Lan and Cheng-Te LiNews recommenders ask users to act on a headline before they see an article. A curiosity gap can therefore earn a click without earning satisfaction, contaminating the implicit feedback used for future ranking. We present \textit{ReBait}, a Chrome demonstration that adds an inspectable, ranking-agnostic intervention to five Taiwanese news feeds. It observes dynamically loaded cards, detects potential clickbait with a pinned XLM-R checkpoint, retrieves the linked article only for flagged cards, and uses Gemini 2.5 Flash to suggest a grounded, neutral headline. A yellow marker and hover comparison preserve the publisher’s title and click target while giving readers information before the click. ReBait turns clickbait mitigation into a deployable choice-time layer for user agency, media literacy, and more trustworthy recommendation feedback.
- Spot EValueVet: Auditing Recommender Systems for Fairness and Value Alignment Using Constitutional AI Principles
by Julia Kharchenko and Chirag ShahRecommender systems increasingly shape user experiences in high-stakes domains such as hiring, lending, and content curation, yet remain susceptible to biases that produce unfair outcomes for underrepresented groups. ValueVet is an interactive web-based platform that enables researchers and practitioners to systematically audit how well LLM-powered recommender systems align with fairness principles, ethical standards, and legal frameworks. Users can define custom constitutional principles—including bias and fairness criteria specific to recommendation contexts—or select from established frameworks (EU AI Act, US regulations, ISO/IEEE standards) and test recommendation outputs across multiple providers. This demonstration showcases ValueVet’s capability to audit recommendation-generating models for disparate treatment and representational harms through interactive scenarios with research-backed fairness assessments.
- Spot FLumi: An LLM-Grounded System for Safety-Aware POI Recommendations
by Maria Alon Kutsaya, Maria Naddaf, Ludovico Boratto and Adir SolomonPoint-of-interest (POI) recommender systems typically optimize for accurate next-POI prediction, while personal safety considerations remain implicit or absent. In this demonstration, we present Lumi, a safety-aware POI recommendation system that integrates city-specific contextual features into an LLM-based recommendation process. Lumi combines crime-related context, street lighting, temporal attributes, weather conditions, holiday indicators, POI category information, and user-defined visit preferences. A key feature of Lumi is a user-controlled cautiousness mechanism, allowing users to explicitly adjust the desired level of safety sensitivity for each search. The system is implemented as a mobile application connected to a WebSocket-based backend that enriches user requests with structured urban context and returns POI recommendations with human-readable safety and relevance explanations. Lumi demonstrates how LLMs can be grounded in city-scale contextual data to support explainable, safety-aware recommendations.
- Spot THow Would You Like Your Recommendations? p-book, a Self-Recommending Living Book That Teaches How Recommender Systems Work
by Pavel Kordik and Eva NecasovaA recommendation platform decides three things for its users: how content is served, how each item is told, and what exists in the catalog. We demonstrate p-book (a personalised book), an open-source living book about recommender systems that hands all three to the reader as explicit, logged, reversible choices; while a production recommender personalises the book itself, so readers watch the mechanisms they are reading about act on their own behaviour. Content is human-contracted concepts told through facet-tagged tellings; readers steer any section, are answered from the curated catalog first, and fill an honestly reported gap by generating or editing a telling that climbs a provenance-preserving trust ladder (private, community, edited, core) an elastic catalog of items on demand with curation-as-a-service. Crucially, the deployment is a research instrument: every paradigm switch, steering action, honest miss, and generated telling is logged under randomised initial assignment. Early telemetry (42 readers, 177 voluntary paradigm switches) is encouraging; we plan to collect interactions at scale across two audiences the RecSys community and the AI detem educational non-profit to study how readers approach the same content and whether LLM-generated tellings are worth curating. Try it live at https://recsys-pbook.vercel.app
- Spot UTasteprint: Cross-Platform Recommendation Agent
by Rongjie Zhu, Tianjun Wei, Cong Zhang, Guibing Guo and Zhu SunRecommender systems model users and rank candidates within individual provider boundaries, fragmenting user context across services. User agents offer a different interaction model: they can act on the user’s behalf and seek recommendations across providers, but only if user context can travel with them. We present Tasteprint, a local-first plugin framework that combines a cold-start questionnaire with authorized activity from 13 shopping, entertainment, social, and conversational services to build portable recommendation context. Tasteprint preserves source-linked platform records, organizes them into domain summaries and a cross-domain profile, and selects task-relevant context for a host agent. This allows signals observed in one domain to inform search and recommendation in another, while profile data remains in a user-owned local directory and requires neither a new model nor a cloud profile service. The demo presents user-side cross-domain recommendation, user correction, and redacted data collection. The demo is available at https://zaodushi.github.io/Tasteprint.skill.
- Spot VWiseFood: A Multi-Task Conversational Recommender for Trustworthy, Integrated, and Personalized Dietary Guidance
by Dimitrios Petrou, Nikolaos Bakatselos, Stylianos Kolidakis, Konstantinos Arvanitis, Konstantinos Andrikos, Elizaveta Kuzmenko, Dimitris Sacharidis and Dimitrios SkoutasWe present WiseFood, a knowledge-grounded conversational recommender system for personalized, healthy, and sustainable dietary decision making. Unlike existing food recommender systems that typically address isolated tasks such as recipe recommendation, meal planning, or nutrition advice, WiseFood supports multi-task recommendation through a unified knowledge base and a household-aware user model. The system provides three conversational interfaces: FoodScholar for evidence-based nutrition guidance, RecipeWrangler for personalized and explainable recipe recommendation, and FoodChat for conversational meal planning. WiseFood combines curated dietary guidelines, scientific evidence, and culinary knowledge to support knowledge-aware reranking, recipe adaptation, and evidence-grounded explanations, while continuously updating user profiles through interaction. The demonstration showcases an end-to-end recommendation journey connecting nutrition learning, recipe discovery, and adaptive meal planning within a unified platform.
- Spot JFusion AutoEncoder: Feature-Grounded Embeddings for Products with Mixed Structured Attributes
by Sreeram Ajay, Abhilash Fulkar, Atharva Kinage, Parag Jain, Akash Khetan, Sambit Sarangi, Piyush Kumar and Sanjay MohanWe present Fusion AutoEncoder, a masked autoencoding architecture for learning compact, feature-grounded embeddings from heterogeneous structured attributes. Product catalogs in industrial recommender systems combine categorical descriptors, numeric commercial signals, multi-value attributes, vector-valued context, behavioral aggregates, and missing values, making reusable representations difficult to learn and costly to refresh with large text encoders. Fusion AutoEncoder addresses this challenge through feature-specific encoders, residual fusion, missing-value-aware masking, and lightweight reconstruction heads that recover clean feature groups from corrupted inputs. The configurable architecture independently represents hotels and users and can generalize beyond them to other product domains with mixed structured attributes. As a downstream evaluation, we test the learned hotel embeddings on temporally bounded bookings using a controlled two-tower ranker with fixed user representations, candidates, and training. Results show that masked reconstruction improves neighborhood feature preservation and that Fusion AutoEncoder improves personalized ranking over structured-text embeddings while retaining a compact, deployment-efficient serving footprint. Our interactive RecSys demo makes these effects tangible through anonymized booking scenarios and side-by-side rankings, while illustrating the embeddings’ applicability to retrieval, ranking, and semantic-ID generation for generative recommendation.
- Spot KAutoRecLab: An Autonomous Recommender Systems Lab
by Moritz Baumgart, Philipp Meister, Justus Krell, Michael Schmidt, Bela Gipp and Joeran BeelRecommender systems (RecSys) research depends on extensive empirical evaluation, yet translating experimental designs into executable code remains a manual, error-prone process. This paper presents AutoRecLab, a Python-based autonomous RecSys lab that automates RecSys experimentation from natural-language prompts. Starting from a research idea, AutoRecLab derives explicit experiment requirements, develops and validates a prototype, and iteratively refines it into the requested full experiment by combining retrieval-augmented generation (RAG) documentation lookup, static type verification, and execution-steered tree search. As a demonstration, AutoRecLab autonomously implements an explicit-to-implicit feedback conversion study and, in a baseline comparison across six algorithms and three datasets, achieves 8 out of 9 successful runs at an average cost of approximately $1 per run using GPT-5.4-mini.
- Spot LFINALLY: A Dataset Recommender for RecSys Experiments
by Louis Owie, Tobias Vente and Joeran BeelFINALLY is a web-based dataset recommender for recommendersystems experiments. It supports a decision that is often made manually or by habit: which datasets should be used to evaluate a recommender algorithm? Users provide optional seed datasets, dataset-level filters, a target set size, and a recommendation strategy. FINALLY then returns a configurable dataset selection that can be inspected through dataset metadata and, where available, in an Algorithm Performance Space (APS), and exported for use in experiments and publications. The system supports APS-based diverse and non-diverse recommendation strategies, as well as random selection, with diversity as the default. FINALLY is publicly accessible, covers more than 90 recommender-system datasets, and reuses the APS Explorer code base while providing a recommendationcentered workflow. To the best of our knowledge, FINALLY is the first operational system that recommends configurable sets of datasets specifically for recommender-systems experiments
- Spot MCORGI: Communal Feed Governance for Bluesky
by Andrew Nordstrom, Anas Buhayh and Robin BurkeRecommender systems curate users’ information environments on social media platforms based on the platforms’ objectives, yet user and community input into these objectives remains limited. With the rise of federated protocols such as the AT Protocol and ActivityPub, researchers have begun investigating ways for users to meaningfully influence and design their recommendations. However, these approaches often treat recommender systems as systems designed for a single user and neglect platform objectives and, more importantly, community objectives. We present CORGI, an open-source AT Protocol feed generator that makes a social-feed ranking policy inspectable and governable by the community of users it jointly serves. This approach treats the recommender system as a community resource shaped by collective input. Our demo incorporates a set of 24 synthetic users to demonstrate the voting mechanism. The demonstration makes the full governance loop observable: a reviewer proposes a policy, synthetic voters respond, their ballots are aggregated, the fixed post set is reranked, and each post’s score can be inspected.
- Spot WSeqPack: Interactive Composable Sequence Compression for Real-time Recommender Systems
by Swapnil Shinde, Aditya Sasanur and Hongyangyang ShiReal-time recommender systems increasingly rely on expansive historical sequences to capture user behaviors, which generates massive feature sequence payloads that saturate cache memory and severely bottleneck network I/O. We present SeqPack, an interactive, model-agnostic compression framework designed to drastically reduce data movement overhead without altering existing model architectures. By utilizing a composable suite of lightweight encoders and quantizers, SeqPack can achieve over 80% payload size reduction across public and proprietary datasets, significantly outperforming standard utilities like gzip, bz2, zlib, and zstd. Crucially, its ultra-fast, vectorized decoding routines decouple network latency from sequence length, maintaining a constant network I/O profile and sustaining peak throughput. Accompanying this framework is the browser-based Compression Playground, enabling users to interactively inspect, compress, and experiment on their own data to see how it can benefit their production pipelines.
- Spot OGroupRec: A Unified Toolkit for Reproducible and Inspectable Group Recommendation Research
by Patrik Dokoupil, Ludovico Boratto and Ladislav PeskaGroup recommender systems (GRS) can be affected by methodological and procedural fragmentation. Methodologically, results-aggregation and profile-aggregation approaches are often studied with different datasets, baselines, and evaluation protocols. Procedurally, GRS experiments repeatedly reimplement dataset preparation, synthetic group generation, evaluation, and attribution in paper-specific code. This makes results difficult to reproduce, compare, and inspect beyond aggregate offline metrics. We present \textit{GroupRec}, a unified toolkit that provides common abstractions for GRS research, including dataset loading, synthetic group generation, evaluation utilities, baselines from both methodological families, and reproducibility support. We also present \textit{GroupRec Inspector}, an interactive web demo built on GroupRec that lets users generate groups under different cohesion regimes, compare recommendation strategies, adjust member importance, and inspect how individual members contribute to recommended items. Together, GroupRec and GroupRec Inspector provide a common experimental ground and an interactive front-end for understanding how methodological choices shape group recommendations.
- Spot PCOMPRESSO: Espresso-Style Sparse Representation Learning for Interpretable Recommender Systems
by Vojtěch Vančura, Giacomo Medda, Martin Spišák and Ladislav PeškaSparse representations can make recommender-system embeddings more compact and inspectable, but developing sparse-learning workflows typically requires substantial engineering around sparsification, training, pruning, storage, and analysis. We present COMPRESSO, an open-source PyTorch framework that exposes this functionality through a simple and modular interface. Inspired by Italian espresso culture, where one orders a caffè while the barista handles the machinery, Compresso lets researchers focus on sparse models rather than infrastructure. The framework provides high-level sparse autoencoder training, reusable sparse tensor representations, differentiable top-k operators, sparse and masked neural parameters, pruning schedules, and composable clustering tools. Our demonstration presents an end-to-end recommender-system workflow in which dense item embeddings are transformed into sparse codes and organized into interpretable product clusters across several domains. Users can inspect sparse factors, representative items, cluster labels, and the source code that generated them. Compresso lowers the barrier to sparse representation learning for compression, interpretation, and the development of sparse recommender models.
- Spot QStreamlitRecommenders: Towards Recommendation Inspectability as a New Reproducibility Standard
by Václav Stibor, Vojtěch Vančura and Ladislav PeškaReproducibility remains a persistent challenge in recommender systems (RS) research. While sharing datasets, source code, or evaluation details is becoming standard, the support for model inspection and understanding often remains limited. As such, while researchers receive artifacts to reproduce aggregate metrics, these do not provide much insight into how the model behaves for individual users, its failure modes, or how its parameters influence recommendations. We argue that these issues may be alleviated through a \textit{visual inspection} of model outputs, yet building such a system from scratch may impose a prohibitive burden on model authors. Therefore, we present \textsc{StreamlitRecommenders}, a lightweight Python library for rapidly building interactive RS demonstrations directly from research code. The framework acts as a thin presentation layer atop existing recommender implementations, enabling researchers to seamlessly create fully interactive web demonstrations with minimal additional code and without expertise in web development or Streamlit internals. By lowering the barrier to publishing inspectable recommendation artifacts alongside research papers, we aim to promote interactive model inspection as a standard component of reproducible recommender systems research.
- Spot RConversational Recommendation over Live E-Commerce Catalogues with Self-Refreshing Retrieval
by Ante Kapetanovic, Tomislav Duricic, Dionizije Fa, Andro Mercep and Emanuel LacicConversational recommender systems based on large language models are often evaluated on static and pre-indexed item collections. E-commerce catalogues, however, change continuously as products are added or removed, prices are updated, and stock availability fluctuates. As such, in this paper we present a multi-turn conversational shopping assistant that is designed to operate over live product catalogues. Our solution includes a self-refreshing retrieval pipeline that ingests a merchant product feed, parses and enriches product records, and synchronizes them with a vector index. In order for a refreshed product catalogue to process only the item delta rather than rebuild the entire index, we use per-item content hashes to identify new, modified, deleted, and unchanged products. A controller-based dialogue architecture applies an LLM for intent classification, preference elicitation, and response generation, while product retrieval, reranking, and diversity selection are handled by dedicated functions. Our demonstration first shows how product additions, price changes, stock updates, and removals are propagated to the index. It then presents a shopping conversation in which the refreshed catalogue is immediately reflected in the recommendations. Access to the WhatsApp chatbot, code documentation, reproduction instructions, and a recorded walkthrough are available at [https://github.com/infobip/infobip-agentic-crs].
RecSys 2026 (Minneapolis)
- About the Conference
- Registration
- Program at Glance
- Program
- Call for Contributions
- Challenge
- Keynotes
- Accepted Contributions
- Presenter Instructions
- Workshops
- Tutorials
- Committees
- Inclusion
- Student Volunteers
- Women in RecSys
- Visa Information
- Addressing Attendance Issues
- Location / Hotel
- Lasting Impact Award
- 60 Milestones for 20 Years




















