Session 5:

5A: Industrial Ranking, CTR & Scalable Models

Date: Wednesday September 30, 14:00 – 15:30 CDT
Session Chair: Kim Falk

  • INDSample Is Feature: Beyond Item-Level, Toward Sample-Level Tokens for Unified Large Recommender Models
    by Shuli Wang, Junwei Yin, Changhao Li, Senjie Kou, Chi Wang, Yinqiu Huang, Yinhua Zhu, Haitao Wang and Xingxing Wang

    Scaling industrial recommender models has followed two parallel paradigms: \textbf{sample information scaling}—enriching the information content of each training sample through deeper and longer behavior sequences—and \textbf{model capacity scaling}—unifying sequence modeling and feature interaction within a single Transformer backbone. However, these two paradigms still face two structural limitations. Firstly, sample information scaling methods encode only a subset of each historical interaction into the sequence token, leaving the majority of the original sample context unexploited and precluding the modeling of sample-level, time-varying features. Secondly, model capacity scaling methods are inherently constrained by the structural heterogeneity between sequential and non-sequential features, preventing the model from fully realizing its representational capacity. To address these issues, we propose \textbf{SIF} (\emph{Sample Is Feature}), which encodes each historical Raw Sample directly into the sequence token—maximally preserving sample information while simultaneously resolving the heterogeneity between sequential and non-sequential features. SIF consists of two key components. The \textbf{Sample Tokenizer} quantizes each historical Raw Sample into a Token Sample via hierarchical group-adaptive quantization (HGAQ), enabling full sample-level context to be incorporated into the sequence efficiently. The \textbf{SIF-Mixer} then performs deep feature interaction over the homogeneous sample representations via token-level and sample-level mixing, fully unleashing the model’s representational capacity. Extensive experiments on a large-scale industrial dataset validate SIF’s effectiveness, and we have successfully deployed SIF on the Meituan food delivery platform.

  • INDAlignment + Accuracy: The Cascade Reward Representation for Preranking
    by Hedi Xia, Dylan Zhou, Yali Bian, Yichu Zhou, Zili Li, Tianyou Wang, Bella Huang, Hongbo Deng, Piyush Maheshwari, Dafang He, Darren Reger, Bowen Deng and James Li

    Prerankers in large-scale recommender systems select candidates for a downstream ranker under strict latency constraints. In practice, teams combine accuracy metrics with alignment losses to train and evaluate prerankers, but what these quantities should target—and how to combine them—remains ad hoc. We derive the \emph{Cascade Reward Representation}: under mild assumptions on a fixed-retrieval, fixed-ranker pipeline, the expected change in user reward for preranker swaps with controlled overlap shift admits a calibrated first-order two-term representation $\mathbb{E}[\mathbf{R}^E – \mathbf{R}^0] = \alpha\,\mathbb{E}[\Delta\hat O] + \beta\,\mathbb{E}[\Delta N] + \mathcal{R}$, where $\Delta\hat O$ is a logged top-fraction overlap shift (\emph{alignment}: agreement with the main ranker’s selections), $\Delta N$ is a threshold-conditioned precision shift (\emph{accuracy}: engagement above a shared ranker threshold), and the remainder is bounded by $O(\mathbb{E}[\Delta^2]) + O_P(1/\sqrt{n})$. Both proxies are computable from production logs without running the main ranker on the full retrieval pool. This representation motivates a calibrated offline metric and a matching two-branch training loss for the model family studied in our production system. We validate the representation in a large-scale industrial recommender system. A calibrated linear combination of the two metrics raises experiment winner prediction from $45$–$50\%$ (accuracy-only) to $85\%$ and Pearson $r$ from ${\le}0.65$ to $0.84$ on held-out experiments. The matching training objective delivers $+1.43\%$ save engagement over an accuracy-only baseline and $+0.62\%$ over a heuristic alignment+accuracy production model in two-week A/B tests.

  • INDAttending to the Core: Core-Task Attention for Recommendation
    by Jingyan Chen, Chenye Sun, Man Zhou, Yunhe Guo, Siyu Gu and Peng Jiang

    Industrial recommender systems typically support multiple busi- ness objectives through the integration of many specialized models. However, each individual model is usually optimized for a single primary target, such as conversions or purchases. For example, in advertising systems, models are often trained for OCPX-style objectives that tightly couple target prediction with bidding and revenue optimization. Since these target signals are often extremely sparse, multi-task learning (MTL) is widely adopted to leverage denser auxiliary tasks for additional supervision. However, exist- ing MTL approaches typically pursue balanced joint optimization across tasks, which may introduce task interference and degrade the performance of the core task. To address this limitation, we propose CoreAtt, a core-task-centric method that employs a novel attention mechanism to adaptively aggregate informative represen- tations from auxiliary tasks into the core task’s prediction tower. CoreAtt consists of two complementary attention pathways: (1) intra-sample attention, which models instance-level interactions among auxiliary tasks to produce context-aware fusion signals, and (2) inter-sample attention, which assesses each auxiliary sig- nal’s global reliability by comparing its prediction score against the population distribution. These two pathways are fused through a lightweight gating mechanism to enrich the representation for the core task. Notably, CoreAtt achieves strong performance even when built upon a simple Shared-Bottom architecture (CoreAtt-SB). We evaluate CoreAtt-SB on three public recommendation datasets, where it consistently outperforms strong MTL baselines while pre- serving auxiliary task performance. Moreover, online A/B tests on a leading short-video platform show that CoreAtt significantly enhanced platform revenue and advertiser value by 3.3% and 2.3%, respectively. Our code is available here1.

  • INDWHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture
    by Renqin Cai, Dawei Sun, Yuanjun Yao, Zhiyong Wang, Velvin Fu, Maggie Zhuang, Yu Shi, Zhongnan Fang, Xuan Cao, Jing Qian and Rui Li

    As scalability becomes increasingly important in recommendation model, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and cross features; and sequence features from user behavior histories. Wukong and HSTU have emerged as representative scalable backbones for these paths: Wukong scales high-order non-sequence feature-interaction modeling, while HSTU scales long user-behavior sequence modeling. Despite their complementary strengths, practical architectures that combine these two types of feature modeling remain underexplored. We present WHALE, a scalable unified recommendation architecture that jointly models non-sequence and sequence features on top of Wukong and HSTU. Each WHALE layer contains a Wukong module, an HSTU module, and an attention-based fusion module in which Wukong-derived interaction representations query HSTU-derived behavior representations. This design keeps both backbones active throughout the network and enables progressive Wukong–HSTU exchange, allowing high-order feature crosses to repeatedly retrieve fine-grained evidence from long user histories. To make WHALE practical for industrial deployment, we introduce customized Triton kernels and other model-systems co-design techniques to improve training and inference efficiency. On large-scale industrial recommendation data, WHALE achieves consistent gains in offline experiments. Additionally, it delivers positive online gains with a modest serving-throughput trade-off. Overall, WHALE provides a practical example of how these two sources of information can be scalably unified in an industrial recommendation model.

  • RESDo We Really Need LLMs to Augment All? A Selective Augmentation Framework with Lightweight Language Models for Multimodal CTR Prediction
    by Ziyun Chen, Yuhan Wang, Honghao Li, Mengzi Tang, Qing Xie and Yongjian Liu

    In recent years, multimodal models and large language models (LLMs) have been increasingly applied to click-through rate (CTR) prediction, owing to their ability to extract recommendation-relevant information from raw content and thereby alleviate the long-standing issue of collaborative signal sparsity. Existing paradigms typically either feed continuous embeddings from pretrained multimodal encoders into CTR models, or employ LLMs for unified knowledge augmentation. While these approaches have demonstrated promising performance, when and why their gains emerge remains unclear, and the inference cost of large LLMs introduces severe latency and scalability challenges in large-scale deployment. In this work, we empirically show that the gains of LLM-augmented CTR are highly item-dependent: uniformly applying stronger augmentation to all items is often inefficient, while the main benefits concentrate on a subset of items with more complex interaction patterns. Motivated by this observation, we propose A Selective Augmentation Framework with Lightweight Language Models for Multimodal CTR Prediction, dubbed SALM, which revisits the data augmentation paradigm and replaces indiscriminate full-data augmentation with a difficulty-aware selective strategy targeting necessary items. Extensive experiments on real-world multimodal datasets, across multiple CTR backbones and augmentation baselines, demonstrate that SALM consistently improves AUC and reduces LogLoss, while simultaneously reducing LLM-based augmentation cost.

  • INDHeterogeneous Ranking in Industrial-Scale Recommender Systems: A Case Study
    by Di Bai, Jintao Liu, Zhenwei Tang, Peifan Wu, Nada Al-Thawr and Luoshu Wang

    Heterogeneous recommendation feeds present complex challenges that extend beyond those found in highly homogeneous environments (e.g., music-only or video-only closed-ecosystem platforms). In Google Discover, a unified feed integrates diverse content sourced from the decentralized open web, including web articles, long-form and short-form videos, user-generated content (UGC), and beyond. Different content types exhibit distinct feature densities and user interaction patterns. Building a unified ranking model that sustains high performance across such heterogeneity, while avoiding negative transfer or majority bias, remains a significant industrial challenge. This paper presents an end-to-end case study on the industrial-scale multi-task ranking of heterogeneous feeds, grounded in real-world deployment. We introduce HA-MoE, a heterogeneity-adaptive multi-gated mixture-of-experts architecture that incorporates explicit heterogeneity context into both gating networks and expert representations. This approach enables effective specialization without significantly increasing operational overhead. To support reliable deployment, we introduce LENS, a lightweight observability framework that provides interpretable diagnostics of expert specialization and tracks this functional heterogeneity across continuous retraining. We evaluate our method using Dual-Level AUC (DL-AUC), a heterogeneity-aware evaluation metric that combines global ranking performance with cross-segment ranking correctness. Offline evaluations on a large-scale industrial dataset demonstrate consistent improvements over baseline models. Furthermore, online A/B testing confirms gains in feed activity and exploration metrics. Together, offline and online results validate the effectiveness of our approach for managing heterogeneity in industrial-scale recommender systems.

  • INDIDProxy: CTR Prediction with Multimodal LLMs for Cold-Start Recommendation at Xiaohongshu
    by Yubin Zhang, Haiming Xu, Guillaume Salha-Galvan, Ruiyan Han, Feiyang Xiao, Yanhua Huang, Li Lin, Luo Yang and Yao Hu

    Content-driven platforms such as Xiaohongshu often leverage clickthrough rate (CTR) prediction models for recommendation. However, these models depend heavily on item ID embeddings, which perform poorly in item cold-start settings. In this paper, we present IDProxy, a production-scale system developed at Xiaohongshu to address this challenge. IDProxy leverages multimodal large language models (MLLMs) to generate proxy embeddings from rich content signals, enabling CTR prediction for new items in the absence of usage data. Through a lightweight coarse-to-fine mechanism, these proxies are aligned with the ID embedding space and trained endto-end with the ranking model, allowing seamless integration into production-facing pipelines. Extensive offline and online experiments demonstrate the effectiveness of the method, which has been deployed in 2025 in Xiaohongshu’s Content Feed and Display Ads features, reaching hundreds of millions of users daily.

Back to program