Pith. sign in

REVIEW 4 major objections 7 minor 42 references

A new ranking architecture, TmallGS, claims that adapting LLM-style Transformers to e-commerce search requires decoupling: keep explicit matching signals out of the attention stream and inject them late via FiLM modulation. The paper report

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 05:16 UTC pith:YZJX4QXO

load-bearing objection A credible, well-engineered industrial scaling architecture with plausible gains, but the Bias Net decoupling proof has a notation bug, and all evidence is proprietary. the 4 major comments →

arxiv 2607.13398 v1 pith:YZJX4QXO submitted 2026-07-15 cs.IR

TMallGS: Scaling Unified Feature and Sequence Modeling for Generative E-commerce Search

classification cs.IR
keywords CTR predictionTransformer rankingdecoupled fusionFiLMfeature heterogeneityscaling lawse-commerce searchindustrial recommender
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper is trying to establish that a Transformer-based ranking model for e-commerce search can outperform both classic DLRMs and previous 'tokenize-everything' Transformer adaptations, provided the architecture explicitly separates heterogeneous feature types. The central claim is that deep self-attention acts as a low-pass filter that dilutes high-frequency explicit matching signals, so the model adds a Decoupled FiLM Late Fusion path that re-injects these hard signals at the output. Offline experiments on a 500-million-sample industrial dataset show the largest gains over baselines, and an online A/B test reports substantial commercial improvements. The authors argue this proves that domain-specific decoupling is critical when scaling LLM-style architectures for search ranking.

Core claim

On the paper's own terms, TmallGS demonstrates that unified Transformer ranking models suffer from signal dilution when explicit cross-features (like query-item match scores) are fed into the attention stack. Its Decoupled FiLM Late Fusion, which modulates the final semantic representation with a dedicated embedding of those heavy cross-features, recovers the lost high-frequency signal and is the single largest contributor to GAUC improvement (-0.32% drop when removed). Across the full offline comparison, TmallGS reaches +1.12% AUC and +1.26% GAUC over the production baseline, and in a 30-day online A/B test it achieves +0.79% PV-AUC, +0.34% Imp-GAUC, +1.38% UCTCVR, and +1.52% GMV with only

What carries the argument

Decoupled FiLM Late Fusion: after the deep Transformer backbone produces a semantic candidate vector, a Feature-wise Linear Modulation layer (FiLM) takes the candidate's comprehensive embedding (including heavy cross-features) and generates affine scaling and shifting coefficients to modulate the unified representation. This creates a separate pathway that preserves high-frequency matching signals without forcing them through the low-pass attention filter. The architecture also uses per-field QKV projections to align heterogeneous feature manifolds, a noise-adaptive gating mechanism to suppress attention noise in long behavior sequences, a context-aware bias net whose bias term cancels in pa

Load-bearing premise

All performance evidence comes from a single proprietary industrial dataset (Tmall App search logs) with no public benchmark, no released code, and no confidence intervals for the offline deltas, so the central claim that decoupled FiLM fusion causes the observed gains is only as strong as that one dataset's representativeness.

What would settle it

Run a head-to-head comparison of TmallGS against OneTrans and other baselines on an independent e-commerce search dataset with long behavior sequences and explicit query-item match features; if removing the FiLM fusion path does not reproduce the reported GAUC drop (or if the overall gains vanish), the decoupling mechanism is not the causal driver claimed.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the decoupling claim holds, future Transformer-based ranking systems should keep explicit matching signals out of the deep attention stack and inject them via a late multiplicative modulation, rather than tokenizing all features uniformly.
  • The reported scaling behavior along width, depth, and sequence length suggests that compute-intensive ranking models can achieve predictable gains beyond the DLRM scaling ceiling, making further investment in Transformer backbones for search economically rational.
  • The context-aware bias net's pairwise-loss cancellation implies a clean separation between global calibration and intra-session ranking, a principle that could generalize to other ranking tasks with request-level biases.
  • The two-stage warm-up result indicates that migrating from memory-bound DLRMs to dense Transformer backbones can be practical without discarding legacy embeddings, lowering the barrier for industrial adoption.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: The reported gains likely depend on the presence of strong explicit cross-features and ultra-long behavior sequences; on datasets without heavy matching signals, the FiLM fusion path may contribute much less, and the architecture's advantage could shrink.
  • Editorial inference: The online lift of +1.52% GMV combined with only +6ms latency suggests that the proposed shallow-but-wide configuration (4 layers, 1024 dimensions) may represent a sweet spot; deeper configurations, while yielding better offline scaling, might not be latency-feasible in production.
  • Editorial inference: A direct test of the paper's mechanism would be to run the same ablation on a public ranking benchmark with explicit user-item features and observe whether FiLM fusion's removal consistently causes the sharpest GAUC drop; if the effect is dataset-specific, the claim of general 'signal preservation' would need to be qualified.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper proposes TmallGS, a Transformer-based ranking architecture for Tmall search, comprising five components: hierarchical distribution-calibrated tokenization (FSR + DCP), a field-adaptive gated Transformer backbone with per-field QKV projections and noise-adaptive gating, decoupled FiLM late fusion, a context-aware bias net, and error-aware progressive training. The authors report offline gains of +1.12% AUC and +1.26% GAUC over a DIN baseline, online A/B lifts of +1.38% UCTCVR and +1.52% GMV, and claim monotonic scaling laws with respect to width, depth, and sequence length. All experiments are conducted on a single proprietary 31-day Tmall dataset. The paper also proposes a two-stage warm-up strategy for deploying Transformer backbones.

Significance. If the reported results hold, TmallGS offers a concrete industrial recipe for scaling Transformer backbones in e-commerce search, addressing feature heterogeneity and high-frequency signal dilution through architectural decoupling. The paper is strong in its detailed architecture description, a full five-component ablation, and a large-scale online A/B test. The key statistical and theoretical underpinnings, however, are not fully supported: the bias-cancellation claim in the pairwise loss is internally inconsistent, and the scaling-law evidence is absent from the manuscript as referenced (Figure 2 is not included). Reproducibility is limited by the exclusive use of proprietary data and no released code. The significance is therefore conditional on correcting these issues and providing stronger statistical evidence.

major comments (4)
  1. [Section 4.6, Eq. (16) vs Eq. (19)] The claimed cancellation of the bias term in the pairwise loss is not valid as written. Eq. (15) defines ŷ = σ(logit_main + logit_bias), while Eq. (19) computes log σ(ŷ(c+) − ŷ(c−)). Because ŷ are probabilities, the difference ŷ(c+) − ŷ(c−) depends on the shared bias term through the sigmoid; the bias does not cancel as claimed in Eq. (16). The cancellation would only hold if the pairwise loss operated on pre-sigmoid logits. This internal inconsistency invalidates the 'orthogonal decoupling' justification for the Bias Net, one of the paper's five core contributions. Please either redefine Eq. (19) in terms of logits or revise the theoretical claim accordingly.
  2. [Section 5.4, Figure 2] Figure 2 is referenced in the text but not present in the manuscript. The scaling-law claim rests entirely on this figure and on only three configurations per dimension (width, depth, length), with no error bars, no fitted scaling curve, and no statistical characterization. This does not support the strong statement that TmallGS 'adheres to scaling laws' or that gains are 'predictable'. The figure and quantitative analysis must be provided, or the claim must be substantially softened.
  3. [Section 5.1.3 and Table 2] Offline results are reported as single-point relative deltas with no confidence intervals, standard errors, or number of repeated runs. The statement in Section 5.1.3 that 'an absolute increase of 0.001 in GAUC is considered statistically significant' is an assertion without supporting test. Given the proprietary dataset and no code release, readers cannot assess whether the +1.12% AUC and +1.26% GAUC gains are within run-to-run noise. Please report repeated-run statistics or, at minimum, the online A/B bucket counts and confidence intervals for the offline deltas.
  4. [Section 5.1.4 and Table 2] Comparison fairness is asserted but not demonstrated. The implementation details state that all Transformer baselines share the same hidden dimension, but this does not equalize parameter counts, FLOPs, depth, training steps, or hyperparameter tuning budgets. Without a clear description of capacity matching and the search protocol for each baseline, the claim that TmallGS outperforms OneTrans and HSTU is not fully supported. Please specify the total parameter/FLOP budgets and the tuning effort per baseline.
minor comments (7)
  1. [Section 4.3/4.6] The 'Context Bias Anchor' in Eq. (5) is called e_bias, but Section 4.6 refers to a 'Context Anchor token'. Please unify the terminology to avoid confusion.
  2. [Eq. (11)] The concatenation operator ⊕ is used in Eq. (11) but not defined. Please state that it denotes vector concatenation.
  3. [Section 5.1.2] The DIN baseline is described as 'combined with RankMixer as our production baseline', yet Table 2 lists DIN and RankMixer separately. Clarify whether DIN in Table 2 is the pure DIN or a DIN+RankMixer hybrid.
  4. [Section 5.2] The text says 'our production baseline is DNN with Rankmixer', but the relative improvements in Table 2 are computed against DIN (Base). It would be helpful to also report gains relative to the actual production baseline (RankMixer).
  5. [Section 5.5.2] The text says 'we prioritize sequence length over model depth', but Section 5.4 reports that depth scaling yields the steepest ROI. This apparent contradiction should be reconciled, perhaps by explaining latency constraints.
  6. [Algorithm 1, line 10] The sequence construction line contains a stray trailing semicolon and inconsistent subscripts (H_ui_h vs H_uih). Please correct these typographical issues.
  7. [Title/Abstract] The paper describes the model as 'generative e-commerce search', but TmallGS is a discriminative ranking model, not a generative model. This terminology may mislead readers; consider rewording to 'generative-inspired' or 'Transformer-based'.

Circularity Check

0 steps flagged

No significant circularity: TmallGS's reported gains are empirical measurements, not predictions derived from fitted inputs; no load-bearing self-citation chain exists.

full rationale

The paper's central claims rest on direct offline/online measurement: TmallGS is trained and evaluated on Tmall logs, compared against baselines in Table 2, ablated in Table 3, and A/B tested in Table 4. No reported gain is a parameter fitted to a subset and then re-reported as a prediction of a closely related quantity. The 'scaling law' (Section 5.4, Figure 2) is an empirical trend over three configurations of the model's own measured performance, not a law fitted to an external distribution and then used to predict the same measurements; it is an analogy to LLM scaling, not a derivation from it. I find no author-overlapping self-citations in the reference list: none of the cited prior works share authors with the present paper, and the only acknowledgment (TorchEasyRec) is engineering infrastructure, not evidence. The Context-Aware Bias Net argument in Section 4.6 contains a real internal inconsistency — Eq. 16 cancels a bias term that does not cancel once Eq. 15's sigmoid is substituted into Eq. 19's pairwise loss — but that is a mathematical/notation defect in the theoretical justification, not a circular equivalence: the component's value is established by ablation (Table 3, w/o Bias Net), independent of the cancellation claim. The abstract also names 'Climber' with no bibliography entry, a citation-hygiene issue rather than a circularity. The use of a single proprietary dataset and the three-point scaling extrapolation are external-validity/overclaiming concerns, not circularity under the definitions of this analysis. Therefore no circular step can be exhibited, and the appropriate score is 0.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 1 invented entities

The central empirical claim depends on two unverifiable inputs: a proprietary dataset and the choice of numerous architectural hyperparameters. No physical entities are introduced; the only 'new' entity is a model-internal context token without independent evidence.

free parameters (5)
  • λ (error-aware progressive weighting) = not specified
    Hyperparameter controlling layer-wise loss weighting (Eq. 18); tuned on validation.
  • γ (pairwise loss weight) = not specified
    Balance between BCE and pairwise objectives (Eq. 20); tuned on validation.
  • Model hyperparameters (layers, dim, heads) = L=4, D=1024, H=8
    Chosen from scaling sweep; not derived.
  • Sequence length truncation = 1500
    Truncation length for lifetime history; average sequence length is 1500.
  • FSR bottleneck dimensions = not specified
    Architecture choices for the saliency reweighting network (Eq. 2).
axioms (5)
  • domain assumption Deep self-attention behaves as a low-pass filter that dilutes high-frequency matching signals
    Invoked in Sections 1 and 4.5 to justify FiLM late fusion; if false, the design rationale for decoupling weakens.
  • domain assumption The proprietary Tmall dataset is representative of industrial e-commerce search and offline gains transfer to online
    All experiments are on this dataset (Section 5.1.1); no public benchmark is used.
  • domain assumption Model FLOPs Utilization (MFU) is the key bottleneck limiting scaling of recommendation models
    Motivates compute-intensive Transformer backbones (Sections 1 and 2).
  • ad hoc to paper Monotonic gains across three configurations imply a scaling law
    Figure 2: only 3 points per dimension, no fitted power law, no error bars, curve not shown in text.
  • standard math Standard backpropagation and Transformer attention are correct
    Background assumptions used throughout.
invented entities (1)
  • Context Anchor token no independent evidence
    purpose: Summarizes request-level context (pagination bias, query context, user state) for the bias net and global attention
    A learned token with no external falsifiable handle; its utility is only evaluated within the model.

pith-pipeline@v1.3.0-alltime-deepseek · 14434 in / 11292 out tokens · 98780 ms · 2026-08-02T05:16:36.927348+00:00 · methodology

0 comments
read the original abstract

In industrial search and ranking systems, Click-Through Rate (CTR) prediction is shifting from traditional Deep Learning Recommendation Models (DLRM) toward unified, compute-intensive Transformer architectures. This transition is driven by the need to improve Model FLOPs Utilization (MFU) and achieve predictable gains through scaling laws. However, existing approaches such as OneTrans and Climber often adopt an all-in-tokenization strategy when adapting Large Language Model (LLM) architectures, overlooking the heterogeneous nature of ranking features. We propose TmallGS, a scalable ranking architecture for Tmall search. TmallGS includes five key components: (1) Hierarchical Distribution-Calibrated Tokenization, which combines Field-wise Saliency Reweighting (FSR) and Distribution-Calibrated Projection (DCP) to map diverse features into optimized subspaces; (2) a Field-Adaptive Gated Transformer Backbone with per-field QKV projections and noise-adaptive gating for refined semantic interaction; (3) Decoupled FiLM Late Fusion to preserve explicit high-frequency signals; (4) a Context-Aware Bias Net to decouple systemic bias from user intent; and (5) Error-Aware Progressive Training with dynamically weighted losses for robust learning. Extensive offline experiments and online A/B tests on Tmall Search show that TmallGS improves training throughput and achieves substantial gains in UCTCVR and GMV.

Figures

Figures reproduced from arXiv: 2607.13398 by Bokang Wang, Guangxin Song, He Guo, Jing Wang, Xing Fang, Yipin Dai, Yufeng Gao, Zhentao Song.

Figure 1
Figure 1. Figure 1: The overall architecture of TmallGS. (Left) The distinct processing flow comprising three stages: Section A uses FSR [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Scaling Laws of TmallGS over DIN baseline. Perfor [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

42 extracted references · 9 linked inside Pith

  1. [1]

    Weijie Bian, Kailun Wu, Lejian Ren, Qi Pi, Yujing Zhang, Can Xiao, Xiang-Rong Sheng, Yong-Nan Zhu, Zhangming Chan, Na Mou, et al . 2022. CAN: feature co-action network for click-through rate prediction. InProceedings of the fifteenth ACM international conference on web search and data mining. 57–65

  2. [2]

    Ben Chen, Xian Guo, Siyuan Wang, Zihan Liang, Yue Lv, Yufei Ma, Xinlong Xiao, Bowen Xue, Xuxin Zhang, Ying Yang, et al . 2025. Onesearch: A preliminary exploration of the unified end-to-end generative framework for e-commerce search.arXiv preprint arXiv:2509.03236(2025)

  3. [3]

    Ben Chen, Siyuan Wang, Yufei Ma, Zihan Liang, Xuxin Zhang, Yue Lv, Ying Yang, Huangyu Dai, Lingtao Mao, Tong Zhao, et al. 2026. OneSearch-V2: The Latent Reasoning Enhanced Self-distillation Generative Search Framework.arXiv preprint arXiv:2603.24422(2026)

  4. [4]

    Qiwei Chen, Huan Zhao, Wei Li, Pipei Huang, and Wenwu Ou. 2019. Behavior sequence transformer for e-commerce recommendation in alibaba. InProceedings of the 1st international workshop on deep learning practice for high-dimensional sparse data. 1–4

  5. [5]

    Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al

  6. [6]

    Sunhao Dai, Jiakai Tang, Jiahua Wu, Kun Wang, Yuxuan Zhu, Bingjun Chen, Bangyang Hong, Yu Zhao, Cong Fu, Kangle Wu, et al. 2025. Onepiece: Bringing context engineering and reasoning to industrial cascade ranking system.arXiv preprint arXiv:2509.18091(2025)

  7. [7]

    Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. Flashat- tention: Fast and memory-efficient exact attention with io-awareness.Advances in neural information processing systems35 (2022), 16344–16359

  8. [8]

    Jiaxin Deng, Shiyao Wang, Kuo Cai, Lejian Ren, Qigen Hu, Weifeng Ding, Qiang Luo, and Guorui Zhou. 2025. Onerec: Unifying retrieve and rank with generative recommender and iterative preference alignment.arXiv preprint arXiv:2502.18965 (2025)

  9. [9]

    Yufei Feng, Fuyu Lv, Weichen Shen, Menghan Wang, Fei Sun, Yu Zhu, and Keping Yang. 2019. Deep session interest network for click-through rate prediction. arXiv preprint arXiv:1905.06482(2019)

  10. [10]

    Huan Gui, Ruoxi Wang, Ke Yin, Long Jin, Maciej Kula, Taibai Xu, Lichan Hong, and Ed H Chi. 2023. Hiformer: Heterogeneous feature interactions learning with transformers for recommender systems.arXiv preprint arXiv:2311.05884(2023)

  11. [11]

    Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: a factorization-machine based neural network for CTR prediction.arXiv preprint arXiv:1703.04247(2017)

  12. [12]

    Ruidong Han, Bin Yin, Shangyu Chen, He Jiang, Fei Jiang, Xiang Li, Chi Ma, Mincong Huang, Xiaoguang Li, Chunzhen Jing, et al . 2025. Mtgr: Industrial- scale generative recommendation framework in meituan. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 5731–5738

  13. [13]

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, DDL Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models.arXiv preprint arXiv:2203.1555610 (2022)

  14. [14]

    Tongwen Huang, Zhiqi Zhang, and Junlin Zhang. 2019. FiBiNET: combining fea- ture importance and bilinear feature interaction for click-through rate prediction. InProceedings of the 13th ACM conference on recommender systems. 169–177

  15. [15]

    Yunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin, Xuanyuan Luo, Zhe Wang, Zheng Chai, Shikang Wu, Yuchao Zheng, and Jingjian Lin. 2026. HyFormer: Revis- iting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction. arXiv preprint arXiv:2601.12681(2026)

  16. [16]

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361(2020)

  17. [17]

    Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xdeepfm: Combining explicit and implicit feature in- teractions for recommender systems. InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. 1754–1763

  18. [18]

    Zida Liang, Changfa Wu, Dunxian Huang, Weiqiang Sun, Ziyang Wang, Yuliang Yan, Jian Wu, Yuning Jiang, Bo Zheng, Ke Chen, et al. 2025. Tbgrecall: A generative retrieval model for e-commerce recommendation scenarios. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 5863–5870

  19. [19]

    Xiao Ma, Liqin Zhao, Guan Huang, Zhi Wang, Zelin Hu, Xiaoqiang Zhu, and Kun Gai. 2018. Entire space multi-task model: An effective approach for estimating post-click conversion rate. InThe 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. 1137–1140

  20. [20]

    Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. 2018. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32

  21. [21]

    Qi Pi, Weijie Bian, Guorui Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Practice on long sequential user behavior modeling for click-through rate prediction. InProceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining. 2671–2679

  22. [22]

    Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, et al. 2026. Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free.Advances in Neural Information Processing Systems38 (2026), 100092–100118

  23. [23]

    Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Tran, Jonah Samost, et al

  24. [24]

    Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. Autoint: Automatic feature interaction learning via self- attentive neural networks. InProceedings of the 28th ACM international conference on information and knowledge management. 1161–1170

  25. [25]

    Xin Song, Xiaochen Li, Jinxin Hu, Hong Wen, Zulong Chen, Yu Zhang, Xiaoyi Zeng, and Jing Zhang. 2025. Lrea: Low-rank efficient attention on modeling long- term user behaviors for ctr prediction. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2843–2847

  26. [26]

    2024.TorchEasyRec: An Easy-to-Use Framework for Recommen- dation

    Alibaba PAI Team. 2024.TorchEasyRec: An Easy-to-Use Framework for Recommen- dation. https://github.com/alibaba/TorchEasyRec

  27. [27]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in neural information processing systems30 (2017)

  28. [28]

    Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi. 2021. Dcn v2: Improved deep & cross network and practical lessons for web-scale learning to rank systems. InProceedings of the web conference 2021. 1785–1797

  29. [29]

    Bencheng Yan, Pengjie Wang, Kai Zhang, Feng Li, Hongbo Deng, Jian Xu, and Bo Zheng. 2022. Apg: Adaptive parameter generation network for click-through rate prediction.Advances in Neural Information Processing Systems35 (2022), 24740–24752

  30. [30]

    Jiahao Yu, Haozhuang Liu, Yeqiu Yang, Lu Chen, Jian Wu, Yuning Jiang, and Bo Zheng. 2026. Transun: A preemptive paradigm to eradicate retransformation bias intrinsically from regression models in recommender systems.Advances in Neural Information Processing Systems38 (2026), 140918–140954

  31. [31]

    Liren Yu, Wenming Zhang, Silu Zhou, Tao Zhang, Zhixuan Zhang, and Dan Ou

  32. [32]

    Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhao- jie Gong, Fangda Gu, Michael He, et al. 2024. Actions speak louder than words: Trillion-parameter sequential transducers for generative recommendations.arXiv preprint arXiv:2402.17152(2024)

  33. [33]

    Buyun Zhang, Liang Luo, Yuxin Chen, Jade Nie, Xi Liu, Daifeng Guo, Yanli Zhao, Shen Li, Yuchen Hao, Yantao Yao, et al. 2024. Wukong: Towards a scaling law for large-scale recommendation.arXiv preprint arXiv:2403.02545(2024)

  34. [34]

    Xuxin Zhang, Di Wang, Dehong Gao, Wen Jiang, Wei Ning, Yang Zhou, and Chen Wang. 2022. Revisiting cold-start problem in ctr prediction: Augmenting embedding via gan. InProceedings of the 31st ACM International Conference on Information & Knowledge Management. 4702–4706

  35. [35]

    Zhaoqi Zhang, Haolei Pei, Jun Guo, Tianyu Wang, Yufei Feng, Hui Sun, Shaowei Liu, and Aixin Sun. 2026. Onetrans: Unified feature interaction and sequence modeling with one transformer in industrial recommender. InProceedings of the ACM Web Conference 2026. 8162–8170

  36. [36]

    Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep interest evolution network for click-through rate prediction. InProceedings of the AAAI conference on artificial intelligence, Vol. 33. 5941–5948

  37. [37]

    Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep interest network for click-through rate prediction. InProceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining. 1059–1068

  38. [38]

    Jie Zhu, Zhifang Fan, Xiaoxie Zhu, Yuchen Jiang, Hangyu Wang, Xintian Han, Haoran Ding, Xinmin Wang, Wenlin Zhao, Zhen Gong, et al. 2025. Rankmixer: Scaling up ranking models in industrial recommenders. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 6309–6316

  39. [39]

    Hua Zong, Qingtao Zeng, Zhengxiong Zhou, Zhihua Han, Zhensong Yan, Mingjie Liu, Hechen Sun, Jiawei Liu, Yiwen Hu, Qi Wang, et al. 2025. RecIS: Sparse to Dense, A Unified Training Framework for Recommendation Models.arXiv preprint arXiv:2509.20883(2025). KDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Zhentao Song et al. A Detailed Training Alg...

  40. [2016]

    InProceedings of the 1st workshop on deep learning for recommender systems

    Wide & deep learning for recommender systems. InProceedings of the 1st workshop on deep learning for recommender systems. 7–10

  41. [2023]

    Recommender systems with generative retrieval.Advances in Neural Information Processing Systems36 (2023), 10299–10315

  42. [2025]

    HHFT: Hierarchical Heterogeneous Feature Transformer for Recommenda- tion Systems.arXiv preprint arXiv:2511.20235(2025)