Pith. sign in

REVIEW 3 major objections 5 minor 102 references

UniRank is an open benchmark that standardizes comparison of unified ranking models, and under that protocol no single architecture wins across platforms or tasks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 11:04 UTC pith:SYSVWUAZ

load-bearing objection A genuinely useful benchmark and toolkit for unified ranking models, but the headline rankings are single-seed, four-decimal comparisons that need error bars before being treated as architecture-level truth. the 3 major comments →

arxiv 2607.19987 v2 pith:SYSVWUAZ submitted 2026-07-22 cs.IR

UniRank: Benchmarking Ranking Models for Unified Sequential Modeling and Feature Interaction

classification cs.IR
keywords ranking modelsbenchmarkingsequential modelingfeature interactionmulti-task learningpointwise autoregressive supervisionscaling lawsrecommender systems
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that ranking models which combine sequential modeling of user behavior with feature interaction can be compared fairly and reproducibly under a single protocol. It proposes UniRank, built on chronological pointwise autoregressive supervision, capacity-controlled hyperparameters, shared training and evaluation pipelines, and efficiency optimizations that lower hardware barriers. Benchmarking 15 representative models on five large public datasets, it finds that neither of the two architectural paradigms (stacked vs. layer-wise unification) dominates and that no model wins everywhere; performance is dataset- and task-dependent. If the protocol is sound, academic researchers gain a reproducible basis for studying scaling laws, long-sequence behavior, and multi-task trade-offs without industrial infrastructure.

Core claim

On its own terms, the paper claims that UniRank is the first open benchmark dedicated to ranking architectures that unify sequential modeling and feature interaction, and that it provides a reproducible foundation for comparing such models. The benchmark's distinctive machinery is a chronological pointwise autoregressive supervision paradigm: every eligible position in a user's full-feedback behavior sequence becomes a supervised sample predicting the target from the strictly preceding events, producing dense gradients that scale with sequence length and number of feedback tasks. Under this protocol, together with six reproducibility requirements spanning data, tasks, code, hyperparameters,

What carries the argument

The load-bearing object is the chronological pointwise autoregressive supervision paradigm: a training objective in which each position in the chronological full-feedback sequence is an independent binary sample predicted from the strictly preceding events, yielding up to K supervised losses per position for K feedback tasks. It is supported by a unified token-based modeling interface that splits architectures into stacked (sequence modeling before feature interaction) and layer-wise (both within each layer) families, and by six reproducibility requirements that align data processing, task definitions, code, capacity, training, and efficiency so that measured differences are attributed to ar

Load-bearing premise

The rankings depend on the assumption that a single one-epoch run, tuned by grid search on a short noisy validation period, produces stable, architecture-driven performance differences rather than tuning luck or run-to-run noise.

What would settle it

Re-run the 15 models with identical hyperparameters but 5–10 different random seeds on the same chronological splits and examine whether the top-three rankings and four-decimal AUC gaps persist. If rank order changes across seeds, the benchmark's claim to provide a reproducible basis for comparing unified ranking models is undermined; if it persists, the claim is supported.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Model rankings from UniRank can be trusted as architecture-driven rather than pipeline-driven, because data, optimization, capacity, and efficiency are controlled.
  • Academic groups can study scaling laws for ranking models under limited compute using five public datasets, including sequences longer than 100,000 interactions.
  • Practitioners can use the handbook results as practical defaults: GeLU and SiLU activations and gated attention are robust general-purpose choices, while optimizer and tokenization choices are dataset-dependent.
  • The chronological pointwise autoregressive paradigm makes negative feedback events both context and supervision, which could improve multi-task ranking models on long behavior histories.
  • The benchmark provides a foundation for extending to multimodal ranking and additional large-scale datasets.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: Because the benchmark runs each configuration once for one epoch with a grid search on a short validation period, the four-decimal AUC gaps separating top models may not survive repeated-seed re-runs; a multi-seed extension would sharpen the benchmark's rank order.
  • Editorial inference: The dataset-dependent winner pattern suggests that claims of a universally superior ranking architecture should be treated skeptically; future evaluation protocols may need to report confidence intervals and cross-platform consistency rather than a single leaderboard.
  • Editorial inference: The chronological pointwise autoregressive supervision is a transferable training paradigm for any sequential multi-label prediction task where negative events are observable, beyond ad ranking and recommendation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces UniRank, a benchmark and PyTorch toolkit for ranking models that unify sequential modeling and feature interaction. It proposes a standardized evaluation protocol based on chronological pointwise autoregressive supervision, capacity-controlled model configuration (embedding dimension 16, depth 3, token dimension 256, sequence length 100), and a unified efficiency-oriented training pipeline. The authors evaluate 15 unified ranking models on five large-scale public datasets (QK-Video, KuaiRand, TAAC-25, Taobao, MerRec), report per-task AUC and Logloss, provide efficiency comparisons, and present handbook-style ablations on tokenization, attention activations, architectural enhancements, optimizers, and scaling. The central claim is that UniRank is the first open benchmark dedicated to these unified architectures and that it provides a reproducible basis for comparing them.

Significance. If the claims are supported, UniRank would be a useful community asset: the preprocessing details in Appendix A are unusually explicit, the capacity controls are sensible, the toolkit reportedly includes substantial efficiency optimizations, and the release of code, processed data, and configurations lowers the barrier to reproducing large-scale ranking experiments. The paper also documents design choices (chronological split, multi-label feedback, one-epoch training) that are often under-reported. The main weakness is statistical: the headline rankings in Table 3 come from a single one-epoch run per configuration with a validation-based grid search and no variance estimates. Because many rank boundaries differ only in the fourth decimal, the paper's qualitative conclusions (e.g., 'HeMix leads', 'neither paradigm dominates', 'dataset-dependent performance') may not be robust to run noise or tuning luck. The benchmark's value as a trustworthy comparison ground therefore depends on addressing this gap.

major comments (3)
  1. [§3.3, Table 3; §3.2.6] The central ranking claim rests entirely on single one-epoch runs with no repeated seeds, confidence intervals, or significance tests. Many rank boundaries are at four-decimal precision, e.g., QK-Video Follow Logloss 0.0049 vs 0.0050, and TAAC-25 Click Logloss 0.0335 vs 0.0365. With one run, these gaps are plausibly within run-to-run variance, and the per-model validation grid search introduces selection bias when reporting test scores. The autoregressive sampling in Eq. (6) makes test samples within a user trajectory strongly autocorrelated, so the effective sample size is far smaller than the row count, further inflating variance. This undermines the qualitative conclusions in §3.3. The authors should report at least 3–5 repeated seeds per configuration (or a statistically justified subsampling procedure) with mean and standard deviation, and/or significance tests, or explicitly temper
  2. [§3.2.6, §3.2.5] The paper claims UniRank provides a 'reproducible basis,' but it does not disclose random seeds, data ordering, or deterministic training settings. DDP, mixed precision, and torch.compile are nondeterministic in general, so two runs of the same configuration can differ. Without seed values and software versions, exact reproduction is not possible, and the reported single-run numbers do not indicate what a user should expect. Please disclose seeds for all randomized components (data loading, initialization, dropout, etc.) and, if feasible, report a same-configuration re-run on at least one dataset to demonstrate stability. This is load-bearing for the 'reproducible' claim.
  3. [§3.3] The bullet conclusions—'Both unified interaction paradigms are competitive', 'No universally superior model exists', 'Model performance is highly dataset-dependent'—are drawn from Table 3. If the underlying rankings are noisy, these conclusions may be artifacts. Even if a full repeated-seed study is infeasible for all 15 models, the authors could run a small subsample or bootstrap on the test set to estimate whether the reported gaps exceed measurement error. Without that, the qualitative claims should be phrased as observations on the specific runs rather than benchmark findings.
minor comments (5)
  1. [Table 1] The ✓/△/✗ ratings for the six reproducibility requirements are subjective. It would help to define explicit criteria for each symbol (e.g., 'full public source code available and runnable' vs 'partial implementations').
  2. [§3.2.3] The critique of user-disjoint splits is somewhat one-sided; such splits are a legitimate evaluation for cold-start, and the temporal split is not always superior. A brief discussion of the trade-off would strengthen the motivation.
  3. [Figure 4] The labels 'All (1 GPU)' and 'All (4 GPUs)' are easy to confuse in the figure; consider renaming to 'All optimizations, 1 GPU' and 'All optimizations, 4 GPUs'. Also, the speedup numbers in the text and figure should be cross-checked.
  4. [References] Reference [13] is titled 'Deep Session Interest Network' but the method is usually called DEIN (Deep Interest Evolution Network); please verify the title and citation.
  5. [Header] The conference placeholder 'Conference acronym 'XX, June 03–05, 2018, Woodstock, NY' should be removed or replaced before submission.

Circularity Check

0 steps flagged

No significant circularity: the benchmark is an empirical protocol and curated results, not a derivation from fitted constants.

full rationale

The paper contains no derivation chain whose output reduces to its own inputs. UniRank's contributions are an evaluation protocol, an optimized toolkit, released datasets, and measured model results; these are empirical contributions rather than claims derived from fitted constants or self-citations. The self-citations ([30,31]) appear only in the related-work survey and in an example list of feature-interaction models; they are not load-bearing for the benchmark's validity or for any reported ranking. The per-model grid search (Section 3.2.6: 'A grid search tunes model-specific hyperparameters') is disclosed and is performed on the validation period, with test performance reported on a chronologically held-out period (Section 3.2.3), so it is standard model selection rather than a fitted input being renamed as a prediction. The chronological pointwise autoregressive objective of Eq. 6 is a training-protocol choice, not a result predicted from that objective. The concern raised by the skeptical reader about single-seed runs and four-decimal AUC gaps is a legitimate statistical-reliability and reproducibility-completeness issue, but it is not circularity: no reported number is forced by construction, and no equation equates a fitted parameter with an output claim. The manuscript also contains no limitation statement asserting circularity or omitted support for a derivation; the absence of a limitation statement on variance estimation is a completeness concern, not a circular step. Therefore, no circular step can be exhibited with a specific quote and reduction, and the appropriate finding is no significant circularity.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 0 invented entities

The benchmark's central claims rest on hand-set experimental controls (sequence length, capacity, epochs, OOV threshold) and on domain assumptions about external validity of public datasets and the fairness of single-run evaluation. No new theoretical entities are introduced.

free parameters (5)
  • Main benchmark sequence length L=100 = 100
    Section 3.2.1 fixes sequence length at 100 for all models; this truncation is necessary for fair comparison but strongly disadvantages models whose design targets long sequences (the paper still advertises 1e5-length support).
  • Capacity controls: embedding dim=16, depth=3, token dim=256 = 16 / 3 / 256
    Section 3.2.6 sets these capacity-matching values by hand; the scaling-law study (Figure 6) shows models respond differently to width vs depth, so these choices shape rankings.
  • Per-model grid search on validation AUC = model-specific (not enumerated)
    Section 3.2.6 'A grid search tunes model-specific hyperparameters.' This is disclosed, but tuning budgets and search spaces are not specified; rankings can depend on tuning effort.
  • OOV threshold min_count=2 = 2
    Appendix A.2 sets min_count=2 for all datasets; increasing to 10 for memory reduction is suggested. This affects vocabulary size and long-tail feature representation.
  • One training epoch = 1
    Section 3.2.6 trains each model for one epoch, citing production practice [86]; no sensitivity analysis for longer training is provided, so slow-converging models may be penalized.
axioms (4)
  • domain assumption The five public datasets, after UniRank preprocessing, are faithful proxies for industrial ranking conditions.
    Central claim about narrowing academia-industry gap rests on external validity of QK-Video, KuaiRand, TAAC-25, Taobao, MerRec; preprocessing rules (e.g., TAAC-25 removes all-negative users) alter the distribution.
  • domain assumption Chronological pointwise autoregressive supervision is an appropriate and unbiased evaluation protocol.
    Section 2.2.2 justifies dense gradients and full-feedback supervision, but no comparison to the traditional new-impression-only protocol is reported; the choice could favor certain architectures.
  • domain assumption Single-epoch training with grid-searched hyperparameters yields rankings that reflect architecture, not tuning.
    Section 3.2.6 asserts this for fairness; no repeated seeds or learning-curve analysis are provided, so rank stability is unmeasured.
  • domain assumption AUC and Logloss, averaged per task, are sufficient to summarize multi-task ranking performance.
    Metrics are standard, but dataset-level conclusions in Table 3 rely on many task-metric pairs; no composite significance testing is used.

pith-pipeline@v1.3.0-alltime-deepseek · 25659 in / 11978 out tokens · 116719 ms · 2026-08-01T11:04:44.026518+00:00 · methodology

0 comments
read the original abstract

Ranking is a core stage in online advertising and recommender systems. Modern ranking models increasingly unify sequential modeling and feature interaction, yet many advances rely on proprietary data, closed implementations, and large-scale industrial infrastructure. This setting limits reproducible comparison and hinders academic study of scaling laws, long-sequence modeling, and multi-task ranking. To address these limitations, this paper proposes UniRank, an open benchmark for ranking models that unify sequential modeling and feature interaction. UniRank uses chronological pointwise autoregressive supervision, standardizes evaluation across feedback tasks, and provides a PyTorch toolkit with Distributed Data Parallel training, operator optimization, mixed-precision training, attention optimization, and other efficiency techniques that reduce hardware requirements. We benchmark 15 representative unified ranking models on five large-scale public datasets from short-video, advertising, and e-commerce platforms, with the largest dataset containing over 700 million instances and the longest behavior sequence exceeding 10^5 interactions. UniRank provides a reproducible basis for comparing unified ranking models, studying scaling laws under limited compute, and narrowing the gap between academic and industrial ranking research. We believe UniRank benefits researchers, practitioners, and beginners through reproducible experiments, production-oriented evaluation, and accessible implementations. Code and data are available at https://github.com/salmon1802/UniRank.

Figures

Figures reproduced from arXiv: 2607.19987 by Honghao Li, Kangyi Lin, Xianquan Wang, Yiwen Zhang, Yi Zhang, Zibin Zhang.

Figure 1
Figure 1. Figure 1: Traditional single-task New-Impression-Only paradigm. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Pointwise autoregressive paradigm. 2.2 The Evolution of Training Paradigms 2.2.1 Traditional Single-Task New-Impression-Only Para￾digm. Traditional ranking pipelines retain only positive behaviors in the user history and use this history to predict a new impression [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Comparison of randomized user-disjoint and per-user [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Training speedup and per-GPU peak memory. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Efficiency comparisons on the KuaiRand dataset (Sequence Length = 100). [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Scaling law study on KuaiRand. and Share. SOAP performs best on MerRec, mainly through im￾provements on Cart, Offer, and Purchase. This difference suggests that optimizer effectiveness depends on the gradient geometry of individual datasets and feedback tasks. Muon provides moderate and relatively stable improvements, whereas RMSProp remains close to AdamW. Scion slightly improves MerRec but substantially … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

102 extracted references · 2 canonical work pages · 2 internal anchors

  1. [1]

    Bruce Croft

    Qingyao Ai, Keping Bi, Jiafeng Guo, and W. Bruce Croft. 2018. Learning A Deep Listwise Context Model for Ranking Refinement. InProceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, Ann Arbor, MI, USA, 135–144. doi:10.1145/3209978.3209985

  2. [2]

    Alibaba. 2018. Taobao Display Ad Click-Log And User Behavior Log. Tianchi dataset

  3. [3]

    Zheng Chai, Qin Ren, Xijun Xiao, Huizhi Yang, Bo Han, Sijun Zhang, Di Chen, Hui Lu, Wenlin Zhao, Lele Yu, Xionghang Xie, Shiru Ren, Xiang Sun, Yaocheng Tan, Peng Xu, Yuchao Zheng, and Di Wu. 2025. LONGER: Scaling up Long Sequence Modeling in Industrial Recommenders.Proceedings of the Nineteenth ACM Conference on Recommender Systems(2025). arXiv:2505.04421

  4. [4]

    Jianxin Chang, Chenbin Zhang, Yiqun Hui, Dewei Leng, Yanan Niu, Yang Song, and Kun Gai. 2023. Pepnet: Parameter And Embedding Personalized Network for Infusing with Personalized Prior Information. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3795–3804

  5. [5]

    Jin Chen, Shangyu Zhang, Bin Hu, Chao Zhou, Junwei Pan, Gengsheng Xue, Wentao Ning, Gengyu Weng, Wang Zheng, Shaohua Liu, Zeen Xu, Chengyuan Mai, Shijie Quan, Tingyu Jiang, Lifeng Wang, Shudong Huang, Chengguo Yin, Haijie Gu, and Jie Jiang. 2026. RankUp: Towards High-Rank Representations for Large Scale Advertising Recommender Systems.CoRRabs/2604.17878 (...

  6. [6]

    Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al

  7. [7]

    Cheng, Y

    Z. Cheng, Y. Liu, J. Zhang, et al . 2026. Beyond Positive Signals: Unlocking Implicit Negative Behaviors for Enhanced Sequential User Modeling.arXiv preprint arXiv:2606.15252(2026)

  8. [8]

    Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep Neural Networks for YouTube Recommendations. InProceedings of the 10th ACM Conference on Recommender Systems. ACM, Boston, MA, USA, 191–198. doi:10.1145/2959100. 2959190

  9. [9]

    Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. InAdvances in Neural Information Processing Systems, Vol. 35. Curran Associates, Inc., Red Hook, NY, USA, 16344–16359

  10. [10]

    Wei Deng, Junwei Pan, Tian Zhou, Deguang Kong, Aaron Flores, and Guang Lin. 2021. Deeplight: Deep Lightweight Feature Interactions for Accelerating CTR Predictions in Ad Serving. InProceedings of the 14th ACM International Conference on Web Search and Data Mining. 922–930

  11. [11]

    Qin Ding, Kevin Course, Linjian Ma, Jianhui Sun, Ruochen Liu, Zhao Zhu, Chunx- ing Yin, Wei Li, Dai Li, Yu Shi, Xuan Cao, Ze Yang, Han Li, Xing Liu, Bi Xue, Hongwei Li, Rui Jian, Daisy Shi He, Jing Qian, Matt Ma, Qunshu Zhang, and Rui Li. 2026. Bending The Scaling Law Curve in Large-Scale Recommendation Systems.arXiv preprint arXiv:2602.16986(2026)

  12. [12]

    Stefan Elfwing, Eiji Uchibe, and Kenji Doya. 2018. Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning. Neural Networks107 (2018), 3–11

  13. [13]

    Yufei Feng, Fuyu Lv, Weichen Shen, Menghan Wang, Fei Sun, Yu Zhu, and Keping Yang. 2019. Deep Session Interest Network for Click-through Rate Prediction. InProceedings of the 28th International Joint Conference on Artificial Intelligence. 2301–2307

  14. [14]

    Chongming Gao, Shijun Li, Wenqiang Lei, et al. 2022. KuaiRand: An Unbiased Sequential Recommendation Dataset with Randomly Exposed Videos. InPro- ceedings of the 31st ACM International Conference on Information & Knowledge Management. 3953–3957

  15. [15]

    Siyu Gu and Xiangrong Sheng. 2022. On Ranking Consistency of Pre-Ranking Stage.CoRRabs/2205.01289 (2022). arXiv:2205.01289 doi:10.48550/arXiv.2205. 01289

  16. [16]

    Huan Gui, Ruoxi Wang, Ke Yin, Long Jin, Maciej Kula, Taibai Xu, Lichan Hong, and Ed H. Chi. 2023. Hiformer: Heterogeneous Feature Interactions Learning with Transformers for Recommender Systems.arXiv preprint arXiv:2311.05884 (2023)

  17. [17]

    Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: A Factorization-Machine Based Neural Network for CTR Prediction. InProceedings of the 26th International Joint Conference on Artificial Intelligence (Melbourne, Australia)(IJCAI’17). AAAI Press, 1725–1731

  18. [18]

    Mingming Ha, Guanchen Wang, Linxun Chen, Xuan Rao, Yuexin Shi, Tianbao Ma, Zhaojie Liu, Yunqian Fan, Zilong Lu, Yanan Niu, Han Li, and Kun Gai. 2026. UniMixer: A Unified Architecture for Scaling Laws in Recommendation Systems. arXiv preprint arXiv:2604.00590(2026)

  19. [20]

    Ruining He, Anirudh Ravula, Bhargav Kanagal, and Joshua Ainslie. 2021. Re- alFormer: Transformer Likes Residual Attention. InFindings of the Association for Computational Linguistics: ACL-IJCNLP 2021. Association for Computational Linguistics, Online, 929–943. doi:10.18653/v1/2021.findings-acl.81

  20. [21]

    Dan Hendrycks and Kevin Gimpel. 2016. Gaussian Error Linear Units (GELUs). arXiv:1606.08415 [cs.LG]

  21. [22]

    Alex Henry, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan Chen. 2020. Query-Key Normalization for Transformers. InFindings of the Associ- ation for Computational Linguistics: EMNLP 2020. Association for Computational Linguistics, Online, 4246–4253. doi:10.18653/v1/2020.findings-emnlp.379

  22. [23]

    Xu Huang, Hao Zhang, Zhifang Fan, Yunwen Huang, Zhuoxing Wei, Zheng Chai, Jinan Ni, Yuchao Zheng, and Qiwei Chen. 2026. MixFormer: Co-Scaling up Dense And Sequence in Industrial Recommenders.arXiv preprint arXiv:2602.14110 (2026)

  23. [24]

    Yunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin, Xuanyuan Luo, Zhe Wang, Zheng Chai, Shikang Wu, Yuchao Zheng, and Jingjian Lin. 2026. HyFormer: Revisiting The Roles of Sequence Modeling And Feature Interaction in CTR Prediction.arXiv preprint arXiv:2601.12681(2026)

  24. [25]

    Huawei. 2021. An Open-Source CTR Prediction Library. https://fuxictr.github.io

  25. [26]

    Yuchen Jiang, Jie Zhu, Xintian Han, Hui Lu, Kunmin Bai, Mingyu Yang, Shikang Wu, Ruihao Zhang, Wenlin Zhao, Shipeng Bai, Sijin Zhou, Huizhi Yang, Tianyi Liu, Wenda Liu, Ziyan Gong, Haoran Ding, Zheng Chai, Deping Xie, Zhe Chen, Yuchao Zheng, and Peng Xu. 2026. TokenMixer-Large: Scaling up Large Ranking Models in Industrial Recommenders.arXiv preprint arXi...

  26. [27]

    Yuchin Juan, Yong Zhuang, Wei-Sheng Chin, and Chih-Jen Lin. 2016. Field- Aware Factorization Machines for CTR Prediction. InProceedings of the 10th ACM Conference on Recommender Systems. 43–50

  27. [28]

    Martin Kaloev and Georgi Krastev. 2021. Comparative Analysis of Activation Functions Used in The Hidden Layers of Deep Neural Networks. In2021 3rd International Congress on Human-Computer Interaction, Optimization and Robotic Applications (HORA). IEEE, 1–5

  28. [29]

    Chao Li, Zhiyuan Liu, Mengmeng Wu, Yuchi Xu, Huan Zhao, Pipei Huang, Guoliang Kang, Qiwei Chen, Wei Li, and Dik Lun Lee. 2019. Multi-Interest Network with Dynamic Routing for Recommendation at Tmall. InProceedings of the 28th ACM International Conference on Information and Knowledge Management. ACM, Beijing, China, 2615–2623. doi:10.1145/3357384.3357814

  29. [30]

    Honghao Li, Yiwen Zhang, Yi Zhang, Hanwei Li, Lei Sang, and Jieming Zhu

  30. [31]

    Honghao Li, Yiwen Zhang, Yi Zhang, Lei Sang, and Jieming Zhu. 2025. Revisiting Feature Interactions from The Perspective of Quadratic Neural Networks for Click-through Rate Prediction. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 1365–1375

  31. [32]

    Kaiyuan Li, Yongxiang Tang, Wenzheng Shu, Yanxiang Zeng, Chao Wang, Yanhua Cheng, Xialong Liu, and Peng Jiang. 2025. Aggregate And Broadcast: Scalable And Efficient Feature Interaction for Recommender Systems.arXiv preprint arXiv:2508.11565(2025)

  32. [33]

    Lichi Li, Zainul Abi Din, Zhen Tan, Sam London, Tianlong Chen, and Ajay Dap- tardar. 2024. MerRec: A Large-Scale Multipurpose Mercari Dataset for Consumer- to-Consumer Recommendation Systems.arXiv preprint arXiv:2402.14230(2024)

  33. [34]

    Xiaopeng Li, Jingtong Gao, Pengyue Jia, Xiangyu Zhao, Yichao Wang, Wanyu Wang, Yejing Wang, Yuhao Wang, Huifeng Guo, and Ruiming Tang. 2025. Scenario-Wise Rec: A Multi-Scenario Recommendation Benchmark. InProceed- ings of the 34th ACM International Conference on Information and Knowledge Management. ACM, 1685–1695. doi:10.1145/3746252.3761137

  34. [35]

    Zekun Li, Zeyu Cui, Shu Wu, Xiaoyu Zhang, and Liang Wang. 2019. FiGNN: Modeling Feature Interactions via Graph Neural Networks for CTR Prediction. InProceedings of the 28th ACM International Conference on Information and Knowledge Management. 539–548

  35. [36]

    Zekun Li, Shu Wu, Zeyu Cui, and Xiaoyu Zhang. 2022. GraphFM: Graph Factoriza- tion Machines for Feature Interaction Modeling.arXiv preprint arXiv:2105.11866 (2022)

  36. [37]

    Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xDeepFM: Combining Explicit And Implicit Feature Interactions for Recommender Systems. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 1754–1763

  37. [38]

    Bin Liu, Niannan Xue, Huifeng Guo, Ruiming Tang, Stefanos Zafeiriou, Xiuqiang He, and Zhenguo Li. 2020. AutoGroup: Automatic Feature Grouping for Modelling Explicit High-Order Feature Interactions in CTR Prediction. InProceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. 199–208

  38. [39]

    Bin Liu, Chenxu Zhu, Guilin Li, Weinan Zhang, Jincai Lai, Ruiming Tang, Xi- uqiang He, Zhenguo Li, and Yong Yu. 2020. AutoFIS: Automatic Feature In- teraction Selection in Factorization Models for Click-through Rate Prediction. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Conference acronym ’XX, June 03–05, 2018, Woodstock, N...

  39. [40]

    Jingyuan Liu, Jianlin Su, Xingcheng Yao, Zhejun Jiang, Guokun Lai, Yulun Du, Yidao Qin, Weixin Xu, Enzhe Lu, Junjie Yan, et al. 2025. Muon Is Scalable for LLM Training. arXiv:2502.16982 [cs.LG]

  40. [41]

    Mingyang Liu, Yong Bai, Zhangming Chan, Sishuo Chen, Xiang-Rong Sheng, Han Zhu, Jian Xu, and Xinyang Chen. 2026. EST: Towards Efficient Scaling Laws in Click-through Rate Prediction via Unified Modeling.arXiv preprint arXiv:2602.10811(2026)

  41. [42]

    Qi Liu, Kai Zheng, Rui Huang, Wuchao Li, Kuo Cai, Yuan Chai, Yanan Niu, Yiqun Hui, Bing Han, Na Mou, Hongning Wang, Wentian Bao, Yunen Yu, Guorui Zhou, Han Li, Yang Song, Defu Lian, and Kun Gai. 2025. RecFlow: An Industrial Full Flow Recommendation Dataset. InThe Thirteenth International Conference on Learning Representations. https://openreview.net/forum...

  42. [43]

    Wang, and Masahito Ueda

    Ziyin Liu, Zhikang T. Wang, and Masahito Ueda. 2020. LaProp: Separating Momentum And Adaptivity in Adam. arXiv:2002.04839 [cs.LG]

  43. [44]

    Ilya Loshchilov and Frank Hutter. 2017. Decoupled Weight Decay Regularization. arXiv preprint arXiv:1711.05101(2017)

  44. [45]

    Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. 2018. Modeling Task Relationships in Multi-Task Learning with Multi-Gate Mixture- of-Experts. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 1930–1939

  45. [46]

    Xiao Ma, Liqin Zhao, Guan Huang, Zhi Wang, Zelin Hu, Xiaoqiang Zhu, and Kun Gai. 2018. Entire Space Multi-Task Model: An Effective Approach for Estimating Post-Click Conversion Rate. InThe 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. 1137–1140

  46. [47]

    Ze Meng, Jinnian Zhang, Yumeng Li, Jiancheng Li, Tanchao Zhu, and Lifeng Sun

  47. [48]

    Diganta Misra. 2019. Mish: A Self Regularized Non-Monotonic Activation Func- tion. arXiv:1908.08681 [cs.LG]

  48. [49]

    Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherni- avskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, Volodymyr Kon- dratenko, Stephanie Pereira, Xianjie Chen, Wenlin Chen, Vijay Rao,...

  49. [50]

    Yongxin Ni, Yu Cheng, Xiangyan Liu, Junchen Fu, Youhua Li, Xiangnan He, Yongfeng Zhang, and Fajie Yuan. 2023. A Content-Driven Micro-Video Recom- mendation Dataset at Scale.arXiv preprint arXiv:2309.15379(2023)

  50. [51]

    Junwei Pan, Jian Xu, Alfonso Lobos Ruiz, Wenliang Zhao, Shengjun Pan, Yu Sun, and Quan Lu. 2018. Field-Weighted Factorization Machines for Click-through Rate Prediction in Display Advertising. InProceedings of the 2018 World Wide Web Conference. 1349–1357

  51. [52]

    Junwei Pan, Wei Xue, Ximei Wang, Haibin Yu, Xun Liu, Shijie Quan, Xueming Qiu, Dapeng Liu, Lei Xiao, and Jie Jiang. 2024. Ads Recommendation in A Collapsed And Entangled World. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 5566–5577

  52. [53]

    Junwei Pan, Wei Xue, Chao Zhou, Xing Zhou, Lunan Fan, Yanbo Wang, Haoran Xin, Zhiyu Hu, Yaozheng Wang, Fengye Xu, Yurong Yang, Xiaotian Li, Junbang Huo, Wentao Ning, Yuliang Sun, Chengguo Yin, Jun Zhang, Shudong Huang, Lei Xiao, Huan Yu, Irwin King, Haijie Gu, and Jie Jiang. 2026. Tencent Advertising Algorithm Challenge 2025: All-Modality Generative Recom...

  53. [54]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al

  54. [55]

    Changhua Pei, Yi Zhang, Yongfeng Zhang, Fei Sun, Xiao Lin, Hanxiao Sun, Jian Wu, Peng Jiang, Junfeng Ge, Wenwu Ou, and Dan Pei. 2019. Personalized Re- Ranking for Recommendation. InProceedings of the 13th ACM Conference on Recommender Systems. ACM, Copenhagen, Denmark, 3–11. doi:10.1145/3298689. 3347000

  55. [56]

    Thomas Pethick, Wanyun Xie, Kimon Antonakopoulos, Zhenyu Zhu, Antonio Silveti-Falls, and Volkan Cevher. 2025. Training Deep Learning Models with Norm-Constrained LMOs. arXiv:2502.07529 [cs.LG]

  56. [57]

    Aleksandr Poslavsky, Alexander D’yakonov, Yuriy Dorn, and Andrey Zimovnov

  57. [58]

    Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, Dayiheng Liu, Jingren Zhou, and Junyang Lin. 2025. Gated Attention for Large Language Models: Non-Linearity, Sparsity, And Attention-Sink-Free. Advances in Neural Information Process- ing Systems 38. https://proceedings.neurips.cc/paper_files/pa...

  58. [59]

    Yanru Qu, Han Cai, Kan Ren, Weinan Zhang, Yong Yu, Ying Wen, and Jun Wang

  59. [60]

    Steffen Rendle. 2010. Factorization Machines. In2010 IEEE International Confer- ence on Data Mining. IEEE, 995–1000

  60. [61]

    Matthew Richardson, Ewa Dominowska, and Robert Ragno. 2007. Predicting Clicks: Estimating The Click-through Rate for New Ads. InProceedings of the 16th International Conference on World Wide Web. 521–530

  61. [62]

    VK-LSVD: A Large-Scale Industrial Dataset for Short-Video Recommenda- tion.arXiv preprint arXiv:2602.04567(2026)

  62. [63]

    Weichen Shen. 2017. DeepCTR: Easy-to-Use, Modular And Extendible Package of Deep-Learning Based CTR Models. https://github.com/shenweichen/deepctr

  63. [64]

    Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. AutoInt: Automatic Feature Interaction Learning via Self- Attentive Neural Networks. InProceedings of the 28th ACM International Confer- ence on Information and Knowledge Management. 1161–1170

  64. [65]

    In2016 IEEE 16th International Conference on Data Mining (ICDM)

    Product-Based Neural Networks for User Response Prediction. In2016 IEEE 16th International Conference on Data Mining (ICDM). IEEE, 1149–1154

  65. [66]

    Yang Sun, Junwei Pan, Alex Zhang, and Aaron Flores. 2021. FM2: Field-Matrixed Factorization Machines for Recommender Systems. InProceedings of the Web Conference 2021. 2828–2837

  66. [67]

    Tijmen Tieleman and Geoffrey Hinton. 2012. RMSProp: Divide The Gradient by A Running Average of Its Recent Magnitude. COURSERA: Neural Networks for Machine Learning, Lecture 6.5

  67. [68]

    Olivier Roy and Martin Vetterli. 2007. The Effective Rank: A Measure of Effective Dimensionality. InProceedings of the 15th European Signal Processing Conference. IEEE, Poznan, Poland, 606–610

  68. [69]

    Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Lucas Janson, and Sham Kakade. 2025. SOAP: Improving And Stabilizing Sham- poo Using Adam for Language Modeling. International Conference on Learn- ing Representations. https://proceedings.iclr.cc/paper_files/paper/2025/hash/ e988664070e9591f93fdcf605f7dc623-Abstract-Conference.html

  69. [70]

    Fangye Wang, Guowei Yang, Xiaojiang Zhou, Song Yang, and Pengjie Wang. 2026. Query-Mixed Interest Extraction And Heterogeneous Interaction: A Scalable CTR Model for Industrial Recommender Systems.arXiv preprint arXiv:2602.09387 (2026)

  70. [71]

    Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2024. RoFormer: Enhanced Transformer with Rotary Position Embedding. Neurocomputing568 (2024), 127063. doi:10.1016/j.neucom.2023.127063

  71. [72]

    Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi. 2021. DCNv2: Improved Deep & Cross Network And Practical Lessons for Web-Scale Learning to Rank Systems. InProceedings of the Web Conference

  72. [73]

    Zhe Wang, Liqin Zhao, Biye Jiang, Guorui Zhou, Xiaoqiang Zhu, and Kun Gai

  73. [74]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need.Advances in Neural Information Processing Systems30 (2017)

  74. [75]

    Yuxin Wu and Kaiming He. 2018. Group Normalization. InProceedings of the European Conference on Computer Vision. Springer, Cham, 3–19. doi:10.1007/978- 3-030-01261-8_1

  75. [76]

    Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. Efficient Streaming Language Models with Attention Sinks. International Con- ference on Learning Representations. arXiv:2309.17453 [cs.CL]

  76. [77]

    Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & Cross Network for Ad Click Predictions. InProceedings of the ADKDD’17. 1–7

  77. [78]

    Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li. 2015. Empirical Evaluation of Rectified Activations in Convolutional Network.arXiv preprint arXiv:1505.00853 (2015)

  78. [79]

    Yantao Yu, Sen Qiao, Lei Shen, Bing Wang, and Xiaoyi Zeng. 2026. Beyond Dense Connectivity: Explicit Sparsity for Scalable Recommendation.arXiv preprint arXiv:2604.08011(2026). Accepted at SIGIR 2026

  79. [80]

    Guanghu Yuan, Jieyu Yang, Shujie Li, Mingjie Zhong, Ang Li, Ke Ding, Yong He, Min Yang, Liang Zhang, Xiaolu Zhang, and Linjian Mo. 2024. MMLRec: A Unified Multi-Task And Multi-Scenario Learning Benchmark for Recommenda- tion. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management(Boise, ID, USA)(CIKM ’24). Associati...

  80. [81]

    Bin Wu, Feifan Yang, Zhangming Chan, Yu-Ran Gu, Jiawei Feng, Chao Yi, Xiang- Rong Sheng, Han Zhu, Jian Xu, Mang Ye, and Bo Zheng. 2025. MUSE: A Simple Yet Effective Multimodal Search-Based Framework for Lifelong User Interest Modeling. arXiv:2512.07216 [cs.IR] https://arxiv.org/abs/2512.07216

Showing first 80 references.