REVIEW 3 major objections 5 minor 102 references
UniRank is an open benchmark that standardizes comparison of unified ranking models, and under that protocol no single architecture wins across platforms or tasks.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 11:04 UTC pith:SYSVWUAZ
load-bearing objection A genuinely useful benchmark and toolkit for unified ranking models, but the headline rankings are single-seed, four-decimal comparisons that need error bars before being treated as architecture-level truth. the 3 major comments →
UniRank: Benchmarking Ranking Models for Unified Sequential Modeling and Feature Interaction
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On its own terms, the paper claims that UniRank is the first open benchmark dedicated to ranking architectures that unify sequential modeling and feature interaction, and that it provides a reproducible foundation for comparing such models. The benchmark's distinctive machinery is a chronological pointwise autoregressive supervision paradigm: every eligible position in a user's full-feedback behavior sequence becomes a supervised sample predicting the target from the strictly preceding events, producing dense gradients that scale with sequence length and number of feedback tasks. Under this protocol, together with six reproducibility requirements spanning data, tasks, code, hyperparameters,
What carries the argument
The load-bearing object is the chronological pointwise autoregressive supervision paradigm: a training objective in which each position in the chronological full-feedback sequence is an independent binary sample predicted from the strictly preceding events, yielding up to K supervised losses per position for K feedback tasks. It is supported by a unified token-based modeling interface that splits architectures into stacked (sequence modeling before feature interaction) and layer-wise (both within each layer) families, and by six reproducibility requirements that align data processing, task definitions, code, capacity, training, and efficiency so that measured differences are attributed to ar
Load-bearing premise
The rankings depend on the assumption that a single one-epoch run, tuned by grid search on a short noisy validation period, produces stable, architecture-driven performance differences rather than tuning luck or run-to-run noise.
What would settle it
Re-run the 15 models with identical hyperparameters but 5–10 different random seeds on the same chronological splits and examine whether the top-three rankings and four-decimal AUC gaps persist. If rank order changes across seeds, the benchmark's claim to provide a reproducible basis for comparing unified ranking models is undermined; if it persists, the claim is supported.
If this is right
- Model rankings from UniRank can be trusted as architecture-driven rather than pipeline-driven, because data, optimization, capacity, and efficiency are controlled.
- Academic groups can study scaling laws for ranking models under limited compute using five public datasets, including sequences longer than 100,000 interactions.
- Practitioners can use the handbook results as practical defaults: GeLU and SiLU activations and gated attention are robust general-purpose choices, while optimizer and tokenization choices are dataset-dependent.
- The chronological pointwise autoregressive paradigm makes negative feedback events both context and supervision, which could improve multi-task ranking models on long behavior histories.
- The benchmark provides a foundation for extending to multimodal ranking and additional large-scale datasets.
Where Pith is reading between the lines
- Editorial inference: Because the benchmark runs each configuration once for one epoch with a grid search on a short validation period, the four-decimal AUC gaps separating top models may not survive repeated-seed re-runs; a multi-seed extension would sharpen the benchmark's rank order.
- Editorial inference: The dataset-dependent winner pattern suggests that claims of a universally superior ranking architecture should be treated skeptically; future evaluation protocols may need to report confidence intervals and cross-platform consistency rather than a single leaderboard.
- Editorial inference: The chronological pointwise autoregressive supervision is a transferable training paradigm for any sequential multi-label prediction task where negative events are observable, beyond ad ranking and recommendation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces UniRank, a benchmark and PyTorch toolkit for ranking models that unify sequential modeling and feature interaction. It proposes a standardized evaluation protocol based on chronological pointwise autoregressive supervision, capacity-controlled model configuration (embedding dimension 16, depth 3, token dimension 256, sequence length 100), and a unified efficiency-oriented training pipeline. The authors evaluate 15 unified ranking models on five large-scale public datasets (QK-Video, KuaiRand, TAAC-25, Taobao, MerRec), report per-task AUC and Logloss, provide efficiency comparisons, and present handbook-style ablations on tokenization, attention activations, architectural enhancements, optimizers, and scaling. The central claim is that UniRank is the first open benchmark dedicated to these unified architectures and that it provides a reproducible basis for comparing them.
Significance. If the claims are supported, UniRank would be a useful community asset: the preprocessing details in Appendix A are unusually explicit, the capacity controls are sensible, the toolkit reportedly includes substantial efficiency optimizations, and the release of code, processed data, and configurations lowers the barrier to reproducing large-scale ranking experiments. The paper also documents design choices (chronological split, multi-label feedback, one-epoch training) that are often under-reported. The main weakness is statistical: the headline rankings in Table 3 come from a single one-epoch run per configuration with a validation-based grid search and no variance estimates. Because many rank boundaries differ only in the fourth decimal, the paper's qualitative conclusions (e.g., 'HeMix leads', 'neither paradigm dominates', 'dataset-dependent performance') may not be robust to run noise or tuning luck. The benchmark's value as a trustworthy comparison ground therefore depends on addressing this gap.
major comments (3)
- [§3.3, Table 3; §3.2.6] The central ranking claim rests entirely on single one-epoch runs with no repeated seeds, confidence intervals, or significance tests. Many rank boundaries are at four-decimal precision, e.g., QK-Video Follow Logloss 0.0049 vs 0.0050, and TAAC-25 Click Logloss 0.0335 vs 0.0365. With one run, these gaps are plausibly within run-to-run variance, and the per-model validation grid search introduces selection bias when reporting test scores. The autoregressive sampling in Eq. (6) makes test samples within a user trajectory strongly autocorrelated, so the effective sample size is far smaller than the row count, further inflating variance. This undermines the qualitative conclusions in §3.3. The authors should report at least 3–5 repeated seeds per configuration (or a statistically justified subsampling procedure) with mean and standard deviation, and/or significance tests, or explicitly temper
- [§3.2.6, §3.2.5] The paper claims UniRank provides a 'reproducible basis,' but it does not disclose random seeds, data ordering, or deterministic training settings. DDP, mixed precision, and torch.compile are nondeterministic in general, so two runs of the same configuration can differ. Without seed values and software versions, exact reproduction is not possible, and the reported single-run numbers do not indicate what a user should expect. Please disclose seeds for all randomized components (data loading, initialization, dropout, etc.) and, if feasible, report a same-configuration re-run on at least one dataset to demonstrate stability. This is load-bearing for the 'reproducible' claim.
- [§3.3] The bullet conclusions—'Both unified interaction paradigms are competitive', 'No universally superior model exists', 'Model performance is highly dataset-dependent'—are drawn from Table 3. If the underlying rankings are noisy, these conclusions may be artifacts. Even if a full repeated-seed study is infeasible for all 15 models, the authors could run a small subsample or bootstrap on the test set to estimate whether the reported gaps exceed measurement error. Without that, the qualitative claims should be phrased as observations on the specific runs rather than benchmark findings.
minor comments (5)
- [Table 1] The ✓/△/✗ ratings for the six reproducibility requirements are subjective. It would help to define explicit criteria for each symbol (e.g., 'full public source code available and runnable' vs 'partial implementations').
- [§3.2.3] The critique of user-disjoint splits is somewhat one-sided; such splits are a legitimate evaluation for cold-start, and the temporal split is not always superior. A brief discussion of the trade-off would strengthen the motivation.
- [Figure 4] The labels 'All (1 GPU)' and 'All (4 GPUs)' are easy to confuse in the figure; consider renaming to 'All optimizations, 1 GPU' and 'All optimizations, 4 GPUs'. Also, the speedup numbers in the text and figure should be cross-checked.
- [References] Reference [13] is titled 'Deep Session Interest Network' but the method is usually called DEIN (Deep Interest Evolution Network); please verify the title and citation.
- [Header] The conference placeholder 'Conference acronym 'XX, June 03–05, 2018, Woodstock, NY' should be removed or replaced before submission.
Circularity Check
No significant circularity: the benchmark is an empirical protocol and curated results, not a derivation from fitted constants.
full rationale
The paper contains no derivation chain whose output reduces to its own inputs. UniRank's contributions are an evaluation protocol, an optimized toolkit, released datasets, and measured model results; these are empirical contributions rather than claims derived from fitted constants or self-citations. The self-citations ([30,31]) appear only in the related-work survey and in an example list of feature-interaction models; they are not load-bearing for the benchmark's validity or for any reported ranking. The per-model grid search (Section 3.2.6: 'A grid search tunes model-specific hyperparameters') is disclosed and is performed on the validation period, with test performance reported on a chronologically held-out period (Section 3.2.3), so it is standard model selection rather than a fitted input being renamed as a prediction. The chronological pointwise autoregressive objective of Eq. 6 is a training-protocol choice, not a result predicted from that objective. The concern raised by the skeptical reader about single-seed runs and four-decimal AUC gaps is a legitimate statistical-reliability and reproducibility-completeness issue, but it is not circularity: no reported number is forced by construction, and no equation equates a fitted parameter with an output claim. The manuscript also contains no limitation statement asserting circularity or omitted support for a derivation; the absence of a limitation statement on variance estimation is a completeness concern, not a circular step. Therefore, no circular step can be exhibited with a specific quote and reduction, and the appropriate finding is no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (5)
- Main benchmark sequence length L=100 =
100
- Capacity controls: embedding dim=16, depth=3, token dim=256 =
16 / 3 / 256
- Per-model grid search on validation AUC =
model-specific (not enumerated)
- OOV threshold min_count=2 =
2
- One training epoch =
1
axioms (4)
- domain assumption The five public datasets, after UniRank preprocessing, are faithful proxies for industrial ranking conditions.
- domain assumption Chronological pointwise autoregressive supervision is an appropriate and unbiased evaluation protocol.
- domain assumption Single-epoch training with grid-searched hyperparameters yields rankings that reflect architecture, not tuning.
- domain assumption AUC and Logloss, averaged per task, are sufficient to summarize multi-task ranking performance.
read the original abstract
Ranking is a core stage in online advertising and recommender systems. Modern ranking models increasingly unify sequential modeling and feature interaction, yet many advances rely on proprietary data, closed implementations, and large-scale industrial infrastructure. This setting limits reproducible comparison and hinders academic study of scaling laws, long-sequence modeling, and multi-task ranking. To address these limitations, this paper proposes UniRank, an open benchmark for ranking models that unify sequential modeling and feature interaction. UniRank uses chronological pointwise autoregressive supervision, standardizes evaluation across feedback tasks, and provides a PyTorch toolkit with Distributed Data Parallel training, operator optimization, mixed-precision training, attention optimization, and other efficiency techniques that reduce hardware requirements. We benchmark 15 representative unified ranking models on five large-scale public datasets from short-video, advertising, and e-commerce platforms, with the largest dataset containing over 700 million instances and the longest behavior sequence exceeding 10^5 interactions. UniRank provides a reproducible basis for comparing unified ranking models, studying scaling laws under limited compute, and narrowing the gap between academic and industrial ranking research. We believe UniRank benefits researchers, practitioners, and beginners through reproducible experiments, production-oriented evaluation, and accessible implementations. Code and data are available at https://github.com/salmon1802/UniRank.
Figures
Reference graph
Works this paper leans on
-
[1]
Qingyao Ai, Keping Bi, Jiafeng Guo, and W. Bruce Croft. 2018. Learning A Deep Listwise Context Model for Ranking Refinement. InProceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, Ann Arbor, MI, USA, 135–144. doi:10.1145/3209978.3209985
arXiv 2018
-
[2]
Alibaba. 2018. Taobao Display Ad Click-Log And User Behavior Log. Tianchi dataset
2018
-
[3]
Zheng Chai, Qin Ren, Xijun Xiao, Huizhi Yang, Bo Han, Sijun Zhang, Di Chen, Hui Lu, Wenlin Zhao, Lele Yu, Xionghang Xie, Shiru Ren, Xiang Sun, Yaocheng Tan, Peng Xu, Yuchao Zheng, and Di Wu. 2025. LONGER: Scaling up Long Sequence Modeling in Industrial Recommenders.Proceedings of the Nineteenth ACM Conference on Recommender Systems(2025). arXiv:2505.04421
Pith/arXiv arXiv 2025
-
[4]
Jianxin Chang, Chenbin Zhang, Yiqun Hui, Dewei Leng, Yanan Niu, Yang Song, and Kun Gai. 2023. Pepnet: Parameter And Embedding Personalized Network for Infusing with Personalized Prior Information. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3795–3804
2023
-
[5]
Jin Chen, Shangyu Zhang, Bin Hu, Chao Zhou, Junwei Pan, Gengsheng Xue, Wentao Ning, Gengyu Weng, Wang Zheng, Shaohua Liu, Zeen Xu, Chengyuan Mai, Shijie Quan, Tingyu Jiang, Lifeng Wang, Shudong Huang, Chengguo Yin, Haijie Gu, and Jie Jiang. 2026. RankUp: Towards High-Rank Representations for Large Scale Advertising Recommender Systems.CoRRabs/2604.17878 (...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2604.17878 2026
-
[6]
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al
- [7]
-
[8]
Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep Neural Networks for YouTube Recommendations. InProceedings of the 10th ACM Conference on Recommender Systems. ACM, Boston, MA, USA, 191–198. doi:10.1145/2959100. 2959190
doi:10.1145/2959100 2016
-
[9]
Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. InAdvances in Neural Information Processing Systems, Vol. 35. Curran Associates, Inc., Red Hook, NY, USA, 16344–16359
2022
-
[10]
Wei Deng, Junwei Pan, Tian Zhou, Deguang Kong, Aaron Flores, and Guang Lin. 2021. Deeplight: Deep Lightweight Feature Interactions for Accelerating CTR Predictions in Ad Serving. InProceedings of the 14th ACM International Conference on Web Search and Data Mining. 922–930
2021
-
[11]
Qin Ding, Kevin Course, Linjian Ma, Jianhui Sun, Ruochen Liu, Zhao Zhu, Chunx- ing Yin, Wei Li, Dai Li, Yu Shi, Xuan Cao, Ze Yang, Han Li, Xing Liu, Bi Xue, Hongwei Li, Rui Jian, Daisy Shi He, Jing Qian, Matt Ma, Qunshu Zhang, and Rui Li. 2026. Bending The Scaling Law Curve in Large-Scale Recommendation Systems.arXiv preprint arXiv:2602.16986(2026)
arXiv 2026
-
[12]
Stefan Elfwing, Eiji Uchibe, and Kenji Doya. 2018. Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning. Neural Networks107 (2018), 3–11
2018
-
[13]
Yufei Feng, Fuyu Lv, Weichen Shen, Menghan Wang, Fei Sun, Yu Zhu, and Keping Yang. 2019. Deep Session Interest Network for Click-through Rate Prediction. InProceedings of the 28th International Joint Conference on Artificial Intelligence. 2301–2307
2019
-
[14]
Chongming Gao, Shijun Li, Wenqiang Lei, et al. 2022. KuaiRand: An Unbiased Sequential Recommendation Dataset with Randomly Exposed Videos. InPro- ceedings of the 31st ACM International Conference on Information & Knowledge Management. 3953–3957
2022
-
[15]
Siyu Gu and Xiangrong Sheng. 2022. On Ranking Consistency of Pre-Ranking Stage.CoRRabs/2205.01289 (2022). arXiv:2205.01289 doi:10.48550/arXiv.2205. 01289
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2205.01289 2022
-
[16]
Huan Gui, Ruoxi Wang, Ke Yin, Long Jin, Maciej Kula, Taibai Xu, Lichan Hong, and Ed H. Chi. 2023. Hiformer: Heterogeneous Feature Interactions Learning with Transformers for Recommender Systems.arXiv preprint arXiv:2311.05884 (2023)
Pith/arXiv arXiv 2023
-
[17]
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: A Factorization-Machine Based Neural Network for CTR Prediction. InProceedings of the 26th International Joint Conference on Artificial Intelligence (Melbourne, Australia)(IJCAI’17). AAAI Press, 1725–1731
2017
-
[18]
Mingming Ha, Guanchen Wang, Linxun Chen, Xuan Rao, Yuexin Shi, Tianbao Ma, Zhaojie Liu, Yunqian Fan, Zilong Lu, Yanan Niu, Han Li, and Kun Gai. 2026. UniMixer: A Unified Architecture for Scaling Laws in Recommendation Systems. arXiv preprint arXiv:2604.00590(2026)
arXiv 2026
-
[20]
Ruining He, Anirudh Ravula, Bhargav Kanagal, and Joshua Ainslie. 2021. Re- alFormer: Transformer Likes Residual Attention. InFindings of the Association for Computational Linguistics: ACL-IJCNLP 2021. Association for Computational Linguistics, Online, 929–943. doi:10.18653/v1/2021.findings-acl.81
-
[21]
Dan Hendrycks and Kevin Gimpel. 2016. Gaussian Error Linear Units (GELUs). arXiv:1606.08415 [cs.LG]
Pith/arXiv arXiv 2016
-
[22]
Alex Henry, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan Chen. 2020. Query-Key Normalization for Transformers. InFindings of the Associ- ation for Computational Linguistics: EMNLP 2020. Association for Computational Linguistics, Online, 4246–4253. doi:10.18653/v1/2020.findings-emnlp.379
-
[23]
Xu Huang, Hao Zhang, Zhifang Fan, Yunwen Huang, Zhuoxing Wei, Zheng Chai, Jinan Ni, Yuchao Zheng, and Qiwei Chen. 2026. MixFormer: Co-Scaling up Dense And Sequence in Industrial Recommenders.arXiv preprint arXiv:2602.14110 (2026)
Pith/arXiv arXiv 2026
-
[24]
Yunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin, Xuanyuan Luo, Zhe Wang, Zheng Chai, Shikang Wu, Yuchao Zheng, and Jingjian Lin. 2026. HyFormer: Revisiting The Roles of Sequence Modeling And Feature Interaction in CTR Prediction.arXiv preprint arXiv:2601.12681(2026)
arXiv 2026
-
[25]
Huawei. 2021. An Open-Source CTR Prediction Library. https://fuxictr.github.io
2021
-
[26]
Yuchen Jiang, Jie Zhu, Xintian Han, Hui Lu, Kunmin Bai, Mingyu Yang, Shikang Wu, Ruihao Zhang, Wenlin Zhao, Shipeng Bai, Sijin Zhou, Huizhi Yang, Tianyi Liu, Wenda Liu, Ziyan Gong, Haoran Ding, Zheng Chai, Deping Xie, Zhe Chen, Yuchao Zheng, and Peng Xu. 2026. TokenMixer-Large: Scaling up Large Ranking Models in Industrial Recommenders.arXiv preprint arXi...
arXiv 2026
-
[27]
Yuchin Juan, Yong Zhuang, Wei-Sheng Chin, and Chih-Jen Lin. 2016. Field- Aware Factorization Machines for CTR Prediction. InProceedings of the 10th ACM Conference on Recommender Systems. 43–50
2016
-
[28]
Martin Kaloev and Georgi Krastev. 2021. Comparative Analysis of Activation Functions Used in The Hidden Layers of Deep Neural Networks. In2021 3rd International Congress on Human-Computer Interaction, Optimization and Robotic Applications (HORA). IEEE, 1–5
2021
-
[29]
Chao Li, Zhiyuan Liu, Mengmeng Wu, Yuchi Xu, Huan Zhao, Pipei Huang, Guoliang Kang, Qiwei Chen, Wei Li, and Dik Lun Lee. 2019. Multi-Interest Network with Dynamic Routing for Recommendation at Tmall. InProceedings of the 28th ACM International Conference on Information and Knowledge Management. ACM, Beijing, China, 2615–2623. doi:10.1145/3357384.3357814
arXiv 2019
-
[30]
Honghao Li, Yiwen Zhang, Yi Zhang, Hanwei Li, Lei Sang, and Jieming Zhu
-
[31]
Honghao Li, Yiwen Zhang, Yi Zhang, Lei Sang, and Jieming Zhu. 2025. Revisiting Feature Interactions from The Perspective of Quadratic Neural Networks for Click-through Rate Prediction. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 1365–1375
2025
-
[32]
Kaiyuan Li, Yongxiang Tang, Wenzheng Shu, Yanxiang Zeng, Chao Wang, Yanhua Cheng, Xialong Liu, and Peng Jiang. 2025. Aggregate And Broadcast: Scalable And Efficient Feature Interaction for Recommender Systems.arXiv preprint arXiv:2508.11565(2025)
arXiv 2025
-
[33]
Lichi Li, Zainul Abi Din, Zhen Tan, Sam London, Tianlong Chen, and Ajay Dap- tardar. 2024. MerRec: A Large-Scale Multipurpose Mercari Dataset for Consumer- to-Consumer Recommendation Systems.arXiv preprint arXiv:2402.14230(2024)
Pith/arXiv arXiv 2024
-
[34]
Xiaopeng Li, Jingtong Gao, Pengyue Jia, Xiangyu Zhao, Yichao Wang, Wanyu Wang, Yejing Wang, Yuhao Wang, Huifeng Guo, and Ruiming Tang. 2025. Scenario-Wise Rec: A Multi-Scenario Recommendation Benchmark. InProceed- ings of the 34th ACM International Conference on Information and Knowledge Management. ACM, 1685–1695. doi:10.1145/3746252.3761137
arXiv 2025
-
[35]
Zekun Li, Zeyu Cui, Shu Wu, Xiaoyu Zhang, and Liang Wang. 2019. FiGNN: Modeling Feature Interactions via Graph Neural Networks for CTR Prediction. InProceedings of the 28th ACM International Conference on Information and Knowledge Management. 539–548
2019
-
[36]
Zekun Li, Shu Wu, Zeyu Cui, and Xiaoyu Zhang. 2022. GraphFM: Graph Factoriza- tion Machines for Feature Interaction Modeling.arXiv preprint arXiv:2105.11866 (2022)
Pith/arXiv arXiv 2022
-
[37]
Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xDeepFM: Combining Explicit And Implicit Feature Interactions for Recommender Systems. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 1754–1763
2018
-
[38]
Bin Liu, Niannan Xue, Huifeng Guo, Ruiming Tang, Stefanos Zafeiriou, Xiuqiang He, and Zhenguo Li. 2020. AutoGroup: Automatic Feature Grouping for Modelling Explicit High-Order Feature Interactions in CTR Prediction. InProceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. 199–208
2020
-
[39]
Bin Liu, Chenxu Zhu, Guilin Li, Weinan Zhang, Jincai Lai, Ruiming Tang, Xi- uqiang He, Zhenguo Li, and Yong Yu. 2020. AutoFIS: Automatic Feature In- teraction Selection in Factorization Models for Click-through Rate Prediction. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Conference acronym ’XX, June 03–05, 2018, Woodstock, N...
2020
-
[40]
Jingyuan Liu, Jianlin Su, Xingcheng Yao, Zhejun Jiang, Guokun Lai, Yulun Du, Yidao Qin, Weixin Xu, Enzhe Lu, Junjie Yan, et al. 2025. Muon Is Scalable for LLM Training. arXiv:2502.16982 [cs.LG]
Pith/arXiv arXiv 2025
-
[41]
Mingyang Liu, Yong Bai, Zhangming Chan, Sishuo Chen, Xiang-Rong Sheng, Han Zhu, Jian Xu, and Xinyang Chen. 2026. EST: Towards Efficient Scaling Laws in Click-through Rate Prediction via Unified Modeling.arXiv preprint arXiv:2602.10811(2026)
arXiv 2026
-
[42]
Qi Liu, Kai Zheng, Rui Huang, Wuchao Li, Kuo Cai, Yuan Chai, Yanan Niu, Yiqun Hui, Bing Han, Na Mou, Hongning Wang, Wentian Bao, Yunen Yu, Guorui Zhou, Han Li, Yang Song, Defu Lian, and Kun Gai. 2025. RecFlow: An Industrial Full Flow Recommendation Dataset. InThe Thirteenth International Conference on Learning Representations. https://openreview.net/forum...
2025
-
[43]
Ziyin Liu, Zhikang T. Wang, and Masahito Ueda. 2020. LaProp: Separating Momentum And Adaptivity in Adam. arXiv:2002.04839 [cs.LG]
Pith/arXiv arXiv 2020
-
[44]
Ilya Loshchilov and Frank Hutter. 2017. Decoupled Weight Decay Regularization. arXiv preprint arXiv:1711.05101(2017)
Pith/arXiv arXiv 2017
-
[45]
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. 2018. Modeling Task Relationships in Multi-Task Learning with Multi-Gate Mixture- of-Experts. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 1930–1939
2018
-
[46]
Xiao Ma, Liqin Zhao, Guan Huang, Zhi Wang, Zelin Hu, Xiaoqiang Zhu, and Kun Gai. 2018. Entire Space Multi-Task Model: An Effective Approach for Estimating Post-Click Conversion Rate. InThe 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. 1137–1140
2018
-
[47]
Ze Meng, Jinnian Zhang, Yumeng Li, Jiancheng Li, Tanchao Zhu, and Lifeng Sun
-
[48]
Diganta Misra. 2019. Mish: A Self Regularized Non-Monotonic Activation Func- tion. arXiv:1908.08681 [cs.LG]
Pith/arXiv arXiv 2019
-
[49]
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherni- avskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, Volodymyr Kon- dratenko, Stephanie Pereira, Xianjie Chen, Wenlin Chen, Vijay Rao,...
Pith/arXiv arXiv 2019
-
[50]
Yongxin Ni, Yu Cheng, Xiangyan Liu, Junchen Fu, Youhua Li, Xiangnan He, Yongfeng Zhang, and Fajie Yuan. 2023. A Content-Driven Micro-Video Recom- mendation Dataset at Scale.arXiv preprint arXiv:2309.15379(2023)
Pith/arXiv arXiv 2023
-
[51]
Junwei Pan, Jian Xu, Alfonso Lobos Ruiz, Wenliang Zhao, Shengjun Pan, Yu Sun, and Quan Lu. 2018. Field-Weighted Factorization Machines for Click-through Rate Prediction in Display Advertising. InProceedings of the 2018 World Wide Web Conference. 1349–1357
2018
-
[52]
Junwei Pan, Wei Xue, Ximei Wang, Haibin Yu, Xun Liu, Shijie Quan, Xueming Qiu, Dapeng Liu, Lei Xiao, and Jie Jiang. 2024. Ads Recommendation in A Collapsed And Entangled World. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 5566–5577
2024
-
[53]
Junwei Pan, Wei Xue, Chao Zhou, Xing Zhou, Lunan Fan, Yanbo Wang, Haoran Xin, Zhiyu Hu, Yaozheng Wang, Fengye Xu, Yurong Yang, Xiaotian Li, Junbang Huo, Wentao Ning, Yuliang Sun, Chengguo Yin, Jun Zhang, Shudong Huang, Lei Xiao, Huan Yu, Irwin King, Haijie Gu, and Jie Jiang. 2026. Tencent Advertising Algorithm Challenge 2025: All-Modality Generative Recom...
Pith/arXiv arXiv 2026
-
[54]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al
-
[55]
Changhua Pei, Yi Zhang, Yongfeng Zhang, Fei Sun, Xiao Lin, Hanxiao Sun, Jian Wu, Peng Jiang, Junfeng Ge, Wenwu Ou, and Dan Pei. 2019. Personalized Re- Ranking for Recommendation. InProceedings of the 13th ACM Conference on Recommender Systems. ACM, Copenhagen, Denmark, 3–11. doi:10.1145/3298689. 3347000
doi:10.1145/3298689 2019
-
[56]
Thomas Pethick, Wanyun Xie, Kimon Antonakopoulos, Zhenyu Zhu, Antonio Silveti-Falls, and Volkan Cevher. 2025. Training Deep Learning Models with Norm-Constrained LMOs. arXiv:2502.07529 [cs.LG]
Pith/arXiv arXiv 2025
-
[57]
Aleksandr Poslavsky, Alexander D’yakonov, Yuriy Dorn, and Andrey Zimovnov
-
[58]
Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, Dayiheng Liu, Jingren Zhou, and Junyang Lin. 2025. Gated Attention for Large Language Models: Non-Linearity, Sparsity, And Attention-Sink-Free. Advances in Neural Information Process- ing Systems 38. https://proceedings.neurips.cc/paper_files/pa...
2025
-
[59]
Yanru Qu, Han Cai, Kan Ren, Weinan Zhang, Yong Yu, Ying Wen, and Jun Wang
-
[60]
Steffen Rendle. 2010. Factorization Machines. In2010 IEEE International Confer- ence on Data Mining. IEEE, 995–1000
2010
-
[61]
Matthew Richardson, Ewa Dominowska, and Robert Ragno. 2007. Predicting Clicks: Estimating The Click-through Rate for New Ads. InProceedings of the 16th International Conference on World Wide Web. 521–530
2007
-
[62]
VK-LSVD: A Large-Scale Industrial Dataset for Short-Video Recommenda- tion.arXiv preprint arXiv:2602.04567(2026)
arXiv 2026
-
[63]
Weichen Shen. 2017. DeepCTR: Easy-to-Use, Modular And Extendible Package of Deep-Learning Based CTR Models. https://github.com/shenweichen/deepctr
2017
-
[64]
Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. AutoInt: Automatic Feature Interaction Learning via Self- Attentive Neural Networks. InProceedings of the 28th ACM International Confer- ence on Information and Knowledge Management. 1161–1170
2019
-
[65]
In2016 IEEE 16th International Conference on Data Mining (ICDM)
Product-Based Neural Networks for User Response Prediction. In2016 IEEE 16th International Conference on Data Mining (ICDM). IEEE, 1149–1154
-
[66]
Yang Sun, Junwei Pan, Alex Zhang, and Aaron Flores. 2021. FM2: Field-Matrixed Factorization Machines for Recommender Systems. InProceedings of the Web Conference 2021. 2828–2837
2021
-
[67]
Tijmen Tieleman and Geoffrey Hinton. 2012. RMSProp: Divide The Gradient by A Running Average of Its Recent Magnitude. COURSERA: Neural Networks for Machine Learning, Lecture 6.5
2012
-
[68]
Olivier Roy and Martin Vetterli. 2007. The Effective Rank: A Measure of Effective Dimensionality. InProceedings of the 15th European Signal Processing Conference. IEEE, Poznan, Poland, 606–610
2007
-
[69]
Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Lucas Janson, and Sham Kakade. 2025. SOAP: Improving And Stabilizing Sham- poo Using Adam for Language Modeling. International Conference on Learn- ing Representations. https://proceedings.iclr.cc/paper_files/paper/2025/hash/ e988664070e9591f93fdcf605f7dc623-Abstract-Conference.html
2025
-
[70]
Fangye Wang, Guowei Yang, Xiaojiang Zhou, Song Yang, and Pengjie Wang. 2026. Query-Mixed Interest Extraction And Heterogeneous Interaction: A Scalable CTR Model for Industrial Recommender Systems.arXiv preprint arXiv:2602.09387 (2026)
arXiv 2026
-
[71]
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2024. RoFormer: Enhanced Transformer with Rotary Position Embedding. Neurocomputing568 (2024), 127063. doi:10.1016/j.neucom.2023.127063
arXiv 2024
-
[72]
Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi. 2021. DCNv2: Improved Deep & Cross Network And Practical Lessons for Web-Scale Learning to Rank Systems. InProceedings of the Web Conference
2021
-
[73]
Zhe Wang, Liqin Zhao, Biye Jiang, Guorui Zhou, Xiaoqiang Zhu, and Kun Gai
-
[74]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need.Advances in Neural Information Processing Systems30 (2017)
2017
-
[75]
Yuxin Wu and Kaiming He. 2018. Group Normalization. InProceedings of the European Conference on Computer Vision. Springer, Cham, 3–19. doi:10.1007/978- 3-030-01261-8_1
doi:10.1007/978- 2018
-
[76]
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024. Efficient Streaming Language Models with Attention Sinks. International Con- ference on Learning Representations. arXiv:2309.17453 [cs.CL]
Pith/arXiv arXiv 2024
-
[77]
Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & Cross Network for Ad Click Predictions. InProceedings of the ADKDD’17. 1–7
2017
-
[78]
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li. 2015. Empirical Evaluation of Rectified Activations in Convolutional Network.arXiv preprint arXiv:1505.00853 (2015)
Pith/arXiv arXiv 2015
-
[79]
Yantao Yu, Sen Qiao, Lei Shen, Bing Wang, and Xiaoyi Zeng. 2026. Beyond Dense Connectivity: Explicit Sparsity for Scalable Recommendation.arXiv preprint arXiv:2604.08011(2026). Accepted at SIGIR 2026
Pith/arXiv arXiv 2026
-
[80]
Guanghu Yuan, Jieyu Yang, Shujie Li, Mingjie Zhong, Ang Li, Ke Ding, Yong He, Min Yang, Liang Zhang, Xiaolu Zhang, and Linjian Mo. 2024. MMLRec: A Unified Multi-Task And Multi-Scenario Learning Benchmark for Recommenda- tion. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management(Boise, ID, USA)(CIKM ’24). Associati...
arXiv 2024
-
[81]
Bin Wu, Feifan Yang, Zhangming Chan, Yu-Ran Gu, Jiawei Feng, Chao Yi, Xiang- Rong Sheng, Han Zhu, Jian Xu, Mang Ye, and Bo Zheng. 2025. MUSE: A Simple Yet Effective Multimodal Search-Based Framework for Lifelong User Interest Modeling. arXiv:2512.07216 [cs.IR] https://arxiv.org/abs/2512.07216
arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.