REVIEW 4 major objections 3 minor 39 references
WHALE demonstrates that a recommender can scale feature-interaction and sequence modeling together, beating either backbone alone.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 19:14 UTC pith:WTYZ2CHJ
load-bearing objection WHALE is a solid industrial architecture paper that gives a real unification result, but the experiments leave the progressive-exchange claim underdetermined. the 4 major comments →
WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that feature-interaction and sequence modeling should be unified by keeping both branches active at every layer with attention-based exchange. Each WHALE layer runs a Wukong module over non-sequence features, an HSTU module over the behavior history, and a fusion module where Wukong queries attend over HSTU keys and values. The attended evidence is fused by a residual MLP and SwiGLU feed-forward network, and the exchange repeats across layers. The paper reports this design beats Wukong-only and HSTU-only models at equal per-example FLOPs, scales with sequence length, depth, and width, and yields positive online gains with a modest throughput cost.
What carries the argument
The load-bearing mechanism is the attention-based fusion module inside each WHALE layer. It projects the Wukong branch's high-order interaction representation into queries and the HSTU branch's sequence representation into keys and values, so each non-sequence interaction can retrieve its own relevant slice of the user history. A residual Fusion MLP and a pre-norm SwiGLU feed-forward network then blend the attended history with the interaction features. Because the same exchange repeats in every layer, the architecture achieves progressive Wukong–HSTU exchange rather than a one-shot shallow hybrid.
Load-bearing premise
The load-bearing premise is that giving the three models the same per-example computational budget also equalizes their real modeling capacity, so the quality edge comes from the progressive-exchange design and not from hidden extra parameters; if that premise fails, the central argument weakens.
What would settle it
Re-run the architecture comparison at matched parameter counts rather than matched FLOPs; if WHALE's edge over Wukong-only and HSTU-only disappears, the fusion design is not carrying the gain. Alternatively, rerun the online A/B test with a pre-registered primary metric and confidence intervals; if the 0.113% primary-metric lift does not reproduce, the online claim is not established.
If this is right
- Unifying the two paradigms beats scaling either one alone: WHALE outperforms Wukong-only and HSTU-only models at 8, 14, and 32 GFLOPs.
- WHALE continues to gain from longer behavior histories (up to 15k steps), more layers (up to 8), and wider embeddings (up to 512), suggesting the fusion mechanism benefits from all three scaling axes.
- Fusion strategy matters: shallow-hybrid fusion and average-pooling fusion cause normalized-entropy regressions of 0.25% and 0.23%, indicating that attention-based layer-wise exchange carries the quality gain.
- With system optimizations, deployment is practical: training throughput improves 30%, inference throughput improves 22%, and online results show a 0.113% primary-metric lift at a 5% serving-throughput cost.
- The architecture provides a template for jointly scaling sequential and non-sequential modeling in large industrial recommenders under strict serving constraints.
Where Pith is reading between the lines
- A testable extension is to swap in other scalable backbones for Wukong and HSTU; if the progressive-exchange mechanism is the true driver, WHALE's gains should transfer to those pairs.
- Because the fusion computes per-candidate attention over the same user history, the design should amortize history computation across candidates within a request, making the reported system optimizations as important as the architecture itself.
- The shared-key/value and shared-gate variants show quality survives aggressive weight tying, hinting that the fusion can be made even cheaper and perhaps support longer histories at fixed serving cost.
- The online claim depends on unlabeled supporting metrics; a fully pre-registered set of metrics and confidence intervals would make the production result easier to compare with future unified recommenders.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes WHALE, a unified industrial recommendation architecture that stacks layers, each containing a Wukong feature-interaction module, an HSTU sequence module, and an attention-based fusion module in which Wukong outputs query HSTU outputs. The design keeps both backbones active across layers and is claimed to enable progressive Wukong–HSTU cross-branch exchange. The authors report offline NE gains over Wukong-only and HSTU-only baselines at matched FLOPs, favorable scaling with sequence length, depth, and width, ablations showing the fusion components matter, and efficiency optimizations including Triton kernels, shared key–value attention, and a shared-gate SwiGLU FFN. A 14-day online A/B test shows a 0.113% primary-metric improvement with a 5% inference QPS regression, and the model is stated to be deployed in production.
Significance. If the empirical claims hold, the paper provides a practical, large-scale demonstration that a recursive dual-branch architecture with per-layer fusion can combine feature-interaction and sequence-modeling strengths, with notable engineering contributions around serving efficiency. The 80B-sample offline study, the 14-day online A/B, and the fact that the model is deployed are real strengths, as are the module-level ablations. However, the central design principle—progressive cross-branch exchange—is not cleanly isolated by the presented comparisons, because the main scaling baselines lack one input modality entirely and no parameter counts are reported. The online evidence also rests on undefined supporting metrics with no significance reporting. The result is promising but requires additional controls and reporting before the paper's load-bearing claim is fully supported.
major comments (4)
- [§5.2, Fig. 3] The headline comparison does not isolate the progressive-exchange mechanism. Wukong-only sees no behavior sequences and HSTU-only sees no non-sequence features, while WHALE sees both, so its NE advantage may reflect additional input information rather than the fusion design. FLOPs matching does not control parameter count, and no parameter counts are given; the extra fusion projections (W_Q, W_K, W_V, W_F, SwiGLU matrices) add parameters beyond FLOPs parity. Please add a both-modality baseline (e.g., a shallow combination without per-layer attention fusion) to the scaling curves and report parameter counts alongside FLOPs.
- [§5.4, Table 1] The shallow-hybrid baseline is the correct architecture-level control, but its FLOPs and parameter count are not reported, and it is absent from the scaling curves in Fig. 3. If shallow-hybrid is substantially smaller than WHALE, the 0.25% NE regression could be a capacity effect rather than evidence against shallow fusion. Please report compute and parameters for Table 1 configurations and include the shallow-hybrid in the FLOPs-sweep comparison.
- [§5.5, Table 2] The online claim is not fully verifiable. 'Metric 1' and 'Metric 2' are undefined, no confidence intervals or significance tests are reported, and the statement that supporting metrics make the primary lift 'directionally reliable' is not checkable. Given the 0.113% primary gain is about 3.8× the stated 0.03% significance threshold, a single standard error could change the conclusion. Please define all metrics, state the pre-specified primary metric, and report variances or significance intervals.
- [§5.1 and §5.4] No error bars or significance intervals are reported for any offline NE numbers, despite the text noting that a 0.05% NE gain is considered noticeable. Some reported module-level differences (e.g., 0.08% for removing SwiGLU FFN in Table 1) are close to that threshold. At minimum, the main WHALE-versus-baseline and shallow-hybrid comparisons should include uncertainty estimates or a statement of how many seeds/experiments were run.
minor comments (3)
- [§4.1] Shared key–value and shared-gate SwiGLU are each justified as 'does not degrade model quality' without showing data. A one-line ablation would make this claim checkable.
- [§2, Eq. (2)] The relative attention bias rab(p,t) is not defined in this paper; since Eq. (2) is a compact summary of HSTU, a brief definition or exact pointer to the original paper would help.
- [§5.4, Table 1] The table does not state the configuration (depth, width, sequence length) used for the fusion-strategy study. Please specify whether all rows share the same 8-layer/15k/256-d setting as the scaling experiments.
Circularity Check
No significant circularity: WHALE's central claims are empirical comparisons against external baselines, with no fitted input renamed as a prediction and no load-bearing self-referential derivation.
full rationale
This is an empirical architecture paper, not a derivation. WHALE's central claims are that the proposed progressive Wukong-HSTU fusion improves offline NE and online metrics relative to strong baselines. These claims are supported by external comparisons (Wukong-only, HSTU-only, shallow-hybrid, and module ablations) and by online A/B tests. No fitted parameter or subset of data is renamed as a prediction: the scaling curves, ablations, and online metrics are all measured outcomes, not quantities determined by construction from the model definition. The Wukong and HSTU components are adopted from prior published work rather than re-derived, and the paper does not invoke a uniqueness theorem or forbid alternatives by self-citation. The related-work citations to HSTU and Wukong include overlapping authors, but these are peer-reviewed external results and are not used as the sole justification for the paper's main empirical conclusion. The skeptical concerns about capacity matching and omitted parameter counts are internal-validity or confound issues, not circularity: they question whether the comparison isolates the proposed mechanism, not whether the result reduces to its own inputs by definition. Under the stated review rules, no circular step can be exhibited from the paper's equations or citation chain, so the appropriate finding is no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (5)
- WHALE depth =
8 layers (2/4/6/8 studied)
- Embedding width D =
256 (64–512 studied)
- Sequence length cap =
15k (3k–15k studied)
- Number of non-sequence tokens N =
64
- Platform significance thresholds =
0.05% offline NE, 0.03% online primary metric
axioms (5)
- domain assumption Wukong and HSTU modules perform as described in the cited papers [32, 33].
- domain assumption Per-example forward FLOPs is a valid complexity measure for fair architecture comparison.
- domain assumption Single-epoch training on 80B samples is sufficient for all compared models to reach comparable optimization states.
- domain assumption The unlabeled online metrics (Metric 1, Metric 2) are valid supporting evidence for the primary lift.
- domain assumption Offline training logs (80B/4B samples) are representative of the production traffic during the 14-day A/B test.
read the original abstract
As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and cross features; and sequence features from user behavior histories. Wukong and HSTU have emerged as representative scalable backbones for these paths: Wukong scales high-order non-sequence feature-interaction modeling, while HSTU scales long user-behavior sequence modeling. Despite their complementary strengths, practical architectures that combine these two types of feature modeling remain underexplored. We present WHALE, a scalable unified recommendation architecture that jointly models non-sequence and sequence features on top of Wukong and HSTU. Each WHALE layer contains a Wukong module, an HSTU module, and an attention-based fusion module in which Wukong-derived interaction representations query HSTU-derived behavior representations. This design keeps both backbones active throughout the network and enables progressive Wukong-HSTU exchange, allowing high-order feature crosses to repeatedly retrieve fine-grained evidence from long user histories. To make WHALE practical for industrial deployment, we introduce customized Triton kernels and other model-systems co-design techniques to improve training and inference efficiency. On large-scale industrial recommendation data, WHALE achieves consistent gains in offline experiments. Additionally, it delivers positive online gains with a modest serving-throughput trade-off. The method has been deployed in production systems. Overall, WHALE provides a practical example of how these two sources of information can be scalably unified in an industrial recommendation model.
Figures
Reference graph
Works this paper leans on
-
[1]
Newsha Ardalani, Carole-Jean Wu, Zeliang Chen, Bhargav Bhushanam, and A. Aziz. 2022. Understanding Scaling Laws for Recommendation Models.arXiv preprint arXiv:2208.08489(2022)
Pith/arXiv arXiv 2022
-
[2]
Alex Beutel, Paul Covington, Sagar Jain, Can Xu, Jia Li, Vincent Gatto, and Ed H. Chi. 2018. Latent Cross: Making Use of Context in Recurrent Recommender Systems. InProceedings of the Eleventh ACM International Conference on Web Search and Data Mining
2018
-
[3]
Fedor Borisyuk, Mingzhou Zhou, Qingquan Song, Siyu Zhu, Birjodh Tiwana, Ganesh Parameswaran, Siddharth Dangi, Lars Hertel, Qiang Charles Xiao, Xi- aochen Hou, et al . 2024. LiRank: Industrial Large Scale Ranking Models at LinkedIn. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 4804–4815
2024
-
[4]
Zheng Chai, Qin Ren, Xijun Xiao, Huizhi Yang, Bo Han, Sijun Zhang, Di Chen, Hui Lu, Wenlin Zhao, Lele Yu, et al. 2025. LONGER: Scaling Up Long Sequence Modeling in Industrial Recommenders. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 247–256
2025
-
[5]
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al
-
[6]
Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep Neural Networks for YouTube Recommendations. InProceedings of the 10th ACM Conference on Recommender Systems. 191–198
2016
-
[7]
Tri Dao. 2024. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. InInternational Conference on Learning Representations
2024
-
[8]
Fu, Stefano Ermon, Atri Rudra, and Christopher R’e
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R’e. 2022. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. InAdvances in Neural Information Processing Systems
2022
-
[9]
Kazushige Goto and Robert A. van de Geijn. 2008. Anatomy of High-Performance Matrix Multiplication.ACM Trans. Math. Software34, 3 (2008), 1–25. doi:10.1145/ 1356052.1356053
arXiv 2008
-
[10]
Audrunas Gruslys, Remi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves
-
[11]
Huan Gui, Ruoxi Wang, Ke Yin, Long Jin, Maciej Kula, Taibai Xu, Lichan Hong, and Ed H. Chi. 2023. Hiformer: Heterogeneous Feature Interactions Learning with Transformers for Recommender Systems.arXiv preprint arXiv:2311.05884 (2023)
Pith/arXiv arXiv 2023
-
[12]
InAdvances in Neural Information Processing Systems 29
Memory-Efficient Backpropagation Through Time. InAdvances in Neural Information Processing Systems 29
-
[13]
Mingming Ha, Guanchen Wang, Linxun Chen, Xuan Rao, Yuexin Shi, Tianbao Ma, Zhaojie Liu, Yunqian Fan, Zilong Lu, Yanan Niu, Han Li, and Kun Gai. 2026. UniMixer: A Unified Architecture for Scaling Laws in Recommendation Systems. arXiv preprint arXiv:2604.00590(2026)
arXiv 2026
-
[14]
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: A Factorization-Machine based Neural Network for CTR Prediction. arXiv preprint arXiv:1703.04247(2017)
Pith/arXiv arXiv 2017
-
[15]
Bojian Hou, Xiaolong Liu, Xiaoyi Liu, Jiaqi Xu, et al . 2026. Kunlun: Establish- ing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design.arXiv preprint arXiv:2602.10016(2026)
Pith/arXiv arXiv 2026
-
[16]
Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, An- toine Atallah, Stuart Bowers, and Joaquin Quiñonero Candela. 2014. Prac- tical Lessons from Predicting Clicks on Ads at Facebook. InProceedings of the Eighth International Workshop on Data Mining for Online Advertising. 1–9. doi:10.1145/2648584.2648589
arXiv 2014
-
[17]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling Laws for Neural Language Models.arXiv preprint arXiv:2001.08361(2020)
Pith/arXiv arXiv 2020
-
[18]
Yunwen Huang, Shiyong Hong, Xijun Xiao, Jinqiu Jin, Xuanyuan Luo, Zhe Wang, Zheng Chai, Shikang Wu, Yuchao Zheng, and Jingjian Lin. 2026. Hy- Former: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction.arXiv preprint arXiv:2601.12681(2026). arXiv:2601.12681 https://arxiv.org/abs/2601.12681
arXiv 2026
-
[19]
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018. Mixed Precision Training. InInternational Confer- ence on Learning Representations
2018
-
[20]
Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. doi:10.1145/ 3219819.3220023
arXiv 2018
-
[21]
NVIDIA. 2026. Matrix Multiplication Background User’s Guide. NVIDIA Docu- mentation. https://docs.nvidia.com/deeplearning/performance/dl-performance- matrix-multiplication/index.html Accessed May 10, 2026
2026
-
[22]
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherni- avskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, Volodymyr Kon- dratenko, Stephanie Pereira, Xianjie Chen, Wenlin Chen, Vijay Rao,...
Pith/arXiv arXiv 2019
-
[23]
PyTorch Contributors. 2025. torch.nonzero. PyTorch Documentation. https: //docs.pytorch.org/docs/stable/generated/torch.nonzero.html When input is on CUDA, torch.nonzero() causes host-device synchronization. Accessed May 10, 2026
2025
-
[24]
PyTorch Contributors. 2025. torch.cuda.set_sync_debug_mode. PyTorch Docu- mentation. https://docs.pytorch.org/docs/2.9/generated/torch.cuda.set_sync_ debug_mode.html Documents PyTorch’s debug mode for CUDA synchronizing operations. Accessed May 10, 2026
2025
-
[25]
Steffen Rendle. 2010. Factorization Machines. In2010 IEEE International Confer- ence on Data Mining. doi:10.1109/ICDM.2010.127
-
[26]
PyTorch Team. 2026. Torch Compiler: AOTInductor. https://docs.pytorch.org/ docs/2.12/user_guide/torch_compiler/torch.compiler_aot_inductor.html. Ac- cessed: 2026-05-21
2026
-
[27]
Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. AutoInt: Automatic Feature Interaction Learning via Self- Attentive Neural Networks. InProceedings of the 28th ACM International Confer- ence on Information and Knowledge Management. doi:10.1145/3357384.3357925
arXiv 2019
-
[28]
Noam Shazeer. 2020. GLU Variants Improve Transformer.arXiv preprint arXiv:2002.05202(2020)
Pith/arXiv arXiv 2020
-
[29]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. InAdvances in Neural Information Processing Systems. https: //arxiv.org/abs/1706.03762
Pith/arXiv arXiv 2017
-
[30]
Philippe Tillet, H. T. Kung, and David Cox. 2019. Triton: An Intermediate Lan- guage and Compiler for Tiled Neural Network Computations. InProceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Program- ming Languages. Association for Computing Machinery, New York, NY, USA, 10–19
2019
-
[31]
Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed H. Chi. 2021. DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems. InProceedings of the Web Conference 2021. arXiv:2008.13535 doi:10.1145/3442381.3450078
Pith/arXiv arXiv 2021
-
[32]
Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & Cross Network for Ad Click Predictions.arXiv preprint arXiv:1708.05123(2017). doi:10.48550/ arXiv.1708.05123
-
[33]
Buyun Zhang, Liang Luo, Yuxin Chen, Jade Nie, Xi Liu, Shen Li, Yanli Zhao, Yuchen Hao, Yantao Yao, Ellie Dingqiao Wen, Jongsoo Park, Maxim Naumov, and Wenlin Chen. 2024. Wukong: Towards a Scaling Law for Large-Scale Recommen- dation. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235)...
2024
-
[34]
Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He, Yinghai Lu, and Yu Shi. 2024. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Rese...
2024
-
[35]
Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep Interest Evolution Network for Click-Through Rate Prediction. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 5941–5948. doi:10.1609/aaai.v33i01.33015941 Conference’17, July 2017, Washington, DC, USA Cai et al
-
[36]
Zhaoqi Zhang, Haolei Pei, Jun Guo, Tianyu Wang, Yufei Feng, Hui Sun, Shaowei Liu, and Aixin Sun. 2025. OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender.arXiv preprint arXiv:2510.26104(2025)
arXiv 2025
-
[37]
Jie Zhu, Zhifang Fan, Xiaoxie Zhu, Yuchen Jiang, Hangyu Wang, Xintian Han, Haoran Ding, Xinmin Wang, Wenlin Zhao, Zhen Gong, Huizhi Yang, Zheng Chai, Zhe Chen, Yuchao Zheng, Qiwei Chen, Feng Zhang, Xun Zhou, Peng Xu, Xiao Yang, Di Wu, and Zuotao Liu. 2025. RankMixer: Scaling Up Ranking Models in Industrial Recommenders.arXiv preprint arXiv:2507.15551(2025)
Pith/arXiv arXiv 2025
-
[38]
Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep Interest Network for Click- Through Rate Prediction. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. doi:10.1145/3219819.3219823
arXiv 2018
-
[2016]
InProceedings of the 1st Workshop on Deep Learning for Recommender Systems
Wide & Deep Learning for Recommender Systems. InProceedings of the 1st Workshop on Deep Learning for Recommender Systems. 7–10
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.