REVIEW 5 major objections 5 minor 1 cited by
Fusion Matters: Learning Fusion in Deep Click-through Rate Prediction Models
T0 review · 5 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A CTR model's fusion design can be learned automatically, beating fixed stacked and parallel fusion on three large ad-click datasets.
desk verdict Useful NAS-for-CTR application with a real protocol flaw — architecture selection on the train set and missing variance — that should be fixed before the precise numbers are trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a one-shot differentiable search over a fusion-only component DAG. The graph contains an embedding block, $n$ shallow and $n$ deep interaction components, and an output block; each component may take inputs from any lower-level component, with connection presence decided by a straight-through unit-step function of $\alpha$ and fusion operation expressed as a softmax-weighted mixture over ADD, PROD, CONCAT, and ATT. During selection, Eq. 16 optimizes model parameters $\Theta$ together with $\alpha$ and $\beta$ on the training loss; after selection, $\alpha^*$ is thresholded to binary connections, $\beta^*$ is kept as a soft mixture (OptFusion-Soft) or hard-coded to the argmax operation (OptFusion-Hard), and $\Theta$ is retrained by Eq. 17. This converts an exponentially large discrete fusion search, of size $O(2^{2n}\cdot k^n)$, into a single continuous optimization.
What would settle it
Run OptFusion's selection stage on one split of a dataset, then evaluate the chosen fusion architecture on a disjoint held-out split against a fixed naive fusion design; if the gap disappears or reverses, the reported gains come from selection overfitting rather than from learning better fusions.
Extended reading notes
Core claim
On its own terms, the paper's claim is that fusion—the choice of which components in a deep CTR model feed into which others, and how their outputs are combined—is a learnable architectural dimension, and that learning it end-to-end improves prediction. OptFusion parameterizes every candidate edge of a component DAG with a connection weight $\alpha$ and every candidate fusion operation with a weight $\beta$, relaxes the discrete choices via a straight-through estimator and a softmax, optimizes both together with model weights on the CTR loss during a selection stage, then fixes the selected architecture and retrains the model. In experiments, both the hard and the soft variant outperform all compared baselines on Criteo, Avazu, and KDD12; the reported AUC improvements over the best baseline are 0.0011, 0.0021, and 0.0036, marked statistically significant at $p < 0.05$. The paper further claims that the learned architectures differ across datasets, with parallel-favoring fusion on Criteo and stacked-favoring fusion on Avazu and KDD12, evidence that a single hand-designed fusion choice leaves performance on the table.
Load-bearing premise
Architecture parameters are selected on the same training set later used to retrain the model, so the method assumes that minimizing the training loss during selection finds fusion designs that generalize to new clicks rather than designs that only fit the training data.
Editorial extensions
If this is right
- Existing fixed fusion designs leave measurable accuracy on the table; selecting fusion automatically yields consistent AUC and LogLoss improvements over stacked, parallel, and expert-designed fusion on all three datasets.
- The best fusion design is data-dependent, since OptFusion's search on Criteo converges to a parallel-leaning architecture while Avazu and KDD12 favor stacked structures; models deployed on new data should therefore search rather than reuse a preset fusion.
- A fusion-only search space is more effective and faster than full neural architecture search for this problem: OptFusion beats the broader NAS baselines while keeping total training time lower.
- Jointly learning connections and operations in one pass beats sequential selection, because the choice of fusion operation changes which connections are useful and vice versa.
- The learned fusion transfers across explicit interaction components, so OptFusion can upgrade models built on different shallow blocks without redesigning them.
Reading between the lines
- A natural transfer test the paper does not run: freeze the fusion architecture found on one dataset and measure how much of the AUC gain survives on another; if most of the gain evaporates, the value of fusion search is dataset-specific tuning rather than a universal architectural improvement.
- Because architecture selection is performed on the same training set used for retraining, a held-out evaluation of the selection procedure itself would clarify whether the reported margins reflect genuine fusion quality or selection overfitting; the paper's significance tests compare final models, not selection robustness.
- The same joint connection-and-operation learning idea could apply wherever hand-set fusion is standard, such as multi-modal or multi-task networks; success there would show the principle generalizes beyond CTR feature interactions.
- The case study's finding that ADD and PROD dominate on two of three datasets suggests the practical gain of operation search comes mostly from choosing among parameter-free combiners, which may be a cheaper hypothesis to test than full operation search.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes OptFusion, a framework for automatically learning fusion design in deep CTR prediction models. It defines a search space over fusion connections and fusion operations, relaxes the discrete choices with continuous architecture parameters α and β, uses a straight-through estimator for connection learning, and jointly optimizes model and architecture parameters in a one-shot selection stage, followed by a retraining stage with fixed architecture. Experiments on Criteo, Avazu, and KDD12 compare OptFusion with a range of CTR baselines and NAS methods; the paper reports AUC and LogLoss gains, an efficiency analysis, and ablation studies on fusion operations, selection algorithms, shallow components, and the number of components. Code is released.
Significance. If the reported gains hold, the paper addresses a genuinely underexplored aspect of CTR model design: fusion connections and operations, rather than component architectures. The focused search space is a sensible alternative to broader NAS methods, and the one-shot selection idea is reasonable. Strengths include public code, evaluation on three large-scale datasets, ablation studies that probe the influence of fusion operations and the selection algorithm, and an efficiency analysis. The main concerns are experimental protocol: architecture selection is performed on the training set without a held-out validation split, the statistical significance claims are not fully supported by the reported details, and several directly related fusion-design baselines are omitted.
major comments (5)
- [§3.4.1 (Eq. 16), Algorithm 1] Architecture parameters α and β are optimized on the same training set D that is later used for retraining (Eq. 17) and for the reported test performance, with no held-out validation split used for architecture selection. Since the reported AUC improvements are only 0.0011–0.0036, selection overfitting on the training loss could plausibly account for part or all of the apparent gains. Please use a held-out validation split for architecture selection, or provide repeated data-split or repeated-run evidence that the selected architectures generalize stably; also report the number of runs and variance for the main results.
- [§4.2, Table 3] The text states that “OptFusion, both soft and hard, outperforms all the SOTA baselines over three datasets” with a “significant margin,” but in Table 3 OptFusion-Hard on KDD12 is not marked with an asterisk, so it is not claimed to be statistically significantly better than the best baseline (EDCN). This internal inconsistency affects the central claim. Please either weaken the claim to match the table or provide the missing significance evidence, and specify the t-test protocol: number of runs, whether tests are paired, and the variance of the reported metrics.
- [§4.1.3 (baselines) and §5.1 (related work)] The related work discusses FinalMLP, EulerNet, and MaskNet as fusion-focused CTR models, yet none of these appears in the baseline comparison in Section 4.1.3 or in Table 3. Since these methods are directly relevant to the paper's claim of outperforming “all SOTA baselines,” the claim is stronger than the evidence. Please add these baselines or explicitly restrict the comparison to the listed methods.
- [§4.4.1, Table 4] The text says that “both Soft and Hard methods exhibit significantly superior performance compared to models with fixed fusion operations,” but Table 4 reports no significance tests or variance. Moreover, on Criteo the ADD-only configuration achieves AUC 0.8111 and LogLoss 0.4422, quite close to Hard (0.8108/0.4413) and Soft (0.8113/0.4408), so the fixed-operation ablation does not uniformly show a large gap. Because the experiments use the searched connections with fixed operations, they also do not isolate the benefit of operation selection from connection learning. Please report repeated runs and tests, and temper or support the claim.
- [§4.4.2, Table 5] The evidence for the one-shot joint selection algorithm, which is one of the paper's stated contributions, is thin: the differences over the sequential selection are 0.0004 in AUC on both Criteo and Avazu (0.8113 vs 0.8109 and 0.7938 vs 0.7934), with no significance tests or repeated-run variance. This small unquantified gap does not convincingly establish that the entanglement between connection and operation selection is beneficial. Please provide variance estimates or a significance test, or frame the one-shot advantage as suggestive rather than established.
minor comments (5)
- [§3.1 (search space analysis)] The identity “2×(1+3+···+2n−1)+2(n+1) = 2n^2+2n+1” appears arithmetically inconsistent: the sum of the first n odd numbers is n^2, so the expression evaluates to 2n^2+2n+2. Please check the derivation and the resulting search-space size.
- [Algorithm 1] The stopping criterion “while not converged” is not defined; please specify the convergence condition or report the number of epochs used in the selection stage.
- [Table 3 footnote] The footnote reports a two-sided t-test with p<0.05 but does not state the number of runs, whether the test is paired, or which variance is used; adding these details would make the significance claims verifiable.
- [§4.1.4 (implementation details)] The statement that the optimal learning rate and L2 regularization from the initial training are reused in retraining is useful, but it is unclear whether all baselines receive the same hyperparameter tuning budget; please clarify the tuning protocol for baselines and OptFusion variants.
- [Figure 3] The case-study figure may be difficult to read in print; please consider larger labels or a table summarizing the selected connections and operations for each dataset.
Circularity Check
No circularity found: the paper's claims are empirical benchmark comparisons; training-set architecture selection is a validity concern, not a circular reduction.
full rationale
The paper does not offer a derivation chain in which a prediction is equivalent to its inputs by construction. Its central claim is empirical: OptFusion is compared against SOTA baselines on three public datasets in Table 3, with AUC/Logloss measured after a separate retraining stage (Eq. 17). The architecture parameters are optimized on the training set in Eq. 16 and Algorithm 1, which is a standard-NAS-protocol concern (no held-out validation split) that could inflate small reported gains; however, the final test metrics are not the training objective itself, so this is an overfitting/validity risk rather than a circularity. The text's claim that both hard and soft variants improve 'by a significant margin' is inconsistent with Table 3, where OptFusion-Hard on KDD12 lacks a significance asterisk, but that is a reporting or statistical-validity issue, not a circular step. Self-citations appear in the related work and references, but they are background material and are not used to justify the central result or to import a uniqueness theorem. No step meets the quote-and-reduction bar required to flag circularity; the paper is self-contained against external benchmarks.
Assumptions & free parameters
free parameters (4)
- default number of shallow and deep component pairs n =
3
- alpha initialization =
0.5
- hard connection threshold =
alpha > 0
- per-dataset learning rate and L2 regularization =
searched from defined grids
assumptions (4)
- domain assumption The DAG level ordering constraint in Eq. 10 prevents cycles and keeps components feed-forward.
- ad hoc to paper The straight-through estimator with backward derivative equal to one provides usable gradients for discrete connection selection.
- ad hoc to paper Architecture parameters selected on the training set transfer to a model retrained from scratch.
- domain assumption Public dataset preprocessing and the reported train and test splits produce comparable evaluation.
Cite this review
Pith. "Pith review of Fusion Matters: Learning Fusion in Deep Click-through Rate Prediction Models." pith.science (2026). https://pith.science/paper/7ZOZXGZA
@misc{pith2026241115731,
author = {Pith},
title = {Pith review of: Fusion Matters: Learning Fusion in Deep Click-through Rate Prediction Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/7ZOZXGZA}},
note = {Machine review of arXiv:2411.15731}
}
read the original abstract
The evolution of previous Click-Through Rate (CTR) models has mainly been driven by proposing complex components, whether shallow or deep, that are adept at modeling feature interactions. However, there has been less focus on improving fusion design. Instead, two naive solutions, stacked and parallel fusion, are commonly used. Both solutions rely on pre-determined fusion connections and fixed fusion operations. It has been repetitively observed that changes in fusion design may result in different performances, highlighting the critical role that fusion plays in CTR models. While there have been attempts to refine these basic fusion strategies, these efforts have often been constrained to specific settings or dependent on specific components. Neural architecture search has also been introduced to partially deal with fusion design, but it comes with limitations. The complexity of the search space can lead to inefficient and ineffective results. To bridge this gap, we introduce OptFusion, a method that automates the learning of fusion, encompassing both the connection learning and the operation selection. We have proposed a one-shot learning algorithm tackling these tasks concurrently. Our experiments are conducted over three large-scale datasets. Extensive experiments prove both the effectiveness and efficiency of OptFusion in improving CTR model performance. Our code implementation is available here\url{https://github.com/kexin-kxzhang/OptFusion}.
Figures
Forward citations
Cited by 1 Pith paper
-
DLF: Enhancing Explicit-Implicit Interaction via Dynamic Low-Order-Aware Fusion for CTR Prediction
DLF is a CTR prediction architecture that combines low-rank, high-rank, and implicit interaction blocks with layer-wise attention fusion, reporting state-of-the-art results on Criteo, Avazu, Movielens, and Frappe.
Reference graph
Works this paper leans on
- [1]
-
[2]
Olivier Chapelle, Eren Manavoglu, and Romer Rosales. 2015. Simple and Scalable Response Prediction for Display Advertising. ACM Trans. Intell. Syst. Technol. 5, 4 (dec 2015), 61. https://doi.org/10.1145/2532128
doi:10.1145/2532128 2015
-
[3]
Bo Chen, Yichao Wang, Zhirong Liu, Ruiming Tang, Wei Guo, Hongkun Zheng, Weiwei Yao, Muyu Zhang, and Xiuqiang He. 2021. Enhancing Explicit and Implicit Feature Interactions via Information Sharing for Parallel Deep CTR Models. In CIKM ’21: The 30th ACM International Conference on Information and Knowledge Management . ACM, Queensland, Australia, 3757–3766...
arXiv 2021
-
[4]
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, Rohan Anil, Zakaria Haque, Lichan Hong, Vihan Jain, Xiaobing Liu, and Hemal Shah
-
[5]
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: A Factorization-Machine based Neural Network for CTR Prediction. In 26th International Joint Conference on Artificial Intelligence, IJCAI 2017 . ijcai.org, Melbourne, Australia, 1725–1731. https://doi.org/10.24963/IJCAI.2017/239
-
[6]
Wei Guo, Ruiming Tang, Huifeng Guo, Jianhua Han, Wen Yang, and Yuzhou Zhang. 2019. Order-aware Embedding Neural Network for CTR Prediction. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2019 . ACM, Paris, France, 1121–1124. https://doi.org/10.1145/3331184.3331332
arXiv 2019
-
[7]
Tongwen Huang, Zhiqi Zhang, and Junlin Zhang. 2019. FiBiNET: combining fea- ture importance and bilinear feature interaction for click-through rate prediction. In Proceedings of the 13th ACM Conference on Recommender Systems, RecSys 2019 . ACM, Copenhagen, Denmark, 169–177. https://doi.org/10.1145/3298689.3347043
arXiv 2019
-
[8]
Eric Jang, Shixiang Gu, and Ben Poole. 2017. Categorical Reparameterization with Gumbel-Softmax. In 5th International Conference on Learning Representations, ICLR 2017. OpenReview.net, Toulon, France. https://openreview.net/forum?id= rkE3y85ee
work page 2017
Show all 56 references
-
[9]
Olivier Chapelle Jean-Baptiste Tien, joycenv. 2014. Display Advertising Challenge. https://kaggle.com/competitions/criteo-display-ad-challenge
2014
-
[10]
Joglekar, Cong Li, Mei Chen, Taibai Xu, Xiaoming Wang, Jay K
Manas R. Joglekar, Cong Li, Mei Chen, Taibai Xu, Xiaoming Wang, Jay K. Adams, Pranav Khaitan, Jiahui Liu, and Quoc V. Le. 2020. Neural Input Search for Large Scale Recommendation Models. In KDD ’20: The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . ACM, C...
2020
-
[11]
Farhan Khawar, Xu Hang, Ruiming Tang, Bin Liu, Zhenguo Li, and Xiuqiang He
-
[12]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti- mization. In 3rd International Conference on Learning Representations, ICLR 2015 . OpenReview.net, San Diego, CA, USA. http://arxiv.org/abs/1412.6980
2015 arXiv
-
[13]
Yujun Li, Xing Tang, Bo Chen, Yimin Huang, Ruiming Tang, and Zhenguo Li
-
[14]
Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mi...
2018
-
[15]
Bin Liu, Ruiming Tang, Yingzhi Chen, Jinkai Yu, Huifeng Guo, and Yuzhou Zhang
-
[16]
Bin Liu, Chenxu Zhu, Guilin Li, Weinan Zhang, Jincai Lai, Ruiming Tang, Xi- uqiang He, Zhenguo Li, and Yong Yu. 2020. AutoFIS: Automatic Feature Inter- action Selection in Factorization Models for Click-Through Rate Prediction. In KDD ’20: The 26th ACM SIGKDD Conference on Kno...
2020
-
[17]
Dugang Liu, Chaohua Yang, Xing Tang, Yejing Wang, Fuyuan Lyu, Weihong Luo, Xiuqiang He, Zhong Ming, and Xiangyu Zhao. 2024. MultiFS: Automated Multi-Scenario Feature Selection in Deep Recommender Systems. In Proceedings of the 17th ACM International Conference on Web Search an...
2024
-
[18]
Hanxiao Liu, Karen Simonyan, and Yiming Yang. 2019. DARTS: Differentiable Architecture Search. In 7th International Conference on Learning Representations, ICLR 2019. OpenReview.net, New Orleans, LA, USA. https://openreview.net/ forum?id=S1eYHoC5FX
2019
-
[19]
Renqian Luo, Fei Tian, Tao Qin, Enhong Chen, and Tie-Yan Liu. 2018. Neural Architecture Optimization. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS
2018
-
[20]
Fuyuan Lyu, Xing Tang, Huifeng Guo, Ruiming Tang, Xiuqiang He, Rui Zhang, and Xue Liu. 2022. Memorize, Factorize, or be Naive: Learning Optimal Feature Interaction Methods for CTR Prediction. In 38th IEEE International Conference on Data Engineering, ICDE 2022 . IEEE, Kuala Lu...
2022
-
[21]
Fuyuan Lyu, Xing Tang, Dugang Liu, Liang Chen, Xiuqiang He, and Xue Liu
-
[22]
Fuyuan Lyu, Xing Tang, Dugang Liu, Chen Ma, Weihong Luo, Liang Chen, Xiuqiang He, and Xue (Steve) Liu. 2023. Towards Hybrid-grained Fea- ture Interaction Selection for Deep Sparse Network. In Advances in Neu- ral Information Processing Systems 36: Annual Conference on Neural I...
2023
-
[23]
Fuyuan Lyu, Xing Tang, Hong Zhu, Huifeng Guo, Yingxue Zhang, Ruiming Tang, and Xue Liu. 2022. OptEmbed: Learning Optimal Embedding Table for Click- through Rate Prediction. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management . ACM, Atl...
2022
-
[24]
Kelong Mao, Jieming Zhu, Liangcai Su, Guohao Cai, Yuru Li, and Zhenhua Dong
-
[25]
Ze Meng, Jinnian Zhang, Yumeng Li, Jiancheng Li, Tanchao Zhu, and Lifeng Sun
-
[26]
Yanru Qu, Han Cai, Kan Ren, Weinan Zhang, Yong Yu, Ying Wen, and Jun Wang
-
[27]
Leon- Suematsu, Jie Tan, Quoc V
Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka I. Leon- Suematsu, Jie Tan, Quoc V. Le, and Alexey Kurakin. 2017. Large-Scale Evolution of Image Classifiers. InProceedings of the 34th International Conference on Machine Learning, ICML 2017 (Proceedings of Mach...
2017
-
[28]
In Proceedings of the ACM Web Conference 2023, WWW 2023
Optimizing Feature Set for Click-Through Rate Prediction. In Proceedings of the ACM Web Conference 2023, WWW 2023 . ACM, Austin, TX, USA, 3386–3395. https://doi.org/10.1145/3543507.3583545
2023
-
[29]
Matthew Richardson, Ewa Dominowska, and Robert Ragno. 2007. Predicting clicks: estimating the click-through rate for new ads. In 16th International Con- ference on World Wide Web, WWW 2007 . ACM, Banff, Alberta, Canada, 521–530. https://doi.org/10.1145/1242572.1242643
2007
-
[30]
Michael S. Ryoo, A. J. Piergiovanni, Juhana Kangaspunta, and Anelia Angelova
-
[31]
Qingquan Song, Dehua Cheng, Hanning Zhou, Jiyan Yang, Yuandong Tian, and Xia Hu. 2020. Towards Automated Neural Interaction Discovery for Click- Through Rate Prediction. In KDD ’20: The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . ACM, Virtual Event, CA,...
2020
-
[32]
In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023
FinalMLP: An Enhanced Two-Stream MLP Model for CTR Prediction. In Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023 . AAAI Press, Washington, DC, USA, 4552–4560. https://doi.org/10.1609/AAAI.V37I4.25577
2023 doi
-
[33]
Ruiming Tang, Bo Chen, Yejing Wang, Huifeng Guo, Yong Liu, Wenqi Fan, and Xiangyu Zhao. 2023. AutoML for Deep Recommender Systems: Fundamentals and Advances. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining, WSDM 2023 . ACM, Singapore,...
2023
-
[34]
Zhen Tian, Ting Bai, Wayne Xin Zhao, Ji-Rong Wen, and Zhao Cao. 2023. Euler- Net: Adaptive Feature Interaction Learning via Euler’s Formula for CTR Predic- tion. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval ...
2023
-
[35]
Fangye Wang, Yingxu Wang, Dongsheng Li, Hansu Gu, Tun Lu, Peng Zhang, and Ning Gu. 2022. Enhancing CTR Prediction with Context-Aware Feature Repre- sentation Learning. In SIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieva...
2022
-
[36]
In 2016 IEEE 16th International Conference on Data Mining (ICDM)
Product-Based Neural Networks for User Response Prediction. In 2016 IEEE 16th International Conference on Data Mining (ICDM) . IEEE, Barcelona, Spain, 1149–1154. https://doi.org/10.1109/ICDM.2016.0151
2016
-
[37]
Ruoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed H. Chi. 2021. DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems. In WWW ’21: The Web Conference 2021. ACM / IW3C2, Ljubljana, Slovenia, ...
2021
-
[38]
Steffen Rendle. 2010. Factorization Machines. In ICDM 2010, The 10th IEEE Inter- national Conference on Data Mining . IEEE Computer Society, Sydney, Australia, 995–1000. https://doi.org/10.1109/ICDM.2010.127
2010 doi
-
[39]
Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Yanghan Wang, Fei Sun, Yiming Wu, Yuandong Tian, Peter Vajda, Yangqing Jia, and Kurt Keutzer. 2019. FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search. In IEEE Conference on Computer Vision and ...
2019
-
[40]
Jun Xiao, Hao Ye, Xiangnan He, Hanwang Zhang, Fei Wu, and Tat-Seng Chua
-
[41]
In Computer Vision - ECCV 2020 - 16th European Conference (Lec- ture Notes in Computer Science, Vol
AssembleNet++: Assembling Modality Representations via Attention Connections. In Computer Vision - ECCV 2020 - 16th European Conference (Lec- ture Notes in Computer Science, Vol. 12365) . Springer, Glasgow, UK, 654–671. https://doi.org/10.1007/978-3-030-58565-5_39
2020 doi
-
[42]
Tunhou Zhang, Dehua Cheng, Yuchen He, Zhengxing Chen, Xiaoliang Dai, Liang Xiong, Feng Yan, Hai Li, Yiran Chen, and Wei Wen. 2023. NASRec: Weight Sharing Neural Architecture Search for Recommender Systems. In Proceedings of the ACM Web Conference 2023, WWW 2023 . ACM, Austin, ...
2023
-
[43]
Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. AutoInt: Automatic Feature Interaction Learning via Self- Attentive Neural Networks. In Proceedings of the 28th ACM International Con- ference on Information and Knowledge Manageme...
2019
-
[44]
Barret Zoph and Quoc V. Le. 2017. Neural Architecture Search with Reinforcement Learning. In 5th International Conference on Learning Representations, ICLR 2017 . OpenReview.net, Toulon, France. https://openreview.net/forum?id=r1Ue8Hcxg
2017
-
[47]
Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & Cross Network for Ad Click Predictions. In ADKDD’17 (ADKDD’17) . Association for Computing Machinery, Halifax, NS, Canada, Article 12, 7 pages. https://doi.org/10.1145/ 3124749.3124754
2017
-
[49]
Zhiqiang Wang, Qingyun She, and Junlin Zhang. 2021. MaskNet: Introducing Feature-Wise Multiplication to CTR Ranking Models by Instance-Guided Mask. CoRR abs/2102.07619 (2021). arXiv:2102.07619 https://arxiv.org/abs/2102.07619
2021 arXiv
-
[51]
https://doi.org/10.1109/CVPR.2019.01099
Computer Vision Foundation / IEEE, Long Beach, CA, USA, 10734–10742. https://doi.org/10.1109/CVPR.2019.01099
2019
-
[54]
Kexin Zhang, Yichao Wang, Xiu Li, Ruiming Tang, and Rui Zhang. 2024. IncMSR: An Incremental Learning Approach for Multi-Scenario Recommendation. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining, WSDM 2024. ACM, Merida, Mexico, 939–948. http...
2024
-
[56]
Weinan Zhang, Tianming Du, and Jun Wang. 2016. Deep Learning over Multi- field Categorical Data - - A Case Study on User Response Prediction. InAdvances in Information Retrieval - 38th European Conference on IR Research, ECIR 2016 (Lecture Notes in Computer Science, Vol. 9626)...
2016 doi
-
[2016]
In Proceedings of the 1st Workshop on Deep Learning for Recommender Systems, DLRS@RecSys 2016
Wide & Deep Learning for Recommender Systems. In Proceedings of the 1st Workshop on Deep Learning for Recommender Systems, DLRS@RecSys 2016 . ACM, Boston, MA, USA, 7–10. https://doi.org/10.1145/2988450.2988454
2016
-
[2017]
In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI 2017
Attentional Factorization Machines: Learning the Weight of Feature Inter- actions via Attention Networks. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI 2017 . ijcai.org, Melbourne, Aus- tralia, 3119–3125. https://doi.org/10...
2017 doi
-
[2018]
https://proceedings
Curran Associates, Montreal, Canada, 7827–7838. https://proceedings. neurips.cc/paper/2018/hash/933670f1ac8ba969f32989c312faba75-Abstract.html
2018
-
[2019]
In The World Wide Web Conference, WWW 2019
Feature Generation by Convolutional Neural Network for Click-Through Rate Prediction. In The World Wide Web Conference, WWW 2019 . ACM, San Francisco, CA, USA, 1119–1129. https://doi.org/10.1145/3308558.3313497
2019
-
[2020]
In CIKM ’20: The 29th ACM International Conference on Information and Knowledge Management
AutoFeature: Searching for Feature Interactions and Their Architectures for Click-through Rate Prediction. In CIKM ’20: The 29th ACM International Conference on Information and Knowledge Management . ACM, Ireland, 625–634. https://doi.org/10.1145/3340531.3411912
-
[2021]
In SIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval
A General Method For Automatic Discovery of Powerful Interactions In Click-Through Rate Prediction. In SIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . ACM, Canada, 1298–1307. https://doi.org/10.1145/3404835.3462842 ...
-
[2023]
In Proceedings of the 17th ACM Conference on Recommender Systems, RecSys 2023
AutoOpt: Automatic Hyperparameter Scheduling and Optimization for Deep Click-through Rate Prediction. In Proceedings of the 17th ACM Conference on Recommender Systems, RecSys 2023 . ACM, Singapore, Singapore, 183–194. https://doi.org/10.1145/3604915.3608800
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.