REVIEW 5 major objections 4 minor 1 cited by
Learning Multi-Branch Cooperation for Enhanced Click-Through Rate Prediction at Taobao
T0 review · 5 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read MBCnet shows that CTR prediction improves when feature-interaction branches cooperate by teaching each other on the samples where they disagree.
desk verdict Useful industrial CTR paper with a genuinely new cooperation scheme, but the reported gains are not cleanly attributed to that scheme. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the cooperation scheme rather than any single branch. Branch co-teaching uses per-sample BCE loss compared against the threshold -log(0.5) to label each branch strong or weak; on disagreement samples the strong branch's prediction (detached via stop-gradient) is the teaching signal for the weak branch, and the loss is normalized by the number of disagreement pairs. Moderate differentiation regularizes branch latent features with an equivalent transformation constraint: z_i W_ij = z_j and W_ij orthogonal, implemented as a two-term Frobenius loss in Eq. (14). The EFGC branch is the other novel component: it groups feature fields by domain intention (e.g., query image with item image, user profile with item attributes) and crosses only within groups, which the paper says improves memorization of intended interactions while discarding redundant crossings.
What would settle it
Train MBCnet's three branches with the cooperation losses removed but with identical architecture and an equal number of training steps, tuning the fusion layer to match the reported AUC; if the no-cooperation ensemble matches or exceeds AUC 0.7522/0.7642 on the same Pailitao splits, the paper's central claim about branch cooperation is not what drives the gain. Alternatively, log the strong/weak assignment per sample during early epochs: if the identities of strong and weak branches flip chaotically from step to step, the teaching signal is mostly noise.
Extended reading notes
Core claim
The central claim is that the limiting factor in multi-branch CTR networks is not branch architecture but branch isolation: when branches are merely averaged or concatenated, strong branches cannot rescue weak ones on individual samples. MBCnet operationalizes rescue with branch co-teaching: on samples where branch i has low BCE loss and branch j has high BCE loss, i's sigmoid output becomes a soft label for j, with gradients stopped, and the teaching direction is symmetric. To keep branches from either collapsing into identical representations or drifting apart, a moderate differentiation loss enforces an orthogonal transformation between each pair of latent branch features: z_i W_ij ≈ z_j with W_ij orthogonal, so representations stay related but distinct. On the paper's two industrial datasets the full model reaches AUC 0.7522 and 0.7642, and the online deployment at Taobao reports +0.09 CTR point, +1.49% deals, +1.62% GMV.
Load-bearing premise
The method assumes that comparing each branch's per-sample BCE loss with the fixed threshold -log(0.5) reliably tells which branch is strong and which is weak, and that the strong branch's prediction is a safe teaching target for the weak branch during training; the paper itself notes this can be unreliable in early training.
Editorial extensions
If this is right
- If branch co-teaching is the effective ingredient, other CTR architectures with two or more branches can adopt the same disagreement-based teaching loss without changing their base networks.
- The cooperation scheme makes the value of each branch measurable: ablations show removing EFGC, CrossNet, Deep, co-teaching, or moderate differentiation each lowers AUC, so the gains do not come from any single component alone.
- Moderate differentiation offers a tunable middle ground between feature collapse and feature divergence; the paper's hyper-parameter study shows a wide range of alpha and beta still beats baselines.
- The reported online gains translate into business metrics: a 0.09-point CTR increase, 1.49% more deals, and 1.62% more GMV in Taobao image2product search, at 1 ms added latency and 2% GPU utility increase.
Reading between the lines
- Inference: the disagreement threshold of -log(0.5) is tied to a balanced-label prior; on skewed CTR data the threshold may need to be recomputed, and the paper's own acknowledgment that early-training loss measurements are unreliable suggests a warm-up schedule or an adaptive threshold as a natural next test.
- Inference: because the cooperation scheme is architecture-agnostic, it could be transferred to other two-tower or multi-expert recommendation models, or to ranking domains with similar binary-outcome structure, where the same strong-teaches-weak dynamic should hold.
- Inference: the paper's Figure 7 suggests different branches specialize by category; a testable extension is to use the disagreement mask itself as a signal for sample weighting or for dynamic branch selection at serving time.
- Inference: a direct falsification experiment would be to compare MBCnet against the same three branches trained with simple average-pool fusion plus equal total compute; if the gap shrinks below the reported 0.61%/0.72%, the cooperation losses, not the EFGC architecture, are the source of the gain.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MBCnet, a three-branch CTR prediction model combining an EFGC branch with low-rank CrossNet and Deep branches, trained with a cooperation scheme based on branch co-teaching (Eq. 12) and moderate differentiation (Eq. 14). The EFGC branch groups feature fields according to domain knowledge and crosses embeddings only within groups. Branch co-teaching transfers soft labels from branches with per-sample BCE loss below -log(0.5) to branches above this threshold; moderate differentiation regularizes branch latent features through an orthogonal transformation constraint. The model is tested offline on two large Taobao datasets (Pailitao-12month and Pailitao-24month) and online in Taobao's image2product search, with reported AUC gains of 0.61% and 0.72% over the best baselines, an absolute +0.09 point CTR, +1.49% deals, and +1.62% GMV. Core TensorFlow loss code is included in Appendix D.
Significance. If substantiated, the paper addresses a real gap: multi-branch CTR models typically combine branches only by late fusion, whereas MBCnet introduces explicit inter-branch training signals. The authors deserve credit for evaluating on very large industrial datasets, deploying the model in a live A/B test, providing executable loss code, and ablating both cooperation principles. The claimed gains are large by CTR standards. However, the current evidence has load-bearing gaps: the reported improvement percentages do not match the AUC table, the online A/B test reports no statistical uncertainty and conflates the EFGC branch with the cooperation scheme, and the strong/weak branch selector is acknowledged to be unreliable early in training. These issues must be fixed before the contribution can be assessed.
major comments (5)
- [Table 3 / §4.2.1] The reported relative AUC improvements do not match the table values. On Pailitao-12month, MBCnet AUC 0.7522 against DCNv2 0.7461 implies +0.82%, not +0.61%. On Pailitao-24month, against DNN 0.7569 (the best public baseline in that column) implies +0.96%, and against DCNv2 0.7559 implies +1.10%, not +0.72%. Comparing against EFGC (0.7490 and 0.7600) gives +0.43% and +0.55%, so no listed baseline reproduces the claimed numbers. Please correct the percentages and explicitly name the comparison baseline used for each dataset.
- [Table 4 / §4.2.2] The online A/B evaluation reports point estimates with no confidence intervals, significance test, run duration, or sample size. Since the ablation in Table 6 shows that removing the EFGC branch (0.7445) and removing both cooperation losses (0.7443) cause comparable drops on Pailitao-12month, the online comparison of full MBCnet against DCNv2 cannot attribute the observed lift to the cooperation scheme; the added EFGC branch alone may explain it. Please provide significance estimates and add a controlled comparison that isolates the cooperation losses, ideally online or at least offline on the same data, e.g., DCNv2+EFGC without L_BCT and L_MDR versus full MBCnet.
- [§3.4.1 and §5] The strong/weak branch selection uses the fixed threshold -log(0.5) on per-sample BCE, and the authors explicitly write that this 'may produce unreliable loss measurements during the early training stage, potentially limiting the model's learning ability.' The threshold was also tuned (footnote 3) and acts as a free parameter of the method. Because the co-teaching loss Eq. (12) is applied from the first update with no warm-up or reliability weighting, the paper should either address this acknowledged limitation (e.g., with a warm-up schedule or curriculum) or provide sensitivity analysis showing that the results are robust to threshold choices and early-training behavior.
- [§3.4.2, Eq. (14)] The L_MDR loss does not explicitly implement 'moderate differentiation.' It has two terms: ||z_i W_ij - z_j||_F^2 and an orthogonality constraint ||z_i W_ij (W_ij)^T - z_i||_F^2. Once W_ij is approximately orthogonal, the loss is minimized by making z_i W_ij close to z_j; it does not penalize branches from becoming either identical or arbitrarily different. The 'max difference' and 'min difference' variants in Table 5 are heuristic opposites rather than an interpolation of the same objective. Please state precisely how Eq. (14) enforces a moderate level of differentiation, or replace/annotate the loss with a formulation that contains explicit bounds on representation distance.
- [Table 5] The principle-1 variants produce extreme AUCs (0.5772 for 'no discrimination' and 0.4933 for 'weak to strong' on Pailitao-12month), the latter below random. These values suggest training collapse, yet the paper presents no convergence curves for these variants and does not explain the mechanism. Since Table 5 is the main evidence that disagreement-based sample selection is essential, this result needs verification and analysis, or a corrected implementation.
minor comments (4)
- [§3.2] There is a typo: 'e-commence search' should be 'e-commerce search.'
- [§4.5] The statement that 'excluding the EFGC branch leads to a 0.94% decrease in AUC on Pailitao-24month' is not reproducible from Table 6: AUC drops from 0.7642 to 0.7548, which is a decrease of 0.94 percentage points of AUC or about 1.24% relative to the MBCnet value. Please make the relative-versus-absolute distinction consistent throughout.
- [Appendix B.1] Figure 6(b) would be easier to interpret if the axis labels and the branch legend were defined in the caption; currently the reader must infer the mapping from the main text.
- [§4.2.2] The text says 'absolute 0.09 point CTR increase'; Table 4 shows a move from 10.37% to 10.46%. Please state explicitly whether this is 0.09 percentage points and report the relative change as well.
Circularity Check
No significant circularity; MBCnet's central performance claims rest on held-out offline and online A/B evaluations, with only a minor non-load-bearing self-citation.
full rationale
The paper's central claim is empirical: MBCnet outperforms baselines on held-out industrial test sets and in a live Taobao A/B test. The offline AUC/LogLoss results in Table 3, the ablations in Tables 5 and 6, and the online CTR/deals/GMV improvements in Table 4 are all evaluated against external data or controlled model variants, not derived from the method's own definitions. The cooperation scheme's two losses (branch co-teaching and moderate differentiation) are stated as explicit design principles with objectives defined in Eqs. 12 and 14; their contribution is assessed by removing them in ablations rather than assumed. The one self-citation, reference [4], is used to name and motivate the 'equivalent transformation' formulation of moderate differentiation, but the actual constraint is defined in Eq. 13 and tested empirically, so the citation is not load-bearing and does not force the result. The fixed threshold -log(0.5) in Eq. 10 is an experimentally chosen hyperparameter, and the authors acknowledge in Section 5 that loss-based strong/weak identification may be unreliable early in training; this is a stated limitation, not a circular derivation. The apparent inconsistency between the reported AUC percentage gains and the raw AUC values, and the absence of a controlled online comparison isolating the cooperation scheme from the EFGC branch, are experimental-attribution and reporting concerns rather than circularity. No step in the paper's argument reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (4)
- alpha =
0.1
- beta =
0.1
- co-teaching threshold =
-log(0.5)
- EFGC feature group definitions =
7 groups listed in Table 1
assumptions (3)
- domain assumption Per-sample BCE loss relative to threshold -log(0.5) is a reliable indicator of branch competence.
- ad hoc to paper Latent features of different branches can be related by an orthogonal transformation that preserves useful information.
- domain assumption Co-teaching with soft labels from the strong branch improves ensemble generalization.
Cite this review
Pith. "Pith review of Learning Multi-Branch Cooperation for Enhanced Click-Through Rate Prediction at Taobao." pith.science (2026). https://pith.science/paper/B6G6HLLQ
@misc{pith2026241113057,
author = {Pith},
title = {Pith review of: Learning Multi-Branch Cooperation for Enhanced Click-Through Rate Prediction at Taobao},
year = {2026},
howpublished = {\url{https://pith.science/paper/B6G6HLLQ}},
note = {Machine review of arXiv:2411.13057}
}
read the original abstract
Existing click-through rate (CTR) prediction works have studied the role of feature interaction through a variety of techniques. Each interaction technique exhibits its own strength, and solely using one type usually constrains the model's capability to capture the complex feature relationships, especially for industrial data with enormous input feature fields. Recent research shows that effective CTR models often combine an MLP network with a dedicated feature interaction network in a two-parallel structure. However, the interplay and cooperative dynamics between different streams or branches remain under-researched. In this work, we introduce a novel Multi-Branch Cooperation Network (MBCnet) which enables multiple branch networks to collaborate with each other for better complex feature interaction modeling. Specifically, MBCnet consists of three branches: the Extensible Feature Grouping and Crossing (EFGC) branch that promotes the model's memorization ability of specific feature fields, the low rank Cross Net branch and Deep branch to enhance explicit and implicit feature crossing for improved generalization. Among these branches, a novel cooperation scheme is proposed based on two principles: Branch co-teaching and moderate differentiation. Branch co-teaching encourages well-learned branches to support poorly-learned ones on specific training samples. Moderate differentiation advocates branches to maintain a reasonable level of difference in their feature representations on the same inputs. This cooperation strategy improves learning through mutual knowledge sharing and boosts the discovery of diverse feature interactions across branches. Experiments on large-scale industrial datasets and online A/B test at Taobao app demonstrate MBCnet's superior performance, delivering a 0.09 point increase in CTR, 1.49% growth in deals, and 1.62% rise in GMV. Core codes are available online.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 1 Pith paper
-
From Collapse to Stability: A Knowledge-Driven Ensemble Framework for Scaling Up Click-Through Rate Prediction Models
KDEF combines knowledge distillation and deep mutual learning with adaptive exam-score weighting so that CTR ensembles with up to ten sub-networks improve instead of collapse.
Reference graph
Works this paper leans on
-
[1]
Gavin Brown. 2010. Ensemble Learning. Springer US, Boston, MA, 312–320. doi:10.1007/978-0-387-30164-8_252
-
[2]
Xu Chen, Siheng Chen, Jiangchao Yao, Huangjie Zheng, Ya Zhang, and Ivor W. Tsang. 2022. Learning on Attribute-Missing Graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 2 (2022), 740–757. doi:10.1109/TPAMI.2020. 3032189
-
[3]
Xu Chen, Zida Cheng, Jiangchao Yao, Chen Ju, Weilin Huang, Jinsong Lan, Xiaoyi Zeng, and Shuai Xiao. 2024. Enhancing Cross-Domain Click-Through Rate Prediction via Explicit Feature Augmentation. In Companion Proceedings of the ACM on Web Conference 2024 (Singapore, Singapore) (WWW ’24). Association for Computing Machinery, New York, NY, USA, 423–432. doi:...
doi:10.1145/3589335 2024
-
[4]
Xu Chen, Ya Zhang, Ivor W Tsang, Yuangang Pan, and Jingchao Su. 2020. Towards equivalent transformation of user preferences in cross domain recommendation. ACM Transactions on Information Systems (TOIS) (2020)
work page 2020
-
[5]
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al
-
[6]
Shayan Doroudi. 2020. The Bias-Variance Tradeoff: How Data Science Can Inform Educational Debates. AERA Open 6, 4 (2020), 2332858420977208. doi:10.1177/ 2332858420977208 arXiv:https://doi.org/10.1177/2332858420977208
-
[7]
Binzong Geng, Zhaoxin Huan, Xiaolu Zhang, Yong He, Liang Zhang, Fajie Yuan, Jun Zhou, and Linjian Mo. 2024. Breaking the length barrier: Llm-enhanced CTR prediction in long textual user behaviors. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2311–2315
2024
-
[8]
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. DeepFM: A Factorization-Machine Based Neural Network for CTR Prediction. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (Melbourne, Australia) (IJCAI’17). AAAI Press, 1725–1731
2017
Show all 47 references
-
[9]
Tsang, and Masashi Sugiyama
Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor W. Tsang, and Masashi Sugiyama. 2018. Co-teaching: robust training of deep neural networks with extremely noisy labels. In Proceedings of the 32nd International Conference on Neural Information Processing Sys...
2018
-
[10]
Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, and Joaquin Quiñonero Candela. 2014. Practical Lessons from Predicting Clicks on Ads at Facebook. In Proceedings of the Eighth International Workshop on Data...
2014
-
[11]
Jim Hefferon. 2018. Linear algebra third edition. (2018)
2018
-
[12]
Zhaoxin Huan, Ke Ding, Ang Li, Xiaolu Zhang, Xu Min, Yong He, Liang Zhang, Jun Zhou, Linjian Mo, Jinjie Gu, et al . 2024. Exploring Multi-Scenario Multi- Modal CTR Prediction with a Large Scale Dataset. In Proceedings of the 47th International ACM SIGIR Conference on Research ...
2024
-
[13]
Rohit Kumar, Sneha Manjunath Naik, Vani D Naik, Smita Shiralli, Sunil V.G, and Moula Husain. 2015. Predicting clicks: CTR estimation of advertisements using Logistic Regression classifier. In2015 IEEE International Advance Computing Conference (IACC). 1134–1138. doi:10.1109/IA...
2015
-
[14]
Pan Li and Alexander Tuzhilin. 2020. DDTCDR: Deep Dual Transfer Cross Domain Recommendation. International conference on web search and data mining (2020). Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Xu Chen et al
2020
-
[15]
Xiangyang Li, Bo Chen, Lu Hou, and Ruiming Tang. 2023. Ctrl: Connect collabo- rative and language model for ctr prediction. ACM Transactions on Recommender Systems (2023)
2023
-
[16]
Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems (KDD ’18). Association for Computing Machinery, New York, NY, USA, 1754–1763. doi:10.1145/321981...
2018
-
[17]
Jianghao Lin, Bo Chen, Hangyu Wang, Yunjia Xi, Yanru Qu, Xinyi Dai, Kangning Zhang, Ruiming Tang, Yong Yu, and Weinan Zhang. 2024. ClickPrompt: CTR Models are Strong Prompt Generators for Adapting Language Models to CTR Prediction. In Proceedings of the ACM on Web Conference 2...
2024
-
[18]
Xiaoliang Ling, Weiwei Deng, Chen Gu, Hucheng Zhou, Cui Li, and Feng Sun
-
[19]
Xiaolong Liu, Zhichen Zeng, Xiaoyi Liu, Siyang Yuan, Weinan Song, Mengyue Hang, Yiqun Liu, Chaofei Yang, Donghyun Kim, Wen-Yen Chen, et al . 2024. A Collaborative Ensemble Framework for CTR Prediction. arXiv preprint arXiv:2411.13700 (2024)
2024 arXiv
-
[20]
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. 2018. Modeling task relationships in multi-task learning with multi-gate mixture-of- experts. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining . 1930–1939
2018
-
[21]
Kelong Mao, Jieming Zhu, Liangcai Su, Guohao Cai, Yuru Li, and Zhenhua Dong
-
[22]
Ibomoiye Domor Mienye and Yanxia Sun. 2022. A Survey of Ensemble Learning: Concepts, Algorithms, Applications, and Prospects. IEEE Access 10 (2022), 99129– 99149. doi:10.1109/ACCESS.2022.3207287
2022
-
[23]
Aashiq Muhamed, Iman Keivanloo, Sujan Perera, James Mracek, Yi Xu, Qingjun Cui, Santosh Rajagopalan, Belinda Zeng, and Trishul Chilimbi. 2021. CTR-BERT: Cost-effective knowledge distillation for billion-parameter teacher models. In NeurIPS Efficient Natural Language and Speech...
2021
-
[24]
Wentao Ouyang, Xiuwu Zhang, Li Li, Heng Zou, Xin Xing, Zhaojie Liu, and Yanlong Du. 2019. Deep spatio-temporal neural networks for click-through rate prediction. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 2078–2086
2019
-
[25]
Wentao Ouyang, Xiuwu Zhang, Lei Zhao, Jinmei Luo, Yu Zhang, Heng Zou, Zhaojie Liu, and Yanlong Du. 2020. Minet: Mixed interest network for cross- domain click-through rate prediction. InProceedings of the 29th ACM international conference on information & knowledge management ...
2020
-
[26]
Steffen Rendle. 2010. Factorization Machines. In 2010 IEEE International Confer- ence on Data Mining . 995–1000. doi:10.1109/ICDM.2010.127
2010 doi
-
[27]
Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. AutoInt: Automatic Feature Interaction Learning via Self-Attentive Neural Networks. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management ...
2019
-
[28]
Jingchao Su, Xu Chen, Ya Zhang, Siheng Chen, Dan Lv, and Chenyang Li. 2020. Collaborative Adversarial Learning for Relational Learning on Multiple Bipartite Graphs. In 2020 IEEE International Conference on Knowledge Graph (ICKG) . 466–
2020
-
[29]
Hongyan Tang, Junning Liu, Ming Zhao, and Xudong Gong. 2020. Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations. In Fourteenth ACM Conference on Recommender Systems . 269– 278
2020
-
[30]
Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing Data using t-SNE. Journal of Machine Learning Research 9, 86 (2008), 2579–2605. http: //jmlr.org/papers/v9/vandermaaten08a.html
2008
-
[31]
Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & Cross Network for Ad Click Predictions. In Proceedings of the ADKDD’17 (Halifax, NS, Canada) (ADKDD’17). Association for Computing Machinery, New York, NY, USA, Article 12, 7 pages. doi:10.1145/3124749.3124754
2017
-
[32]
Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & cross network for ad click predictions. In Proceedings of the ADKDD’17 . 1–7
2017
-
[33]
Ruoxi Wang, Rakesh Shivanna, Derek Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed Chi. 2021. Dcn v2: Improved deep & cross network and practical lessons for web-scale learning to rank systems. In Proceedings of the Web Conference 2021. 1785–1797
2021
-
[34]
Xinfei Wang. 2020. A Survey of Online Advertising Click-Through Rate Prediction Models. In 2020 IEEE International Conference on Information Technology,Big Data and Artificial Intelligence (ICIBA), Vol. 1. 516–521. doi:10.1109/ICIBA50161.2020. 9277337
2020
-
[35]
Yongquan Yang, Haijun Lv, and Ning Chen. 2022. A Survey on ensemble learning under the era of deep learning. Artificial Intelligence Review 56, 6 (Nov. 2022), 5545–5589. doi:10.1007/s10462-022-10283-5
2022 doi
-
[36]
Yanwu Yang and Panyu Zhai. 2022. Click-through rate prediction in online advertising: A literature review. Information Processing & Management 59, 2 (2022), 102853. doi:10.1016/j.ipm.2021.102853
2022
-
[37]
Feng Yu, Zhaocheng Liu, Qiang Liu, Haoli Zhang, Shu Wu, and Liang Wang
-
[38]
Weinan Zhang, Tianming Du, and Jun Wang. 2016. Deep Learning over Multi- field Categorical Data. In Advances in Information Retrieval , Nicola Ferro, Fabio Crestani, Marie-Francine Moens, Josiane Mothe, Fabrizio Silvestri, Giorgio Maria Di Nunzio, Claudia Hauff, and Gianmaria ...
2016
-
[39]
Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep interest evolution network for click-through rate prediction. In Proceedings of the AAAI conference on artificial intelligence , Vol. 33. 5941–5948
2019
-
[40]
Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep interest network for click-through rate prediction. In Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining ...
2018
-
[41]
Zhi-Hua Zhou. 2025. Ensemble methods: foundations and algorithms . CRC press
2025
-
[42]
EFGC” branch exhibits greater consistency with “CrossNet
Jieming Zhu, Jinyang Liu, Weiqi Li, Jincai Lai, Xiuqiang He, Liang Chen, and Zibin Zheng. 2020. Ensembled CTR Prediction via Knowledge Distillation. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management (Virtual Event, Ireland) (CIKM ’20...
2020
-
[473]
doi:10.1109/ICBK50248.2020.00072
2020
-
[2016]
In Proceedings of the 1st workshop on deep learning for recommender systems
Wide & deep learning for recommender systems. In Proceedings of the 1st workshop on deep learning for recommender systems . 7–10
-
[2017]
In Proceedings of the 26th International Conference on World Wide Web Companion(Perth, Australia) (WWW ’17 Companion)
Model Ensemble for Click Prediction in Bing Search Ads. In Proceedings of the 26th International Conference on World Wide Web Companion(Perth, Australia) (WWW ’17 Companion) . International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE, 689–...
-
[2020]
In Proceedings of the 29th ACM International Conference on Information & Knowledge Management (Virtual Event, Ireland) (CIKM ’20)
Deep Interaction Machine: A Simple but Effective Model for High-order Feature Interactions. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management (Virtual Event, Ireland) (CIKM ’20). Association for Computing Machinery, New York, NY, USA...
-
[2023]
Proceedings of the AAAI Conference on Artificial Intelligence 37, 4 (Jun
FinalMLP: An Enhanced Two-Stream MLP Model for CTR Prediction. Proceedings of the AAAI Conference on Artificial Intelligence 37, 4 (Jun. 2023), 4552–4560. doi:10.1609/aaai.v37i4.25577
2023 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.