REVIEW 5 major objections 7 minor 1 cited by
DGenCTR: Towards a Universal Generative Paradigm for Click-Through Rate Prediction via Discrete Diffusion
T0 review · 5 major / 7 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper claims that a discrete-diffusion pre-training stage, which reconstructs masked CTR features including the click label, followed by fine-tuning the same scoring network, improves CTR prediction and exhibits scaling behavior.
desk verdict A genuinely new use of discrete diffusion for CTR with consistent gains, but the 'direct and lossless transfer' claim rests on a mishandled 1/λ weighting in Eq 18. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is an absorbing-state discrete diffusion process over the unordered set of a sample's feature fields plus its click label. Each feature field flips independently toward a [MASK] absorbing state at a field-specific rate; the reverse score simplifies, via the absorbing-state structure, to the conditional distribution of the clean feature given the unmasked features, so a single time-independent scoring network (an HSTU-style sequence transducer) learns all denoising transitions. The label-aware joint modeling, the reparameterized denoising objective with sampled softmax, and the algebraic equivalence of the label-only masked loss to a calibration loss for CTR are what m
What would settle it
Run the same two-stage pipeline on Criteo with the feature-reconstruction terms removed from the pre-training objective, keeping only the label-masked term. If the resulting AUC does not drop below the full DGenCTR model, the joint feature-distribution modeling that the paper credits for the gains is not the operative mechanism. Equivalently, if an identically sized from-scratch SFT run with the same compute matches the full framework, the pre-training stage is not the cause of the reported gains.
Extended reading notes
Core claim
The central claim is that a sample-level generative pre-training objective and the downstream CTR objective can be made so close that the CTR task is a special case of the denoising task. The forward process independently masks each feature field with field-specific noise schedules, driving toward an absorbing [MASK] state; the reverse process is reparameterized so the learned network predicts clean feature values conditional on unmasked features, independent of timestep. Because the click label is included as one of the features, the model learns separate joint distributions for positives and negatives. When only the label is masked, the pre-training loss coincides with a calibration-style
Load-bearing premise
The paper assumes that because the label-masking component of the pre-training loss equals a standard CTR calibration loss (Eq. 18), the whole pre-training objective is directly aligned with CTR prediction, so every parameter can be transferred losslessly; the full objective also contains feature-reconstruction terms and a weighting over mask levels, so the alignment is an empirical assumption rather than a proven identity.
Editorial extensions
If this is right
- If the central claim holds, generative pre-training at sample level is a viable alternative to discriminative CTR training rather than an architectural patch: it changes the training objective, and the same inference-time architecture serves.
- CTR models trained this way can be made larger with consistent returns, unlike the plateau reported for discriminative models; the paper observes a power-law relationship between compute increase and GAUC gain.
- Sequence-generation generative recommenders, when adapted to CTR, lose user-item cross-features and underperform; the sample-level framing avoids that loss by preserving all fields.
- Transfer should be full: transferring only embeddings or only the scoring network loses a measurable part of the gain, so pre-training and fine-tuning share the same network.
- Label-conditioned pre-training matters: without label modeling, transfer degrades; so generative pre-training for CTR should model positive and negative samples separately.
Reading between the lines
- A testable extension the paper leaves implicit is to interpolate between generation and classification by varying which fields are masked: if Eq. 18 is taken literally, the SFT stage is one particular denoising task, so a schedule that mixes label-masking with feature-masking might trade off calibration and representation quality.
- The same two-stage recipe should transfer to neighboring tabular prediction tasks (e.g., conversion rate, churn, fraud) where samples are unordered and cross-features matter; the Malware out-of-domain result is a weak hint in that direction.
- Because gains are measured relative to an HSTU-based discriminative baseline of similar architecture, the paper's comparison isolates the training paradigm; an even cleaner test would control total compute, since generative pre-training costs extra training FLOPs.
- The online report (higher CTR and revenue) suggests production feasibility follows from keeping inference-time architecture unchanged, but those numbers are deployment-specific; a multi-platform replication would show how universal the paradigm is.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DGenCTR, a two-stage CTR training framework. In the first stage, a discrete absorbing-state diffusion model is trained to reconstruct masked feature fields of a CTR sample; the click label is treated as an additional feature so that the model learns separate distributions for positive and negative samples. The scoring network is an HSTU-style transformer and the training loss is a λ-weighted denoising cross-entropy with sampled softmax. In the second stage, all pre-trained parameters are transferred to a logistic-style CTR head and fine-tuned with binary cross-entropy. The authors argue that the label-reconstruction component is mathematically equivalent to a calibration loss and therefore the pre-training is 'directly optimized' for CTR, enabling 'direct and lossless' transfer. They report offline improvements over discriminative and generative baselines on Criteo, Avazu, Malware, and a 913M-impression industrial dataset, an ablation study, a scaling study, and a 10-day online A/B test.
Significance. If the results hold, the framework is a useful empirical contribution: it is a sample-level generative pretraining scheme that preserves cross-features, shows consistent AUC/logloss gains across four datasets, and gives some evidence that generative pretraining can improve with model capacity. The paper's strengths are the breadth of the empirical study, the inclusion of an online deployment, and the careful application of the discrete-diffusion machinery from prior work. The main weakness is that the paper's central theoretical claim — that the pretraining objective is exactly aligned with the CTR objective via Eq. (18) — is not correct as stated; the empirical value of the method does not depend on that equivalence, but the manuscript's framing and the 'universal paradigm' novelty claim do. With the theoretical claim repaired or downgraded and statistical details added, the empirical story is of interest to the CTR/recsys community.
major comments (5)
- [Section 4.2, Eq. (18)] The claimed equivalence between the label-masking component of L_DP and the calibration loss is not established and, as written, is ill-posed. Starting from Eq. (15), the expectation is over X_λ ~ p_λ(·|X_0); when only the label is masked, the outer expectation includes both the event that the label is masked (probability λ) and the event that it is not. The sum over B[M] conditions on the label being masked, so keeping the 1/λ factor inside this conditional expectation double-counts the mask probability; the integral ∫_0^1 (1/λ) dλ diverges unless a cutoff is imposed. A correct derivation must either drop the 1/λ or use a conditional expectation defined without it. Furthermore, even a correct label-only equivalence would not imply that the full pre-training objective is 'directly optimized' for the downstream CTR objective, because Eq. (15) contains feature-reconstruction terms for all
- [Section 4.1.4, Eq. (15)] The change of variables from t to λ(t) = 1 - exp(-σ̄_k(t)) is introduced per feature through σ̄_k(t), but the transformed loss in Eq. (15) uses a single λ outside the sum over k. If the noise schedules differ across feature fields, there is no single λ and the integral is not well-defined. The transition from Eq. (14) to Eq. (15) also appears to drop the H_3 factors and the σ_k(t) weighting without explanation. Please provide a step-by-step derivation or state explicitly that Eq. (15) defines a new weighted training objective rather than an exact change of variables. This matters because Eq. (18) is derived from Eq. (15).
- [Section 5.7, Fig. 4] The scaling-law claim is based on a small number of model sizes (four or five HSTU blocks) on a single industrial dataset. No error bars, repeated seeds, or fitted functional form are reported; the y-axis is a performance gain relative to one baseline, not an absolute testable quantity. A 'clear power-law relationship' with this level of evidence is not established. Please report the number of points, a fitted exponent with uncertainty, goodness of fit, and ideally multiple datasets, or downgrade the claim to 'consistent with scaling behavior in this setting.'
- [Section 5.3 and Table 2] The main empirical claims rest on small AUC differences (e.g., 0.8167 vs. 0.8129 on Criteo). Table 2 reports no variance or confidence intervals, and Table 4, which says 'mean over 5 runs,' does not give standard deviations. The Mann-Whitney U test is mentioned but without details on the unit of analysis, the number of comparisons, or whether the test is paired. Without this information, the word 'conclusive' in the abstract and Section 5.8 is too strong. Please add standard deviations/confidence intervals and a precise significance-testing protocol.
- [Section 5.8] The online A/B test reports large lifts (6.9% cumulative revenue, 5.8% CTR) over 10 days, but omits the basic experimental setup: number of users/impressions in control and treatment, randomization unit, how the baseline was chosen, confidence intervals or a significance test for the lifts, and whether the reported values are within confidence. This is important because the industrial dataset and online test are presented as the strongest evidence for the framework. Please include these details or clearly label the results as directional.
minor comments (7)
- [Throughout] Placeholder elements from the ACM Woodstock template remain (e.g., 'Woodstock ’18', DOI '10.1145/1122445.1122456', copyright line). The reference to 'two tables' in Section 5.3 is confusing; the comparison table is Table 2.
- [Table 2] The header 'Genrative Paradigm' should be 'Generative Paradigm'.
- [Sections 5.6 and 5.7] Section 5.6 heading says '(Q4)' but should be '(RQ4)'. Section 5.7 has the typo 'incresinging'.
- [Table 4] In the 'w/o Fea' row, '0.8143 4375' and '0,5881' contain formatting errors. Also, 'w/o Fea' removes the per-feature score function, not a feature, so the name is misleading.
- [Section 5.6] The final sentence, 'Consequently, we set the number of diffusion steps for our model,' is incomplete — no value is given.
- [Section 5.2] Discriminative models use batch size 2048 while generative models use batch size 96. This asymmetry should be justified, as batch size can affect model quality and the comparison's fairness.
- [Figures 3 and 4] The captions are too terse; the axes and curves are not described, and the main text does not state the number of points in Figure 4 or the fitted curve equation.
Circularity Check
No significant circularity; Eq 18 is an algebraic identity overclaimed as full-objective equivalence, but the core claims rest on independent empirical evaluation.
full rationale
The paper's derivation chain is not circular in the sense of making a prediction that reduces to its inputs by construction. The discrete-diffusion forward/reverse process (Eqs. 3-13) and the denoising-score-entropy objective (Eqs. 14-17) are imported from prior external work (Lou et al.; Ou et al.), and applying an external derivation is not circular. Eq. 18 is a genuine algebraic identity: for a binary label, the softmax over the two label states equals a sigmoid/BCE form; thus the label-reconstruction component of the pretraining loss is the same loss family as the SFT objective. However, the paper's stronger conclusion in Sec. 4.2 that this makes the whole pretraining stage 'directly optimized for the downstream CTR prediction task' and justifies 'direct and lossless transfer' is not established by Eq. 18 alone, because the actual pretraining objective also contains feature-reconstruction terms and a 1/lambda weighting (Eqs. 15, 17). This is an overclaim or missing proof, not a circularity: the transfer benefit is then tested empirically in Table 3, the ablations in Table 4, and the online A/B test, all against independent baselines. The only self-citations are to EDCN [4] and SAFPN [60] in related work as examples of gating mechanisms; they are not load-bearing for the method or the main results. The scaling-law analysis is a descriptive power-law fit to the authors' own measurements, not a first-principles prediction, so it does not create a circular validation. Overall, the central empirical claims are evaluated on external datasets and against strong discriminative and generative baselines, so there is no forced equivalence between the framework's design and its reported outcomes.
Assumptions & free parameters
free parameters (4)
- Per-feature noise schedule sigma_k(t) =
not disclosed; learned or chosen per feature field
- Number of diffusion steps T =
approximately 500
- Number of pre-training epochs =
3
- Generative model batch size =
96
assumptions (5)
- standard math Absorbing discrete diffusion score reparameterization and time-invariance from Lou et al. (2024) and Ou et al. (2025)
- domain assumption Independence of feature fields in the forward masking process
- domain assumption Sampled softmax with in-batch negatives approximates the full softmax
- domain assumption The click label can be modeled as an additional discrete token in the sample and reconstructed through the same scoring network
- domain assumption HSTU is an appropriate shared backbone for both reconstruction and CTR scoring
Cite this review
Pith. "Pith review of DGenCTR: Towards a Universal Generative Paradigm for Click-Through Rate Prediction via Discrete Diffusion." pith.science (2026). https://pith.science/paper/BSIJECJA
@misc{pith2026250814500,
author = {Pith},
title = {Pith review of: DGenCTR: Towards a Universal Generative Paradigm for Click-Through Rate Prediction via Discrete Diffusion},
year = {2026},
howpublished = {\url{https://pith.science/paper/BSIJECJA}},
note = {Machine review of arXiv:2508.14500}
}
read the original abstract
Recent advances in generative models have inspired the field of recommender systems to explore generative approaches, but most existing research focuses on sequence generation, a paradigm ill-suited for click-through rate (CTR) prediction. CTR models critically depend on a large number of cross-features between the target item and the user to estimate the probability of clicking on the item, and discarding these cross-features will significantly impair model performance. Therefore, to harness the ability of generative models to understand data distributions and thereby alleviate the constraints of traditional discriminative models in label-scarce space, diverging from the item-generation paradigm of sequence generation methods, we propose a novel sample-level generation paradigm specifically designed for the CTR task: a two-stage Discrete Diffusion-Based Generative CTR training framework (DGenCTR). This two-stage framework comprises a diffusion-based generative pre-training stage and a CTR-targeted supervised fine-tuning stage for CTR. Finally, extensive offline experiments and online A/B testing conclusively validate the effectiveness of our framework.
Forward citations
Cited by 1 Pith paper
-
MATT-CTR: Unleashing a Model-Agnostic Test-Time Paradigm for CTR Prediction with Confidence-Guided Inference Paths
MATT is a model-agnostic test-time method that estimates feature-combination frequency from training data and uses it to sample and average multiple masked-input CTR predictions.
Reference graph
Works this paper leans on
-
[8]
In Proceedings of the 1st Workshop on Deep Learning for Recommender Systems (DLRS@RecSys)
Wide & Deep Learning for Recommender Systems. In Proceedings of the 1st Workshop on Deep Learning for Recommender Systems (DLRS@RecSys) . (Sep. 2016), 7-10
work page 2016
-
[14]
Xingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen, Jie Jiang, and Mingsheng Long. 2024. On the Embedding Collapse when Scaling up Recommendation Models. In Proceedings of the 41th International Conference on Machine Learning (ICLR). (Jul. 2024)
work page 2024
-
[1]
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al
-
[2]
Keqin Bao, Jizhi Zhang, Wenjie Wang, Yang Zhang, Zhengyi Yang, Yancheng Luo, Chong Chen, Fuli Feng, and Qi Tian. 2023. A bi-step grounding paradigm for large language models in recommendation systems. ACM Transactions on Recommender Systems (2023)
work page 2023
-
[3]
Andrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth, George Deligiannidis, and Arnaud Doucet. 2022. A Continuous Time Framework for Discrete Denoising Models. In Proceedings of the 36th Advances in Neural Infor- mation Processing Systems 35: Annual Conference on Neural Information Processing (NeurIPS). (Nov. 2022)
work page 2022
-
[4]
Bo Chen, Yichao Wang, Zhirong Liu, Ruiming Tang, Wei Guo, Hongkun Zheng, Weiwei Yao, Muyu Zhang, Xiuqiang He. 2021. Enhancing Explicit and Implicit Feature Interactions via Information Sharing for Parallel Deep CTR Models. In Proceedings of the 30th ACM International Conference on Information and Knowledge Management (CIKM). (Nov. 2021), 3757-3766
work page 2021
-
[5]
Jianxin Chang, Chenbin Zhang, Yiqun Hui, Dewei Leng, Yanan Niu, Yang Song, and Kun Gai. 2023. PEPNet: Parameter and Embedding Personalized Network for Infusing with Personalized Prior Information. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . (Aug. 2023), 3795-3804
2023
-
[6]
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, Rohan Anil, Zakaria Haque, Lichan Hong, Vihan Jain, Xiaobing Liu, and Hemal Shah
Show all 70 references
-
[7]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...
2019
-
[9]
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. Deepfm: a factorization-machine based neural network for ctr prediction. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (IJCAI). Melbourne, Australia., 2782–2788
2017
-
[11]
Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics . 249–256
2010
-
[12]
Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang. 2022. Recommendation as language processing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5). In Proceedings of the 16th ACM Conference on Recommender Systems. 299–315
2022
-
[13]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising Diffusion Probabilistic Models. In Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing (NeurIPS)
2020
-
[15]
Tongwen Huang, Zhiqi Zhang, and Junlin Zhang. 2019. FiBiNET: combining fea- ture importance and bilinear feature interaction for click-through rate prediction. In Proceedings of ACM Conference on Recommender Systems (RecSys) . 169–177
2019
-
[16]
Ruidong Han, Bin Yin, Shangyu Chen, He Jiang, Fei Jiang, Xiang Li, Chi Ma, Mincong Huang, Xiaoguang Li, Chunzhen Jing, Yueming Han, Menglei Zhou, Lei Yu, Chuan Liu, and Wei Lin. 2025. MTGR: Industrial-Scale Generative Recom- mendation Framework in Meituan. arXiv preprint arXiv...
2025 arXiv
-
[17]
Jianchao Ji, Zelong Li, Shuyuan Xu, Wenyue Hua, Yingqiang Ge, Juntao Tan, and Yongfeng Zhang. 2024. Genrec: Large language model for generative recommen- dation. In European Conference on Information Retrieval . Springer, 494–502
2024
-
[18]
Wenyue Hua, Shuyuan Xu, Yingqiang Ge, and Yongfeng Zhang. 2023. How to index item ids for recommendation foundation models. In Proceedings of the Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). 195–204
2023
-
[19]
Kaggle. 2015. Avazu Click-Through Rate Prediction. https://www.kaggle.com/c/avazu-ctr-prediction
2015
-
[20]
Kaggle. 2014. Criteo Display Advertising Challenge. https://www.kaggle.com/c/criteo-display-ad-challenge
2014
-
[21]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti- mization. In ICLR
2015
-
[22]
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 (2020)
2020 arXiv
-
[23]
Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xDeepFM: Combining Explicit and Implicit Feature In- teractions for Recommender Systems. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data ...
2018
-
[24]
Aaron Lou, Chenlin Meng, and Stefano Ermon. 2024. Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution. In Proceedings of the 41st International Conference on Machine Learning (ICML) . (Jul. 2024)
2024
-
[25]
Zihan Liu, Yupeng Hou, and Julian McAuley. 2024. Multi-behavior generative recommendation. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM) . 1575–1585
2024
-
[26]
Yaoyiran Li, Xiang Zhai, Moustafa Alzantot, Keyi Yu, Ivan Vulić, Anna Korhonen, and Mohamed Hammad. 2024. Calrec: Contrastive alignment of generative llms for sequential recommendation. In Proceedings of the 18th ACM Conference on Recommender Systems. 422–432
2024
-
[27]
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. 2018. Modeling task relationships in multi-task learning with multi-gate mixture-of- experts. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD) . Lond...
2018
-
[28]
Microsoft. 2019. Microsoft Malware Prediction. https://www.kaggle.com/c/microsoft-malware-prediction
2019
-
[29]
Jingyang Ou, Shen Nie, Kaiwen Xue, Fengqi Zhu, Jiacheng Sun, Zhenguo Li, and Chongxuan Li. 2025. Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data. In Proceedings of the 13th International Conference on Learning Representations (ICLR...
2025
-
[30]
Simon J Mason and Nicholas E Graham. 2002. Areas beneath the relative operat- ing characteristics (ROC) and relative operating levels (ROL) curves: Statistical significance and interpretation. Quarterly Journal of the Royal Meteorological DGenCTR: Towards a Universal Generativ...
2002
-
[31]
Junwei Pan, Wei Xue, Ximei Wang, Haibin Yu, Xun Liu, Shijie Quan, Xueming Qiu, Dapeng Liu, Lei Xiao, and Jie Jiang. 2024. Ads Recommendation in a Collapsed and Entangled World. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . (Aug...
2024
-
[32]
Junwei Pan, Jian Xu, Alfonso Lobos Ruiz, Wenliang Zhao, Shengjun Pan, Yu Sun, and Quan Lu. 2018. Field-weighted Factorization Machines for Click-Through Rate Prediction in Display Advertising. In Proceedings of the 2018 World Wide Web Conference on World Wide Web (WWW) . (Apr....
2018
-
[33]
William Peebles and Saining Xie. 2023. Scalable diffusion models with transform- ers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 4195–4205
2023
-
[34]
Qi Pi, Guorui Zhou, Yujing Zhang, Zhe Wang, Lejian Ren, Ying Fan, Xiaoqiang Zhu, and Kun Gai. 2020. Search-based user interest modeling with lifelong se- quential behavior data for click-through rate prediction. InProceedings of the 29th ACM International Conference on Informa...
2020
-
[35]
Improving Language Understanding by Generative Pre-Training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving Language Understanding by Generative Pre-Training. OpenAI blog, 2018
2018
-
[36]
Yanru Qu, Han Cai, Kan Ren, Weinan Zhang, Yong Yu, Ying Wen, and Jun Wang
-
[37]
In Proceed- ings of the 16th International Conference on Data Mining (ICDM) .(Dec
Product-Based Neural Networks for User Response Prediction. In Proceed- ings of the 16th International Conference on Data Mining (ICDM) .(Dec. 2016), 1149-1154
2016
-
[38]
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Tran, Jonah Samost, et al. 2023. Recommender systems with generative retrieval. Advances in Neural Information Processing Systems 36 (2023) , 10299–10315
2023
-
[39]
Language Models are Unsupervised Multi-Task Learners.OpenAI blog, 1(8):9, 2019
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language Models are Unsupervised Multi-Task Learners.OpenAI blog, 1(8):9, 2019
2019
-
[40]
Steffen Rendle. 2010. Factorization Machines. In Proceedings of the 10th IEEE International Conference on Data Mining (ICDM) . (Dec. 2020), 995-1000
2010
-
[41]
Haoran Sun, Lijun Yu, Bo Dai, Dale Schuurmans, and Hanjun Dai. 2023. Score- based Continuous-time Discrete Diffusion Models. In Proceedings of the 11th International Conference on Learning Representations (ICLR) . (May. 2023)
2023
-
[42]
Badrul Sarwar, George Karypis, Joseph Konstan, and John Riedl. 2001. Item-based collaborative filtering recommendation algorithms. In Proceedings of the 10th international conference on World Wide Web. 285–295
2001
-
[43]
Dario Shariatian, Umut Simsekli, and Alain Oliviero Durmus. 2025. Heavy-Tailed Diffusion with Denoising Levy Probabilistic Models. In Proceedings of the 13th International Conference on Learning Representations (ICLR) . (Apr. 2025)
2025
-
[44]
Juntao Tan, Shuyuan Xu, Wenyue Hua, Yingqiang Ge, Zelong Li, and Yongfeng Zhang. 2024. Idgenrec: Llm-recsys alignment with textual id learning. In Proceed- ings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) . 355–364
2024
-
[45]
Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. AutoInt: Automatic Feature Interaction Learning via SelfAt- tentive Neural Networks. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management...
2019
-
[46]
Xiang-Rong Sheng, Jingyue Gao, Yueyao Cheng, Siran Yang, Shuguang Han, Hongbo Deng, Yuning Jiang, Jian Xu, and Bo Zheng. 2023. Joint Optimization of Ranking and Calibration with Contextualized Hybrid Model. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discover...
2023
-
[47]
Fangye Wang, Yingxu Wang, Dongsheng Li, Hansu Gu, Tun Lu, Peng Zhang, and Ning Gu. 2022. Enhancing CTR Prediction with Context-Aware Feature Representation Learning. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrie...
2022
-
[48]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser and Illia Polosukhin. 2017. Attention is All you Need. In Proceedings of 30th Conference on Advances in Neural Information Processing Systems (NIPS). (Dec. 2017), 5998–6008
2017
-
[49]
Fangye Wang, Hansu Gu, Dongsheng Li, Tun Lu, Peng Zhang, and Ning Gu. 2023. Towards Deeper, Lighter and Interpretable Cross Network for CTR Prediction. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management (CIKM). (Oct. 2023), 2523-2533
2023
-
[50]
Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & Cross Network for Ad Click Predictions. In Proceedings of the ADKDD’17 . (Aug. 2017), 12:1-12:7
2017
-
[51]
Hong Wen, Jing Zhang, Fuyu Lv, Wentian Bao, Tianyi Wang, and Zulong Chen
-
[52]
Wenjie Wang, Honghui Bao, Xinyu Lin, Jizhi Zhang, Yongqi Li, Fuli Feng, SeeKiong Ng, and Tat-Seng Chua. 2024. Learnable item tokenization for genera- tive recommendation. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM) . 2400–2409
2024
-
[53]
Hong Wen, Jing Zhang, Yuan Wang, Fuyu Lv, Wentian Bao, Quan Lin, and Keping Yang. 2020. Entire space multi-task modeling via post-click behavior decom- position for conversion rate prediction. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Develo...
2020
-
[54]
Zhiqiang Wang, Qingyun She, Junlin Zhang. 2021. MaskNet: Introducing Feature- Wise Multiplication to CTR Ranking Models by Instance-Guided Mask. In Pro- ceedings of DLP-KDD
2021
-
[55]
Ruoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed H. Chi. 2021. DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems. In Proceedings of the 30th Web Conference (WWW) . (Apr. 2021), 1785-1797
2021
-
[56]
Xuanhua Yang, Xiaoyu Peng, Penghui Wei, Shaoguo Liu, Liang Wang and Bo Zheng. 2022. AdaSparse: Learning Adaptively Sparse Structures for Multi-Domain Click-Through Rate Prediction. In Proceedings of the 31st ACM International Conference on Information and Knowledge Management ...
2022
-
[57]
Ye Wang, Jiahao Xun, Minjie Hong, Jieming Zhu, Tao Jin, Wang Lin, Haoyuan Li, Linjun Li, Yan Xia, Zhou Zhao, et al. 2024. EAGER: Two-Stream Generative Rec- ommender with Behavior-Semantic Collaboration. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and...
2024
-
[58]
Bowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen, Wayne Xin Zhao, Ming Chen, and Ji-Rong Wen. 2024. Adapting large language models by integrating collaborative semantics for recommendation. In 2024 IEEE 40th International Conference on Data Engineering (ICDE) . IEEE, 1435–1448
2024
-
[59]
Xiaoxiao Xu, Chen Yang, Qian Yu, Zhiwei Fang, Jiaxing Wang, Chaosheng Fan, Yang He, Changping Peng, Zhangang Lin, and Jingping Shao. 2022. Alleviating Cold-start Problem in CTR Prediction with A Variational Embedding Learning Framework. In Proceedings of the ACM Web Conference...
2022
-
[60]
Moyu Zhang, Yongxiang Tang, Jinxin Hu, and Yu Zhang. 2024. Scenario-Adaptive Fine-Grained Personalization Network: Tailoring User Behavior Representation to the Scenario Context. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Infor...
2024
-
[61]
Yuhao Yang, Zhi Ji, Zhaopeng Li, Yi Li, Zhonglin Mo, Yue Ding, Kai Chen, Zijian Zhang, Jie Li, Shuanglong Li, and Lin Liu. 2025. Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense Representations. CoRR abs/2503.02453 (2025)
2025 arXiv
-
[62]
Kexin Zhang, Fuyuan Lyu, Xing Tang, Dugang Liu, Chen Ma, Kaize Ding, Xi- uqiang He, and Xue Liu. 2025. Fusion Matters: Learning Fusion in Deep Click- through Rate Prediction Models. In Proceedings of the Eighteenth ACM Interna- tional Conference on Web Search and Data Mining (...
2025
-
[63]
Jing Zhang and Dacheng Tao. 2021. Empowering Things With Intelligence: A Survey of the Progress, Challenges, and Opportunities in Artificial Intelligence of Things. IEEE Internet of Things Journal . 8(10), 7789–7817
2021
-
[64]
Weinan Zhang, Jiarui Qin, Wei Guo, Ruiming Tang, and Xiuqiang He. 2021. Deep Learning for Click-Through Rate Estimation. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence (IJCAI) . (Aug. 2021), 4695- 4703
2021
-
[65]
Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He, Yinghai Lu, and Yu Shi. 2024. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations. In Proceedings of the 41st I...
2024
-
[66]
Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep Interest Evolution Network for Click-Through Rate Prediction. In Proceedings of the 31rd AAAI Conference on Artificial Intelligence (AAAI). (Jan. 2019), 5941-5948
2019
-
[67]
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. 2022. Scaling vision transformers. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12104–12113
2022
-
[69]
Shengyu Zhang, Lingxiao Yang, Dong Yao, Yujie Lu, Fuli Feng, Zhou Zhao, Tat-seng Chua, and Fei Wu. 2022. Re4: Learning to Re-contrast, Re-attend, Re- construct for Multi-interest Recommendation. In Proceedings of the ACM Web Conference 2022 (WWW) . (Apr. 2022), 2216-2226
2022
-
[71]
Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep interest network for click-through rate prediction. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining ...
2018
-
[2016]
In 12th USENIX symposium on operating systems design and implementation (OSDI 16)
Tensorflow: A system for large-scale machine learning. In 12th USENIX symposium on operating systems design and implementation (OSDI 16)
-
[2021]
In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR)
Hierarchically Modeling Micro and Macro Behaviors via Multi-Task Learn- ing for Conversion Rate Prediction. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR)
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.