REVIEW 3 major objections 6 minor 77 references
A model-agnostic test-time wrapper that masks low-confidence features and averages multiple inference paths can improve trained CTR models without retraining.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 10:40 UTC pith:WUXJ734N
load-bearing objection Promising test-time wrapper for CTR models, but the masking semantics and the confidence derivation are both shaky; the empirical gains are real but need controls. the 3 major comments →
MATT-CTR: Unleashing a Model-Agnostic Test-Time Paradigm for CTR Prediction with Confidence-Guided Inference Paths
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
MATT's central claim is that the predictive failures of trained CTR models on infrequent feature combinations can be repaired at inference time by exploiting the correlation between a combination's training-set frequency and the model's confidence in it. MATT quantifies confidence with a hierarchical probabilistic hashing scheme that records high-frequency combinations exactly in a min-heap while bounding low-frequency estimates through Chebyshev's inequality, then iteratively samples features proportional to the confidence of the combination they would form, producing K parallel paths. Each path's prediction is computed after masking unselected features, and the paths are aggregated by conf
What carries the argument
The central object is the confidence score H(f, F) for a feature combination, defined as an occurrence-frequency estimate obtained by hierarchical probabilistic hashing. A min-heap pins exact counts for high-frequency combinations; multiple hash tables plus a Chebyshev lower bound give conservative confidence for low-frequency ones. These scores drive a sequential Bernoulli sampling process that constructs instance-specific feature subsets, and the final prediction is a confidence-weighted ensemble over K sampled paths.
Load-bearing premise
The load-bearing assumption is that setting a feature's value to 0 at inference is the same as removing it; trained CTR models never saw zero-masked inputs during training, so zero may instead inject a different learned embedding or real value, not an absent feature.
What would settle it
Run MATT on a CTR model that was trained with zero-padding or masking augmentation, or that natively supports missing features, and compare it with MATT on the same model trained normally. If the gains disappear or reverse when zeros were seen during training, the effect is not removal of low-confidence features but the novelty of zero inputs; alternatively, measure the calibration of model outputs on fully and partially zeroed instances and check whether the aggregated score remains a valid probability.
If this is right
- Any trained CTR model can be improved at inference time without retraining or architecture changes, reducing the cost of model upgrades.
- Test-time compute can be traded for parallel CPU resources rather than added latency, making the approach deployable in real-time ranking systems.
- Rare feature combinations, not just rare individual features, become the actionable target for inference-time intervention.
- The method appears compatible across diverse CTR architectures, including sequential, multi-expert, and neural-architecture-search-derived models.
- Frequency-based confidence offers a calibration-free proxy for model uncertainty in binary prediction tasks.
Where Pith is reading between the lines
- The frequency-as-confidence proxy could transfer to other sparse, high-dimensional binary prediction tasks such as fraud detection, search ranking, or ad bidding, where occurrence counts are cheap to collect.
- The zero-masking assumption deserves scrutiny: if zero corresponds to a learned embedding or real value rather than absence, the observed gains might reflect a different mechanism, such as implicit input perturbation, rather than removal of low-confidence features.
- A direct test would compare zero-masking against models trained with masked or missing-feature support, or against feeding only selected features through an architecture that handles variable-length input, to isolate 'removal' from 'zero-embedding' effects.
- The multi-path aggregation is a form of test-time ensembling; its variance-reduction benefit could be separated from the confidence-guidance benefit by ablating the confidence weights.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MATT, a model-agnostic test-time inference paradigm for CTR prediction. MATT estimates confidence scores for feature combinations as empirical occurrence frequencies in the training set, using a hierarchical probabilistic hashing scheme: exact counts in a min-heap for high-frequency combinations and a hash-table-based lower bound (derived via Chebyshev's inequality) for low-frequency ones. At inference, MATT iteratively samples features into multiple instance-specific 'inference paths' using confidence-based sampling probabilities, then constructs a sparse input by zeroing out unselected features, scores each path with a frozen base CTR model, and aggregates the path scores with confidence-based weights. Offline experiments on Criteo, Avazu, KDD12, and an industrial dataset show consistent AUC/LogLoss improvements over three base models (HSTU, PLE, OptFu), and a seven-day online A/B test reports a 5.3% CTR lift.
Significance. If the mechanism is valid, MATT would be a novel and practically valuable contribution: it is the first CTR-specific test-time scaling method that is model-agnostic, requires no retraining, and yields meaningful offline and online gains. The paper also provides a concrete algorithmic framework and large-scale experiments, including an online deployment, which is a strength. However, the validity of the core mechanism is currently not established: the zero-masking operation in Eq. (14) is not argued to be equivalent to feature removal for trained models, and the confidence-score derivation in §4.1 is mathematically unsupported. These are load-bearing issues because the reported gains and the interpretation of the method depend on them. With additional controls and a corrected or re-framed confidence score, the idea could be salvageable, but as presented the evidence does not yet support the central claims of 'removing low-confidence features' and 'unleashing predictive potential.'
major comments (3)
- [§4.2, Eq. (14)] Setting unselected fields to 0 is not 'masking' or 'removing' them for a trained CTR model. For categorical fields, 0 is an ordinary embedding index (often a real category or learned default); for continuous fields, 0 is a numeric value. The model has learned parameters for these values, not for 'missingness.' The paper provides no evidence that partially-zeroed inputs produce calibrated predictions for the original instance, nor any control experiment (e.g., random zero-masking, a true missing-value encoding, or analysis of which fields are selected). The central claim in §5.3 that MATT 'mitigates the influence of low-confidence features' therefore rests on an unvalidated equivalence. The Table 2 gains and the 5.3% online lift could be an out-of-distribution artifact of injecting zeros rather than a confidence-guided effect. This issue is load-bearing and must be addressed with explicit
- [§4.1, Eqs. (8)–(9)] The derivation of the low-frequency confidence lower bound does not follow from the stated inequalities. Chebyshev's inequality gives P(|X−μ| ≥ k) ≤ σ²/k², which does not directly produce the conditional probability expression in Eq. (9), and the intermediate manipulations involving P(μX−x<k1) and P(μX−x>k2) are not justified. The choice k1 = 1/sqrt(1/σ_X² − α/k2²) appears without a valid algebraic basis, and the final claim 'lower bound probability higher than 1−α' is not a standard confidence statement; it is unclear whether x is the unknown true count or a random variable and what distribution is being used. Since this lower bound is the confidence score that drives path sampling and weighting, the theoretical foundation of the method is not sound. The authors must either supply a correct derivation or explicitly re-frame the score as a heuristic and validate it empirically.
- [§5.3, Table 2] The claim that MATT 'consistently outperforms all baseline models across all four datasets' is not supported with uncertainty quantification. No error bars, standard deviations, or repeated-seed results are reported. Statistical significance (asterisks, p<0.05) is only shown for OptFu+MATT, not for HSTU+MATT or PLE+MATT, although the latter two are also used to support the compatibility claim. Given that AUC differences are on the order of 0.001–0.005, the reader cannot judge whether the improvements are stable. The authors should report means and variances over multiple runs and significance tests for all MATT variants.
minor comments (6)
- [§4.2, Eqs. (10)–(13)] Because a feature that fails a Bernoulli trial remains in the candidate set, it may be sampled and selected at a later step. The text says the process 'converges' to a high-confidence feature set, but there is no analysis of convergence or of the probability of selecting low-confidence features at later steps. Please clarify the stochastic behavior or soften the convergence claim.
- [§5.4] Typo: 'MARR-RME' should be 'MATT-RME' or 'MATT-RMR' as defined in the variant list.
- [§5.6] The statement that 'MATT's overall wall-clock inference time remains equivalent to that of the base model' holds only if all K paths run fully in parallel. Please state this assumption explicitly and report the actual CPU/GPU resource cost per query, given that the paper acknowledges the trade-off only in terms of parallel computing resources.
- [References] Reference [12] is a duplicate of [10]; references [6] and [13] share the same arXiv identifier (2502.18965), which appears incorrect. Please verify all citations.
- [§4.1, Eq. (5)] The symbol n is used both for the number of feature fields (Section 3.1) and for the number of hash table values in X(c_i). Please disambiguate, e.g., use n_c or L_m'.
- [Throughout] The paper uses 'posterior occurrence frequency' to describe empirical training-set counts. This terminology is misleading; a posterior would involve a prior. Please rename to 'empirical frequency' or 'training frequency.'
Circularity Check
No significant circularity: MATT's gains are measured against external AUC/LogLoss/CTR metrics; the frequency-based 'confidence' is a design choice, not a fitted target.
full rationale
MATT's central claim is that adding test-time confidence-guided feature masking improves prediction accuracy of existing CTR backbones. The confidence scores are obtained from training-set occurrence frequencies (§4.1: 'we propose using the posterior occurrence count of a feature combination as its confidence score'). Using those same scores as sampling probabilities (Eq. 10) and aggregation weights (Eq. 16) makes 'confidence-guided' literally mean 'frequency-guided'; this is a naming/design choice, not a prediction derived from the method. Crucially, validation is against external AUC/LogLoss (Table 2) and online CTR (§5.6), metrics not constructed from the frequency estimates; no fitted value is relabeled as a prediction. The paper's self-citations ([2], [62], [71]) are background/baseline references and are not used to justify MATT's mechanism. Concerns about Eq. (14) zero-masking producing out-of-distribution inputs and about the Chebyshev lower-bound algebra (Eqs. 8-9) are correctness/validity issues, not circularity. I therefore find no load-bearing circular step and assign score 0.
Axiom & Free-Parameter Ledger
free parameters (5)
- alpha (α) =
0.05
- Min-heap capacity |Z_m| =
0.1% of combinations with frequency > 10 at each order
- K (number of parallel paths) =
8
- T (number of path iterations) =
10 or 15
- L_m (number of hash tables per order) =
Not reported
axioms (4)
- domain assumption Training-set occurrence frequency of a feature combination is a valid proxy for the model's prediction confidence
- ad hoc to paper Zeroing unselected features is equivalent to removing them for arbitrary trained CTR models
- ad hoc to paper The Chebyshev-based lower bound in Eqs. (8)–(9) is a valid confidence score
- domain assumption Iterative Bernoulli sampling with p_i^t = 0 for all candidates is either avoided or benign
read the original abstract
Recently, a growing body of research has focused on either optimizing CTR model architectures to better model feature interactions or refining training objectives to aid parameter learning, thereby achieving better predictive performance. However, previous efforts have primarily focused on the training phase, largely neglecting opportunities for optimization during the inference phase. Infrequently occurring feature combinations, in particular, can degrade prediction performance, leading to unreliable or low-confidence outputs. To unlock the predictive potential of trained CTR models, we propose a Model-Agnostic Test-Time paradigm (MATT), which leverages the confidence scores of feature combinations to guide the generation of multiple inference paths, thereby mitigating the influence of low-confidence features on the final prediction. Specifically, to quantify the confidence of feature combinations, we introduce a hierarchical probabilistic hashing method to estimate the occurrence frequencies of feature combinations at various orders, which serve as their corresponding confidence scores. Then, using the confidence scores as sampling probabilities, we generate multiple instance-specific inference paths through iterative sampling and subsequently aggregate the prediction scores from multiple paths to conduct robust predictions. Finally, extensive offline experiments and online A/B tests strongly validate the compatibility and effectiveness of MATT across existing CTR models.
Figures
Reference graph
Works this paper leans on
-
[1]
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al
-
[2]
Bo Chen, Yichao Wang, Zhirong Liu, Ruiming Tang, Wei Guo, Hongkun Zheng, Weiwei Yao, Muyu Zhang, Xiuqiang He. 2021. Enhancing Explicit and Implicit Feature Interactions via Information Sharing for Parallel Deep CTR Models. InProceedings of the 30th ACM International Conference on Information and Knowledge Management (CIKM). (Nov. 2021), 3757-3766
2021
-
[3]
Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan. 2024. AlphaMath Almost Zero: Process Supervision without Process. InProceedings of the Advances in Neu- ral Information Processing Systems 38: Annual Conference on Neural Information Processing Systems (NIPS). (Dec. 2024)
2024
-
[4]
Jianxin Chang, Chenbin Zhang, Yiqun Hui, Dewei Leng, Yanan Niu, Yang Song, and Kun Gai. 2023. PEPNet: Parameter and Embedding Personalized Network for Infusing with Personalized Prior Information. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD). (Aug. 2023), 3795-3804
2023
-
[5]
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, Rohan Anil, Zakaria Haque, Lichan Hong, Vihan Jain, Xiaobing Liu, and Hemal Shah
-
[7]
InProceedings of the 1st Workshop on Deep Learning for Recommender Systems (DLRS@RecSys)
Wide & Deep Learning for Recommender Systems. InProceedings of the 1st Workshop on Deep Learning for Recommender Systems (DLRS@RecSys). (Sep. 2016), 7-10
2016
-
[8]
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017. Deepfm: a factorization-machine based neural network for ctr prediction. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (IJCAI). Melbourne, Australia., 2782–2788
2017
-
[9]
Xidong Feng, Ziyu Wan, Muning Wen, Stephen Marcus McAleer, Ying Wen, Weinan Zhang, and Jun Wang. 2024. AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and Training. InProceedings of the 41st International Conference on Machine Learning (ICML). (Jul. 2024)
2024
-
[10]
Xingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen, Jie Jiang, Mingsheng Long
-
[11]
Shijie Geng, Shuchang Liu, Zuohui Fu, Yingqiang Ge, and Yongfeng Zhang. 2022. Recommendation as language processing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5). InProceedings of the 16th ACM Conference on Recommender Systems. 299–315
2022
-
[12]
On the Embedding Collapse when Scaling up Recommendation Models
Xingzhuo Guo, Junwei Pan, Ximei Wang, Baixu Chen, Jie Jiang, and Mingsheng Long. On the Embedding Collapse when Scaling up Recommendation Models
-
[13]
Ruidong Han, Bin Yin, Shangyu Chen, He Jiang, Fei Jiang, Xiang Li, Chi Ma, Mincong Huang, Xiaoguang Li, Chunzhen Jing, Yueming Han, Menglei Zhou, Lei Yu, Chuan Liu, and Wei Lin. 2025. MTGR: Industrial-Scale Generative Recom- mendation Framework in Meituan.arXiv preprint arXiv:2502.18965(2025)
Pith/arXiv arXiv 2025
-
[14]
Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. InProceedings of the thirteenth international conference on artificial intelligence and statistics. 249–256
2010
-
[15]
Wenyue Hua, Shuyuan Xu, Yingqiang Ge, and Yongfeng Zhang. 2023. How to index item ids for recommendation foundation models. InProceedings of the Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). 195–204
2023
-
[16]
InProceedings of the 41th International Conference on Machine Learning (ICLR). (Jul. 2024)
2024
-
[17]
Kaggle. 2015. Avazu Click-Through Rate Prediction. https://www.kaggle.com/c/avazu-ctr-prediction
2015
-
[18]
Tongwen Huang, Zhiqi Zhang, and Junlin Zhang. 2019. FiBiNET: combining fea- ture importance and bilinear feature interaction for click-through rate prediction. InProceedings of ACM Conference on Recommender Systems (RecSys). 169–177
2019
-
[19]
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361(2020)
Pith/arXiv arXiv 2020
-
[20]
Kaggle. 2014. Criteo Display Advertising Challenge. https://www.kaggle.com/c/criteo-display-ad-challenge
2014
-
[21]
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra
-
[22]
Jikun Kang, Xin Zhe Li, Xi Chen, Amirreza Kazemi, Qianyi Sun, Boxing Chen, Dong Li, Xu He, Quan He, Feng Wen, Jianye Hao, and Jun Yao. 2024. MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time.arXiv preprint arXiv: 2405.16265
Pith/arXiv arXiv 2024
-
[23]
Jianxun Lian, Xiaohuan Zhou, Fuzheng Zhang, Zhongxia Chen, Xing Xie, and Guangzhong Sun. 2018. xDeepFM: Combining Explicit and Implicit Feature In- teractions for Recommender Systems. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD). (Aug. 2018), 1754-1763
2018
-
[24]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti- mization. In ICLR
2015
-
[25]
Pengfei Liu, Xipeng Qiu, and Xuanjing Huang. 2017. Adversarial Multi-task Learning for Text Classification. InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics. Vancouver, Canada, 1–10
2017
-
[26]
Yaoyiran Li, Xiang Zhai, Moustafa Alzantot, Keyi Yu, Ivan Vulić, Anna Korhonen, and Mohamed Hammad. 2024. Calrec: Contrastive alignment of generative llms for sequential recommendation. InProceedings of the 18th ACM Conference on Recommender Systems. 422–432
2024
-
[27]
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. Let’s Verify Step by Step. InProceedings of the 12th International Conference on Learning Representations (ICLR). (May. 2024)
2024
-
[28]
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi. 2018. Modeling task relationships in multi-task learning with multi-gate mixture-of- experts. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD). London, UK, 1930–1939
2018
-
[29]
Jinyun Li, Huiwen Zheng, Yuanlin Liu, Minfang Lu, Lixia Wu, and Haoyuan Hu. 2023. ADL: Adaptive Distribution Learning Framework for Multi-Scenario CTR Prediction. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). (Jul. 2023), 1786-1790
2023
-
[30]
Junwei Pan, Wei Xue, Ximei Wang, Haibin Yu, Xun Liu, Shijie Quan, Xueming Qiu, Dapeng Liu, Lei Xiao, and Jie Jiang. 2024. Ads Recommendation in a Collapsed and Entangled World. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD). (Aug. 2024), 5566-5577
2024
-
[31]
Qi Pi, Guorui Zhou, Yujing Zhang, Zhe Wang, Lejian Ren, Ying Fan, Xiaoqiang Zhu, and Kun Gai. 2020. Search-based user interest modeling with lifelong se- quential behavior data for click-through rate prediction. InProceedings of the 29th ACM International Conference on Information & Knowledge Management (CIKM). 2685–2692
2020
-
[32]
Zihan Liu, Yupeng Hou, and Julian McAuley. 2024. Multi-behavior generative recommendation. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM). 1575–1585
2024
-
[33]
Ruihong Qiu, Zi Huang, Hongzhi Yin, and Zijian Wang. 2022. Contrastive Learn- ing for Representation Degeneration Problem in Sequential Recommendation. InProceedings of the Fifteenth ACM International Conference on Web Search and Data Mining (WSDM). (Feb. 2022), 813-823
2022
-
[34]
Junwei Pan, Jian Xu, Alfonso Lobos Ruiz, Wenliang Zhao, Shengjun Pan, Yu Sun, and Quan Lu. 2018. Field-weighted Factorization Machines for Click-Through Rate Prediction in Display Advertising. InProceedings of the 2018 World Wide Web Conference on World Wide Web (WWW). (Apr. 2018), 1349-1357
2018
-
[35]
Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Tran, Jonah Samost, et al
-
[36]
Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J
Avi Singh, John D. Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J. Liu, James Harrison, Jaehoon Lee, Kelvin Xu, Aaron Parisi, Abhishek Kumar, Alex Alemi, Alex Rizkowsky, Azade Nova, Ben Adlam, Bernd Bohnet, Gamaleldin Elsayed, Hanie Sedghi, Igor Mordatch, Isabelle Simpson, Izzeddin Gur, Jasper Snoek, Jeffrey Pennington, Jiri...
2024
-
[37]
William Peebles and Saining Xie. 2023. Scalable diffusion models with transform- ers. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 4195–4205
2023
-
[38]
Pawel Swietojanski, Jinyu Li, and Steve Renals. 2016. Learning hidden unit con- tributions for unsupervised acoustic model adaptation.IEEE/ACM Transactions on Audio, Speech, and Language Processing. 24, 8 (2016), 1450–1463
2016
-
[39]
Steffen Rendle. 2010. Factorization Machines. InProceedings of the 10th IEEE International Conference on Data Mining (ICDM). (Dec. 2020), 995-1000
2010
-
[40]
Weiping Song, Chence Shi, Zhiping Xiao, Zhijian Duan, Yewen Xu, Ming Zhang, and Jian Tang. 2019. AutoInt: Automatic Feature Interaction Learning via SelfAt- tentive Neural Networks. InProceedings of the 28th ACM International Conference on Information and Knowledge Management (CIKM). 1161–1170
2019
-
[41]
Xiang-Rong Sheng, Feifan Yang, Litong Gong, Biao Wang, Zhangming Chan, Yujing Zhang, Yueyao Cheng, Yong-Nan Zhu, Tiezheng Ge, Han Zhu, Yuning Jiang, Jian Xu, Bo Zheng. 2024. Enhancing Taobao Display Advertising with Multimodal Representations: Challenges, Approaches and Insights. InProceed- ings of the 33rd ACM International Conference on Information and ...
2024
-
[42]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.arXiv preprint arXiv:2402.03300
Pith/arXiv arXiv 2024
-
[43]
Badrul Sarwar, George Karypis, Joseph Konstan, and John Riedl. 2001. Item-based collaborative filtering recommendation algorithms. InProceedings of the 10th international conference on World Wide Web. 285–295
2001
-
[44]
Hongyan Tang, Junning Liu, Ming Zhao, and Xudong Gong. 2020. Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations. InFourteenth ACM Conference on Recommender Systems. 269– 278
2020
-
[45]
Qingquan Song, Dehua Cheng, Hanning Zhou, Jiyan Yang, Yuandong Tian, and Xia Hu. 2020. Towards Automated Neural Interaction Discovery for Click- Through Rate Prediction. InProceedings of the 26th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KDD). 945–955
2020
-
[46]
Ye Tian, Baolin Peng, Linfeng Song, Lifeng Jin, Dian Yu, Haitao Mi, and Dong Yu. 2024. Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing. InProceedings of the Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems (NIPS). (Dec. 2024)
2024
-
[47]
Fangye Wang, Hansu Gu, Dongsheng Li, Tun Lu, Peng Zhang, and Ning Gu. 2023. Towards Deeper, Lighter and Interpretable Cross Network for CTR Prediction. InProceedings of the 32nd ACM International Conference on Information and Knowledge Management (CIKM). (Oct. 2023), 2523-2533
2023
-
[48]
Fangye Wang, Yingxu Wang, Dongsheng Li, Hansu Gu, Tun Lu, Peng Zhang, and Ning Gu. 2022. Enhancing CTR Prediction with Context-Aware Feature Representation Learning. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). (Jul. 2022), 343-352
2022
-
[49]
Gemini 1.5: 2024
Gemini Team. Gemini 1.5: 2024. Unlocking Multimodal Understanding Across Millions of Tokens of Context
2024
-
[50]
Hong Wen, Jing Zhang, Yuan Wang, Fuyu Lv, Wentian Bao, Quan Lin, and Keping Yang. 2020. Entire space multi-task modeling via post-click behavior decom- position for conversion rate prediction. InProceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). 2377–2386
2020
-
[51]
Juntao Tan, Shuyuan Xu, Wenyue Hua, Yingqiang Ge, Zelong Li, and Yongfeng Zhang. 2024. Idgenrec: Llm-recsys alignment with textual id learning. InProceed- ings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). 355–364
2024
-
[52]
Ruoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain, Dong Lin, Lichan Hong, and Ed H. Chi. 2021. DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems. InProceedings of the 30th Web Conference (WWW). (Apr. 2021), 1785-1797
2021
-
[53]
Wenjie Wang, Honghui Bao, Xinyu Lin, Jizhi Zhang, Yongqi Li, Fuli Feng, SeeKiong Ng, and Tat-Seng Chua. 2024. Learnable item tokenization for genera- tive recommendation. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM). 2400–2409
2024
-
[54]
Zhiqiang Wang, Qingyun She, Junlin Zhang. 2021. MaskNet: Introducing Feature- Wise Multiplication to CTR Ranking Models by Instance-Guided Mask. InPro- ceedings of DLP-KDD
2021
-
[55]
Hong Wen, Jing Zhang, Fuyu Lv, Wentian Bao, Tianyi Wang, and Zulong Chen
-
[56]
Mingjia Yin, Junwei Pan, Hao Wang, Ximei Wang, Shangyu Zhang, Jie Jiang, Defu Lian, and Enhong Chen. 2025. From Feature Interaction to Feature Generation: A Generative Paradigm of CTR Prediction Models. InProceedings of the 42nd International Conference on Machine Learning (ICML). (Jul. 2025)
2025
-
[57]
Shenghao Yang, Weizhi Ma, Peijie Sun, Qingyao Ai, Yiqun Liu, Mingchen Cai, and Min Zhang. 2024. Sequential Recommendation with Latent Relations based on Large Language Model. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). (Jul. 2024), 335-344
2024
-
[58]
Ruoxi Wang, Bin Fu, Gang Fu, and Mingliang Wang. 2017. Deep & Cross Network for Ad Click Predictions. InProceedings of the ADKDD’17. (Aug. 2017), 12:1-12:7
2017
-
[59]
Bowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen, Wayne Xin Zhao, Ming Chen, and Ji-Rong Wen. 2024. Adapting large language models by integrating collaborative semantics for recommendation. In2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 1435–1448
2024
-
[60]
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. 2022. STaR: Boot- strapping Reasoning With Reasoning. InAdvances in Neural Information Pro- cessing Systems 35: Annual Conference on Neural Information Processing Systems (NIPS). (Nov. 2022)
2022
-
[61]
Jing Zhang and Dacheng Tao. 2021. Empowering Things With Intelligence: A Survey of the Progress, Challenges, and Opportunities in Artificial Intelligence of Things.IEEE Internet of Things Journal. 8(10), 7789–7817
2021
-
[62]
Xiaoxiao Xu, Chen Yang, Qian Yu, Zhiwei Fang, Jiaxing Wang, Chaosheng Fan, Yang He, Changping Peng, Zhangang Lin, and Jingping Shao. 2022. Alleviating Cold-start Problem in CTR Prediction with A Variational Embedding Learning Framework. InProceedings of the ACM Web Conference 2022 (WWW). (Apr. 2022), 27-35
2022
-
[63]
Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He, Yinghai Lu, and Yu Shi. 2024. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations. InProceedings of the 41st International Conference on Machine Learning (ICML). (Jul. 2024)
2024
-
[64]
Kexin Zhang, Fuyuan Lyu, Xing Tang, Dugang Liu, Chen Ma, Kaize Ding, Xi- uqiang He, and Xue Liu. 2025. Fusion Matters: Learning Fusion in Deep Click- through Rate Prediction Models. InProceedings of the Eighteenth ACM Interna- tional Conference on Web Search and Data Mining (WSDM). (Mar. 2025), 744-753
2025
-
[65]
Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Keming Lu, Chuanqi Tan, Chang Zhou, and Jingren Zhou. 2023. Scaling Relationship on Learning Math- ematical Reasoning with Large Language Models.arxiv preprint arXiv:2308.01825
Pith/arXiv arXiv 2023
-
[66]
Weinan Zhang, Jiarui Qin, Wei Guo, Ruiming Tang, and Xiuqiang He. 2021. Deep Learning for Click-Through Rate Estimation. InProceedings of the Thirtieth International Joint Conference on Artificial Intelligence (IJCAI). (Aug. 2021), 4695- 4703
2021
-
[67]
Shengyu Zhang, Lingxiao Yang, Dong Yao, Yujie Lu, Fuli Feng, Zhou Zhao, Tat-seng Chua, and Fei Wu. 2022. Re4: Learning to Re-contrast, Re-attend, Re- construct for Multi-interest Recommendation. InProceedings of the ACM Web Conference 2022 (WWW). (Apr. 2022), 2216-2226
2022
-
[68]
Guorui Zhou, Na Mou, Ying Fan, Qi Pi, Weijie Bian, Chang Zhou, Xiaoqiang Zhu, and Kun Gai. 2019. Deep Interest Evolution Network for Click-Through Rate Prediction. InProceedings of the 31rd AAAI Conference on Artificial Intelligence (AAAI). (Jan. 2019), 5941-5948
2019
-
[69]
Moyu Zhang, Yongxiang Tang, Jinxin Hu, and Yu Zhang. 2024. Scenario-Adaptive Fine-Grained Personalization Network: Tailoring User Behavior Representation to the Scenario Context. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR). (Jul. 2024), 1557-1566
2024
-
[70]
Jie Zhu, Zhifang Fan, Xiaoxie Zhu, Yuchen Jiang, Hangyu Wang, Xintian Han, Haoran Ding, Xinmin Wang, Wenlin Zhao, Zhen Gong, Huizhi Yang, Zheng Chai, Zhe Chen, Yuchao Zheng, Qiwei Chen, Feng Zhang, Xun Zhou, Peng Xu, Xiao Yang, Di Wu, Zuotao Liu. 2025. RankMixer: Scaling Up Ranking Models in Industrial Recommenders.arXiv preprint arXiv:2507.15551(2025)
Pith/arXiv arXiv 2025
-
[71]
Moyu Zhang, Yun Chen, Yujun Jin, Jinxin Hu, and Yu Zhang. 2025. DGenCTR:Towards a Universal Generative Paradigm for Click-Through Rate Prediction via Discrete Diffusion.arXiv preprint arXiv:2508.14500(2025)
Pith/arXiv arXiv 2025
-
[72]
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. 2022. Scaling vision transformers. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12104–12113
2022
-
[76]
Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep interest network for click-through rate prediction. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD). 1059–1068
2018
-
[2016]
In12th USENIX symposium on operating systems design and implementation (OSDI 16)
Tensorflow: A system for large-scale machine learning. In12th USENIX symposium on operating systems design and implementation (OSDI 16)
-
[2021]
InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR)
Hierarchically Modeling Micro and Macro Behaviors via Multi-Task Learn- ing for Conversion Rate Prediction. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR)
-
[2022]
InPro- ceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems (NIPS)
Solving Quantitative Reasoning Problems with Language Models. InPro- ceedings of the Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems (NIPS). (Nov. 2022)
2022
-
[2023]
Recommender systems with generative retrieval. InAdvances in Neural Woodstock ’18, June 03–05, 2018, Woodstock, NY Moyu Zhang, Yun Chen, Yujun Jin, Jinxin Hu, Yu Zhang, and Xiaoyi Zeng Information Processing Systems 36 (2023), 10299–10315
2018
-
[2024]
InProceedings of the 41st International Conference on Machine Learning (ICML)
On the Embedding Collapse when Scaling up Recommendation Models. InProceedings of the 41st International Conference on Machine Learning (ICML). 2024
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.