REVIEW 2 major objections 5 minor 40 references
A training-free news recommender that mixes recency with article similarity can beat or match deep models offline and nearly match them online while running over 600 times faster.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 08:21 UTC pith:K2TFBBCD
load-bearing objection Clean, shippable training-free news baseline that beats neural models offline, nearly matches them online at 600 imes speed, and honestly shows the offline–online and distribution gaps. the 2 major comments →
ZoRRO: A Zero-Weight Personalized Recommender System for Scalable News Recommendation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A simple product of exponential recency and cosine similarity over fixed article embeddings and category vectors is enough to outperform or match trained neural news recommenders on offline ranking metrics and to reach nearly the same online click-through rate, all while remaining training-free and more than 600 times faster at inference.
What carries the argument
The ZoRRO relevance score: for a candidate article c and user history H, r(c) equals the candidate’s intrinsic recency weight times the sum over history of relational similarities each multiplied by the history item’s own recency weight. Relational similarity is just the sum of cosine similarities of pre-computed document embeddings and one-hot category vectors.
Load-bearing premise
Fixed, pre-trained article embeddings plus simple category labels, combined only by cosine similarity and exponential decay, already capture enough of a user’s news interests that no further training on the recommendation objective is needed.
What would settle it
An online A/B test on the same publisher where a neural model that is allowed to update its embeddings on live click data substantially widens the CTR gap over ZoRRO, or where removing either the embedding or the category term collapses ZoRRO’s offline ranking lead.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ZoRRO, a zero-weight, training-free news recommender that scores candidates by combining exponential recency (intrinsic relevance) with cosine similarity over fixed article embeddings and one-hot categories (relational relevance), as formalized in Eqs. (1)–(4). On the EB-NeRD hidden test set it outperforms NRMS, LSTUR, and NPA on AUC/MRR/nDCG (Table 1), while delivering >600× higher inference throughput (Table 3). A six-day live A/B test on Ekstra Bladet shows CTR nearly on par with NRMS (4.19 % vs 4.33 %) and well above a Popular baseline (Table 7). Beyond-accuracy analyses offline and online (Tables 2, 6) further show that models with similar accuracy induce different topical and sentiment distributions. The authors conclude that lightweight, training-free methods remain competitive for real-time news recommendation and that evaluation should go beyond ranking/CTR alone.
Significance. If the reported numbers hold, the work supplies a strong, reproducible practical baseline for large-scale news recommendation: a simple, training-free scorer that matches or exceeds widely used neural models offline, remains competitive online, and is two orders of magnitude faster to serve. The honest documentation of the offline–online gap and the beyond-accuracy distributional differences are themselves useful contributions for the community. Code is released, the method is fully specified, and the online experiment involves hundreds of thousands of users, which together raise the bar for systems papers in this domain.
major comments (2)
- Table 7 reports point estimates of CTR (4.19 % ZoRRO vs 4.33 % NRMS) with only per-method variance and impression counts; no confidence intervals, hypothesis test, or power analysis is given. Because the central online claim is “nearly on par,” a formal statistical comparison (or at least bootstrap CIs) is needed to substantiate that the 0.14 pp gap is negligible rather than a real under-performance.
- §4.3 states that all methods share the same rule-based candidate set of the top-250 popular recent articles. This design choice is reasonable for a controlled A/B test, but it limits the claim that ZoRRO is a complete end-to-end recommender; the method never has to solve retrieval or cold-start candidate generation. The manuscript should either (a) quantify how much of the performance is attributable to the shared candidate pool, or (b) clearly scope the contribution as a re-ranker rather than a full recommender.
minor comments (5)
- §3, Eq. (4): the unweighted sum of cosine similarities is presented without discussion of possible scaling or temperature parameters; a one-sentence justification or ablation would help readers understand why equal weighting is preferred.
- Figure 1 caption and axis labels could more clearly indicate that LSTUR/NPA were evaluated with masked user IDs; the current text buries this detail in the body.
- Table 4 reports mean±std over three runs for embeddings, but Tables 1 and 5 do not; consistency of reporting would strengthen the ranking claims.
- The abstract and introduction claim “more than 600× faster”; Table 3 shows ~667× versus NRMS. Rounding is fine, but citing the exact throughput ratio once would avoid any impression of exaggeration.
- A few typographical inconsistencies appear (e.g., “600 ×” vs “600×”, “0 .5 ms”); a light copy-edit pass would polish the camera-ready version.
Circularity Check
No significant circularity: ZoRRO is an empirical systems method whose scoring rule does not embed the evaluation metrics.
full rationale
The paper proposes a training-free scoring function (Eqs. 1–4) that multiplies exponential recency by a sum of cosine similarities over fixed article embeddings and one-hot categories. Hyper-parameters (decay rates, embedding choice) are selected on a held-out validation split via Optuna and then frozen; the reported offline ranking metrics (Table 1) and online CTR (Table 7) are measured on unseen test data and live traffic. No quantity that is later called a “prediction” or “result” is algebraically identical to a fitted input, and no uniqueness theorem or load-bearing self-citation is invoked to force the form of the model. Self-citations appear only as ordinary prior work on the same dataset and beyond-accuracy metrics; they do not close a definitional loop. The derivation chain is therefore self-contained against external benchmarks, and the circularity score is zero.
Axiom & Free-Parameter Ledger
free parameters (3)
- λc (candidate recency decay) =
0.015
- λh (history recency decay) =
0.0
- article embedding choice =
contrastive (in-domain)
axioms (3)
- ad hoc to paper User interest in a candidate can be approximated by a sum of cosine similarities to historical clicks, re-weighted only by publication recency.
- domain assumption News relevance decays exponentially with hours since publication.
- domain assumption Cosine similarity on fixed text embeddings plus one-hot categories is an adequate relational relevance measure.
invented entities (1)
-
ZoRRO scoring framework
independent evidence
read the original abstract
We present ZoRRO (Zero-Weight Personalized Recommender System), a zero-weight, training-free framework for personalized news recommendation designed for scalable real-world deployment. ZoRRO outperforms strong neural baselines in offline ranking evaluations and achieves click-through rate performance in online A/B testing that is nearly on par with a state-of-the-art deep learning model, while operating more than 600 times faster. Our experiments reveal gaps between offline and online performance and demonstrate that models with similar click-through rate outcomes can produce markedly different recommendation distributions, thereby influencing the overall news flow. These findings position ZoRRO as a practical and efficient solution for large-scale news recommendation and highlight the importance of evaluating recommender systems using metrics beyond accuracy alone.
Figures
Reference graph
Works this paper leans on
-
[1]
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019. Optuna: A Next-generation Hyperparameter Optimization Framework. InProceedings of the 25th ACM SIGKDD International Confer- ence on Knowledge Discovery & Data Mining(Anchorage, AK, USA)(KDD ’19). Association for Computing Machinery, New York, NY, USA, 2623–2631. doi:10.1...
-
[2]
Mingxiao An, Fangzhao Wu, Chuhan Wu, Kun Zhang, Zheng Liu, and Xing Xie. 2019. Neural News Recommendation with Long- and Short-term User Representations. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 336–345. doi:10.18653/v1/P19-1033
-
[3]
Andreas Argyriou, Miguel González-Fierro, and Le Zhang. 2020. Microsoft Recommenders: Best Practices for Production-Ready Recommendation Systems. InCompanion Proceedings of the Web Conference 2020(Taipei, Taiwan)(WWW ’20). Association for Computing Machinery, New York, NY, USA, 50–51. doi:10. 1145/3366424.3382692
arXiv 2020
-
[4]
Keshav Balasubramanian, Abdulla Alshabanah, Joshua D Choe, and Murali Annavaram. 2021. cDLRM: Look Ahead Caching for Scalable Training of Recom- mendation Models. InProceedings of the 15th ACM Conference on Recommender Systems(Amsterdam, Netherlands)(RecSys ’21). Association for Computing Ma- chinery, New York, NY, USA, 263–272. doi:10.1145/3460231.347424...
-
[5]
Das, Mayur Datar, Ashutosh Garg, and Shyam Rajaram
Abhinandan S. Das, Mayur Datar, Ashutosh Garg, and Shyam Rajaram. 2007. Google News Personalization: Scalable Online Collaborative Filtering. InPro- ceedings of the 16th International Conference on World Wide Web(Banff, Alberta, Canada)(WWW ’07). Association for Computing Machinery, New York, NY, USA, 271–280. doi:10.1145/1242572.1242610
-
[6]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL]
Pith/arXiv arXiv 2019
-
[7]
Maurizio Ferrari Dacrema, Paolo Cremonesi, and Dietmar Jannach. 2019. Are we really making much progress? A worrying analysis of recent neural recommen- dation approaches. InProceedings of the 13th ACM Conference on Recommender Systems(Copenhagen, Denmark)(RecSys ’19). Association for Computing Ma- chinery, New York, NY, USA, 101–109. doi:10.1145/3298689.3347058
-
[8]
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2022. SimCSE: Simple Contrastive Learning of Sentence Embeddings. arXiv:2104.08821 [cs.CL]
Pith/arXiv arXiv 2022
-
[9]
Mouzhi Ge, Carla Delgado-Battenfeld, and Dietmar Jannach. 2010. Beyond Accuracy: Evaluating Recommender Systems by Coverage and Serendipity. In Proceedings of the Fourth ACM Conference on Recommender Systems(Barcelona, Spain)(RecSys ’10). Association for Computing Machinery, New York, NY, USA, 257–260. doi:10.1145/1864708.1864761
-
[10]
Udit Gupta, {Carole Jean} Wu, Xiaodong Wang, Maxim Naumov, Brandon Reagen, David Brooks, Bradford Cottel, Kim Hazelwood, Mark Hempstead, Bill Jia, {Hsien Hsin S.} Lee, Andrey Malevich, Dheevatsa Mudigere, Mikhail Smelyanskiy, Liang Xiong, and Xuan Zhang. 2020. The architectural implications of facebook’s DNN- based personalized recommendation. InProceedin...
-
[11]
Andreea Iana, Goran Glavaš, and Heiko Paulheim. 2024. Train Once, Use Flex- ibly: A Modular Framework for Multi-Aspect Neural News Recommendation. InFindings of the Association for Computational Linguistics: EMNLP 2024, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Com- putational Linguistics, Miami, Florida, USA, 9555–9571. do...
doi:10.18653/v1/2024 2024
-
[12]
Marius Kaminskas and Derek Bridge. 2016. Diversity, Serendipity, Novelty, and Coverage: A Survey and Empirical Analysis of Beyond-Accuracy Objectives in Recommender Systems.ACM Trans. Interact. Intell. Syst.7, 1, Article 2 (Dec. 2016), 42 pages. doi:10.1145/2926720
doi:10.1145/2926720 2016
-
[13]
Liu Ke, Udit Gupta, Mark Hempstead, Carole-Jean Wu, Hsien-Hsin S. Lee, and Xuan Zhang. 2022. Hercules: Heterogeneity-Aware Inference Serving for At- Scale Personalized Recommendation. arXiv:2203.07424 [cs.DC] https://arxiv. org/abs/2203.07424
Pith/arXiv arXiv 2022
-
[14]
Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Opti- mization. arXiv:1412.6980 [cs.LG] https://arxiv.org/abs/1412.6980
Pith/arXiv arXiv 2014
-
[15]
Johannes Kruse, Kasper Lindskow, Michael Riis Andersen, and Jes Frellsen. 2023. Creating the next generation of news experience on ekstrabladet.dk with rec- ommender systems. InProceedings of the 17th ACM Conference on Recommender Systems(Singapore, Singapore)(RecSys ’23). Association for Computing Machin- ery, New York, NY, USA, 1067–1070. doi:10.1145/36...
-
[16]
Johannes Kruse, Kasper Lindskow, Michael Riis Andersen, and Jes Frellsen
-
[17]
doi:10.1038/s42256-025-01043-5 In press
Why Design Choices Matter in Recommender Systems.Nature Machine Intelligence7, 6 (2025), 979–980. doi:10.1038/s42256-025-01043-5 In press
-
[18]
Johannes Kruse, Kasper Lindskow, Michael Riis Andersen, Ryotaro Shimizu, Julian McAuley, Pierre-Alexandre Mattei, and Jes Frellsen. 2025. Normative Alignment of Recommender Systems via Internal Label Shift. InProceedings of the Nineteenth ACM Conference on Recommender Systems (RecSys ’25). Association for Computing Machinery, New York, NY, USA, 1240–1245....
doi:10.1145/3705328 2025
-
[19]
Johannes Kruse, Kasper Lindskow, Saikishore Kalloori, Marco Polignano, Clau- dio Pomo, Abhishek Srivastava, Anshuk Uppal, Michael Riis Andersen, and Jes Frellsen. 2024. EB-NeRD a large-scale dataset for news recommendation. In Proceedings of the Recommender Systems Challenge 2024(Bari, Italy)(RecSysChal- lenge ’24). Association for Computing Machinery, Ne...
-
[20]
Johannes Kruse, Kasper Lindskow, Saikishore Kalloori, Marco Polignano, Clau- dio Pomo, Abhishek Srivastava, Anshuk Uppal, Michael Riis Andersen, and Jes Frellsen. 2024. RecSys Challenge 2024: Balancing Accuracy and Editorial Values in News Recommendations. InProceedings of the 18th ACM Conference on Recom- mender Systems(Bari, Italy)(RecSys ’24). Associat...
-
[21]
Jian Li, Jieming Zhu, Qiwei Bi, Guohao Cai, Lifeng Shang, Zhenhua Dong, Xin Jiang, and Qun Liu. 2022. MINER: Multi-Interest Matching Network for News Recommendation. InFindings of the Association for Computational Linguistics: ACL 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Irel...
2022
-
[22]
Shiwei Li, Huifeng Guo, Xing Tang, Ruiming Tang, Lu Hou, Ruixuan Li, and Rui Zhang. 2024. Embedding Compression in Recommender Systems: A Survey. ACM Comput. Surv.56, 5, Article 130 (Jan. 2024), 21 pages. doi:10.1145/3637841
-
[23]
Defu Lian, Haoyu Wang, Zheng Liu, Jianxun Lian, Enhong Chen, and Xing Xie. 2020. LightRec: A Memory and Search-Efficient Recommender System. In Proceedings of The Web Conference 2020(Taipei, Taiwan)(WWW ’20). Association for Computing Machinery, New York, NY, USA, 695–705. doi:10.1145/3366423. 3380151
doi:10.1145/3366423 2020
-
[24]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692
Pith/arXiv arXiv 2019
-
[25]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781 [cs.CL] https://arxiv.org/abs/1301.3781
Pith/arXiv arXiv 2013
-
[26]
Tao Qi, Fangzhao Wu, Chuhan Wu, and Yongfeng Huang. 2022. News Rec- ommendation with Candidate-aware User Modeling. arXiv:2204.04726 [cs.IR] https://arxiv.org/abs/2204.04726
Pith/arXiv arXiv 2022
-
[27]
S. Raza and C. Ding. 2022. News recommender system: a review of recent progress, challenges, and opportunities.Artificial Intelligence Review55 (2022), 749–800. doi:10.1007/s10462-021-10043-x
-
[28]
Steffen Rendle, Walid Krichene, Li Zhang, and John Anderson. 2020. Neural Collaborative Filtering vs. Matrix Factorization Revisited. InProceedings of the 14th ACM Conference on Recommender Systems(Virtual Event, Brazil)(RecSys ’20). Association for Computing Machinery, New York, NY, USA, 240–248. doi:10. 1145/3383313.3412488
arXiv 2020
-
[29]
Barry Smyth and Paul McClave. 2001. Similarity vs. Diversity. InProceedings of the 4th International Conference on Case-Based Reasoning: Case-Based Reason- ing Research and Development (ICCBR ’01). Springer-Verlag, Berlin, Heidelberg, 347–361
2001
-
[30]
Harald Steck. 2019. Embarrassingly Shallow Autoencoders for Sparse Data. InThe World Wide Web Conference(San Francisco, CA, USA)(WWW ’19). Association for Computing Machinery, New York, NY, USA, 3251–3257. doi:10.1145/3308558. 3313710
doi:10.1145/3308558 2019
-
[31]
Robin Verachtert, Olivier Jeunen, and Bart Goethals. 2023. Scheduling on a budget: Avoiding stale recommendations with timely updates.Machine Learning with Applications11 (2023), 100455. doi:10.1016/j.mlwa.2023.100455
-
[32]
Sanne Vrijenhoek, Lien Michiels, Johannes Kruse, Alain Starke, {Jordi Viader} Guerrero, and Nava Tintarev. 2023. Report on NORMalize: The First Workshop on the Normative Design and Evaluation of Recommender Systems.CEUR Workshop Proceedings3639 (2023)
2023
-
[33]
Sanne Vrijenhoek, Lien Michiels, Johannes Kruse, Alain Starke, Nava Tintarev, and Jordi Viader Guerrero. 2023. NORMalize: The First Workshop on Normative Design and Evaluation of Recommender Systems. InProceedings of the 17th ACM Conference on Recommender Systems(Singapore, Singapore)(RecSys ’23). Association for Computing Machinery, New York, NY, USA, 12...
arXiv 2023
-
[34]
Shuhei Watanabe. 2023. Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles for Better Empirical Performance. arXiv:2304.11127 [cs.LG] https://arxiv.org/abs/2304.11127
Pith/arXiv arXiv 2023
-
[35]
Chuhan Wu, Fangzhao Wu, Mingxiao An, Jianqiang Huang, Yongfeng Huang, and Xing Xie. 2019. NPA: Neural News Recommendation with Personalized Attention. InProceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining(Anchorage, AK, USA)(KDD ’19). Association for Computing Machinery, New York, NY, USA, 2576–2584. doi:10.114...
doi:10.1145/3292500 2019
-
[36]
Chuhan Wu, Fangzhao Wu, Suyu Ge, Tao Qi, Yongfeng Huang, and Xing Xie. 2019. Neural News Recommendation with Multi-Head Self-Attention. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP- IJCNLP). Association for Computational Linguistics...
2019
-
[37]
doi:10.18653/v1/D19-1671
-
[38]
Chuhan Wu, Fangzhao Wu, Yongfeng Huang, and Xing Xie. 2023. Personalized News Recommendation: Methods and Challenges.ACM Trans. Inf. Syst.41, 1, Article 24 (Jan. 2023), 50 pages. doi:10.1145/3530257
doi:10.1145/3530257 2023
-
[39]
Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, and Ming Zhou. 2020. MIND: A Large-scale Dataset for News Recommendation. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). ...
-
[40]
I must have clicked on something
Árni Már Einarsson, Elisabetta Petrucci, Jannie Møller Hartley, Stine Lomborg, and Johannes Kruse. 2025. “I must have clicked on something” – Users ´ Experiences and Evaluations of News Recommender Systems.Journalism Practice0, 0 (2025), 1–20. arXiv:https://doi.org/10.1080/17512786.2025.2572972 doi:10.1080/17512786.2025.2572972
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.