Pith. sign in

REVIEW 2 major objections 5 minor 40 references

A training-free news recommender that mixes recency with article similarity can beat or match deep models offline and nearly match them online while running over 600 times faster.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 08:21 UTC pith:K2TFBBCD

load-bearing objection Clean, shippable training-free news baseline that beats neural models offline, nearly matches them online at 600 imes speed, and honestly shows the offline–online and distribution gaps. the 2 major comments →

arxiv 2607.10910 v1 pith:K2TFBBCD submitted 2026-07-12 cs.IR cs.LG

ZoRRO: A Zero-Weight Personalized Recommender System for Scalable News Recommendation

classification cs.IR cs.LG
keywords news recommendationtraining-free recommenderzero-weight modelonline A/B testingbeyond-accuracy evaluationrecencyarticle embeddingsscalability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

News sites must recommend brand-new articles to users whose tastes keep shifting, yet heavy neural recommenders are slow to retrain and costly to serve. This paper offers ZoRRO, a zero-weight, training-free scorer that multiplies each candidate’s recency by a weighted sum of how similar it is to articles the user already clicked. The similarity uses fixed embeddings and category labels; nothing is learned on the recommendation task. On a large offline news log ZoRRO outranks strong neural baselines. In a live six-day A/B test it delivers click-through rates almost as high as the best neural model while answering more than 600 times more requests per second. The same experiments also show that two systems with nearly identical click rates can still push very different mixes of topics and sentiment, so accuracy alone does not decide what news actually reaches readers.

Core claim

A simple product of exponential recency and cosine similarity over fixed article embeddings and category vectors is enough to outperform or match trained neural news recommenders on offline ranking metrics and to reach nearly the same online click-through rate, all while remaining training-free and more than 600 times faster at inference.

What carries the argument

The ZoRRO relevance score: for a candidate article c and user history H, r(c) equals the candidate’s intrinsic recency weight times the sum over history of relational similarities each multiplied by the history item’s own recency weight. Relational similarity is just the sum of cosine similarities of pre-computed document embeddings and one-hot category vectors.

Load-bearing premise

Fixed, pre-trained article embeddings plus simple category labels, combined only by cosine similarity and exponential decay, already capture enough of a user’s news interests that no further training on the recommendation objective is needed.

What would settle it

An online A/B test on the same publisher where a neural model that is allowed to update its embeddings on live click data substantially widens the CTR gap over ZoRRO, or where removing either the embedding or the category term collapses ZoRRO’s offline ranking lead.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces ZoRRO, a zero-weight, training-free news recommender that scores candidates by combining exponential recency (intrinsic relevance) with cosine similarity over fixed article embeddings and one-hot categories (relational relevance), as formalized in Eqs. (1)–(4). On the EB-NeRD hidden test set it outperforms NRMS, LSTUR, and NPA on AUC/MRR/nDCG (Table 1), while delivering >600× higher inference throughput (Table 3). A six-day live A/B test on Ekstra Bladet shows CTR nearly on par with NRMS (4.19 % vs 4.33 %) and well above a Popular baseline (Table 7). Beyond-accuracy analyses offline and online (Tables 2, 6) further show that models with similar accuracy induce different topical and sentiment distributions. The authors conclude that lightweight, training-free methods remain competitive for real-time news recommendation and that evaluation should go beyond ranking/CTR alone.

Significance. If the reported numbers hold, the work supplies a strong, reproducible practical baseline for large-scale news recommendation: a simple, training-free scorer that matches or exceeds widely used neural models offline, remains competitive online, and is two orders of magnitude faster to serve. The honest documentation of the offline–online gap and the beyond-accuracy distributional differences are themselves useful contributions for the community. Code is released, the method is fully specified, and the online experiment involves hundreds of thousands of users, which together raise the bar for systems papers in this domain.

major comments (2)
  1. Table 7 reports point estimates of CTR (4.19 % ZoRRO vs 4.33 % NRMS) with only per-method variance and impression counts; no confidence intervals, hypothesis test, or power analysis is given. Because the central online claim is “nearly on par,” a formal statistical comparison (or at least bootstrap CIs) is needed to substantiate that the 0.14 pp gap is negligible rather than a real under-performance.
  2. §4.3 states that all methods share the same rule-based candidate set of the top-250 popular recent articles. This design choice is reasonable for a controlled A/B test, but it limits the claim that ZoRRO is a complete end-to-end recommender; the method never has to solve retrieval or cold-start candidate generation. The manuscript should either (a) quantify how much of the performance is attributable to the shared candidate pool, or (b) clearly scope the contribution as a re-ranker rather than a full recommender.
minor comments (5)
  1. §3, Eq. (4): the unweighted sum of cosine similarities is presented without discussion of possible scaling or temperature parameters; a one-sentence justification or ablation would help readers understand why equal weighting is preferred.
  2. Figure 1 caption and axis labels could more clearly indicate that LSTUR/NPA were evaluated with masked user IDs; the current text buries this detail in the body.
  3. Table 4 reports mean±std over three runs for embeddings, but Tables 1 and 5 do not; consistency of reporting would strengthen the ranking claims.
  4. The abstract and introduction claim “more than 600× faster”; Table 3 shows ~667× versus NRMS. Rounding is fine, but citing the exact throughput ratio once would avoid any impression of exaggeration.
  5. A few typographical inconsistencies appear (e.g., “600 ×” vs “600×”, “0 .5 ms”); a light copy-edit pass would polish the camera-ready version.

Circularity Check

0 steps flagged

No significant circularity: ZoRRO is an empirical systems method whose scoring rule does not embed the evaluation metrics.

full rationale

The paper proposes a training-free scoring function (Eqs. 1–4) that multiplies exponential recency by a sum of cosine similarities over fixed article embeddings and one-hot categories. Hyper-parameters (decay rates, embedding choice) are selected on a held-out validation split via Optuna and then frozen; the reported offline ranking metrics (Table 1) and online CTR (Table 7) are measured on unseen test data and live traffic. No quantity that is later called a “prediction” or “result” is algebraically identical to a fitted input, and no uniqueness theorem or load-bearing self-citation is invoked to force the form of the model. Self-citations appear only as ordinary prior work on the same dataset and beyond-accuracy metrics; they do not close a definitional loop. The derivation chain is therefore self-contained against external benchmarks, and the circularity score is zero.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 1 invented entities

The central empirical claim rests on a small number of free decay rates, the choice of fixed embeddings, and standard domain assumptions about news (recency matters, cosine similarity is a usable proxy for interest). No new physical or mathematical entities are postulated.

free parameters (3)
  • λc (candidate recency decay) = 0.015
    Exponential decay rate for candidate articles; best value 0.015 found by grid search on validation AUC (Figure 2).
  • λh (history recency decay) = 0.0
    Exponential decay rate for history articles; best value 0.0 (no decay) found by the same grid search.
  • article embedding choice = contrastive (in-domain)
    Four embedding families compared; contrastive in-domain embedding selected for ZoRRO because it gave highest AUC (Table 4).
axioms (3)
  • ad hoc to paper User interest in a candidate can be approximated by a sum of cosine similarities to historical clicks, re-weighted only by publication recency.
    Core scoring equation (Eq. 1–2) in Section 3; not derived from a user model but postulated as sufficient.
  • domain assumption News relevance decays exponentially with hours since publication.
    Standard modeling choice in news recommendation; instantiated as Eq. 3.
  • domain assumption Cosine similarity on fixed text embeddings plus one-hot categories is an adequate relational relevance measure.
    Eq. 4; common content-based practice, not re-validated beyond the ablation in Table 5.
invented entities (1)
  • ZoRRO scoring framework independent evidence
    purpose: Zero-weight, training-free personalization for news via the product of intrinsic and relational relevance.
    The specific combination and the claim that it needs no learned parameters are introduced by the paper; independent evidence is the empirical performance itself.

pith-pipeline@v1.1.0-grok45 · 16554 in / 2351 out tokens · 26347 ms · 2026-07-14T08:21:12.429476+00:00 · methodology

0 comments
read the original abstract

We present ZoRRO (Zero-Weight Personalized Recommender System), a zero-weight, training-free framework for personalized news recommendation designed for scalable real-world deployment. ZoRRO outperforms strong neural baselines in offline ranking evaluations and achieves click-through rate performance in online A/B testing that is nearly on par with a state-of-the-art deep learning model, while operating more than 600 times faster. Our experiments reveal gaps between offline and online performance and demonstrate that models with similar click-through rate outcomes can produce markedly different recommendation distributions, thereby influencing the overall news flow. These findings position ZoRRO as a practical and efficient solution for large-scale news recommendation and highlight the importance of evaluating recommender systems using metrics beyond accuracy alone.

Figures

Figures reproduced from arXiv: 2607.10910 by Jes Frellsen, Johannes Kruse, Jon Tofteskov, Julian McAuley, Kasper Lindskow, Michael Riis Andersen, Ryotaro Shimizu.

Figure 1
Figure 1. Figure 1: Effect of history size on AUC performance. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: ZoRRO’s AUC as a function of 𝜆ℎ and 𝜆𝑐 . 4.3 Online A/B test We conducted a six-day online A/B test from November 12 to 18, 2024, on Ekstra Bladet’s website2 , comparing ZoRRO, NRMS, and a Most Popular baseline. All models were served through a FastAPI application on AWS Fargate connected to a PostgreSQL database that continuously updated article content and embeddings. The sys￾tem was deployed across two … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

40 extracted references · 4 canonical work pages

  1. [1]

    Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019. Optuna: A Next-generation Hyperparameter Optimization Framework. InProceedings of the 25th ACM SIGKDD International Confer- ence on Knowledge Discovery & Data Mining(Anchorage, AK, USA)(KDD ’19). Association for Computing Machinery, New York, NY, USA, 2623–2631. doi:10.1...

  2. [2]

    Mingxiao An, Fangzhao Wu, Chuhan Wu, Kun Zhang, Zheng Liu, and Xing Xie. 2019. Neural News Recommendation with Long- and Short-term User Representations. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 336–345. doi:10.18653/v1/P19-1033

  3. [3]

    Andreas Argyriou, Miguel González-Fierro, and Le Zhang. 2020. Microsoft Recommenders: Best Practices for Production-Ready Recommendation Systems. InCompanion Proceedings of the Web Conference 2020(Taipei, Taiwan)(WWW ’20). Association for Computing Machinery, New York, NY, USA, 50–51. doi:10. 1145/3366424.3382692

  4. [4]

    Keshav Balasubramanian, Abdulla Alshabanah, Joshua D Choe, and Murali Annavaram. 2021. cDLRM: Look Ahead Caching for Scalable Training of Recom- mendation Models. InProceedings of the 15th ACM Conference on Recommender Systems(Amsterdam, Netherlands)(RecSys ’21). Association for Computing Ma- chinery, New York, NY, USA, 263–272. doi:10.1145/3460231.347424...

  5. [5]

    Das, Mayur Datar, Ashutosh Garg, and Shyam Rajaram

    Abhinandan S. Das, Mayur Datar, Ashutosh Garg, and Shyam Rajaram. 2007. Google News Personalization: Scalable Online Collaborative Filtering. InPro- ceedings of the 16th International Conference on World Wide Web(Banff, Alberta, Canada)(WWW ’07). Association for Computing Machinery, New York, NY, USA, 271–280. doi:10.1145/1242572.1242610

  6. [6]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL]

  7. [7]

    Maurizio Ferrari Dacrema, Paolo Cremonesi, and Dietmar Jannach. 2019. Are we really making much progress? A worrying analysis of recent neural recommen- dation approaches. InProceedings of the 13th ACM Conference on Recommender Systems(Copenhagen, Denmark)(RecSys ’19). Association for Computing Ma- chinery, New York, NY, USA, 101–109. doi:10.1145/3298689.3347058

  8. [8]

    Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2022. SimCSE: Simple Contrastive Learning of Sentence Embeddings. arXiv:2104.08821 [cs.CL]

  9. [9]

    Mouzhi Ge, Carla Delgado-Battenfeld, and Dietmar Jannach. 2010. Beyond Accuracy: Evaluating Recommender Systems by Coverage and Serendipity. In Proceedings of the Fourth ACM Conference on Recommender Systems(Barcelona, Spain)(RecSys ’10). Association for Computing Machinery, New York, NY, USA, 257–260. doi:10.1145/1864708.1864761

  10. [10]

    Udit Gupta, {Carole Jean} Wu, Xiaodong Wang, Maxim Naumov, Brandon Reagen, David Brooks, Bradford Cottel, Kim Hazelwood, Mark Hempstead, Bill Jia, {Hsien Hsin S.} Lee, Andrey Malevich, Dheevatsa Mudigere, Mikhail Smelyanskiy, Liang Xiong, and Xuan Zhang. 2020. The architectural implications of facebook’s DNN- based personalized recommendation. InProceedin...

  11. [11]

    Andreea Iana, Goran Glavaš, and Heiko Paulheim. 2024. Train Once, Use Flex- ibly: A Modular Framework for Multi-Aspect Neural News Recommendation. InFindings of the Association for Computational Linguistics: EMNLP 2024, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Com- putational Linguistics, Miami, Florida, USA, 9555–9571. do...

  12. [12]

    Marius Kaminskas and Derek Bridge. 2016. Diversity, Serendipity, Novelty, and Coverage: A Survey and Empirical Analysis of Beyond-Accuracy Objectives in Recommender Systems.ACM Trans. Interact. Intell. Syst.7, 1, Article 2 (Dec. 2016), 42 pages. doi:10.1145/2926720

  13. [13]

    Lee, and Xuan Zhang

    Liu Ke, Udit Gupta, Mark Hempstead, Carole-Jean Wu, Hsien-Hsin S. Lee, and Xuan Zhang. 2022. Hercules: Heterogeneity-Aware Inference Serving for At- Scale Personalized Recommendation. arXiv:2203.07424 [cs.DC] https://arxiv. org/abs/2203.07424

  14. [14]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Opti- mization. arXiv:1412.6980 [cs.LG] https://arxiv.org/abs/1412.6980

  15. [15]

    Johannes Kruse, Kasper Lindskow, Michael Riis Andersen, and Jes Frellsen. 2023. Creating the next generation of news experience on ekstrabladet.dk with rec- ommender systems. InProceedings of the 17th ACM Conference on Recommender Systems(Singapore, Singapore)(RecSys ’23). Association for Computing Machin- ery, New York, NY, USA, 1067–1070. doi:10.1145/36...

  16. [16]

    Johannes Kruse, Kasper Lindskow, Michael Riis Andersen, and Jes Frellsen

  17. [17]

    doi:10.1038/s42256-025-01043-5 In press

    Why Design Choices Matter in Recommender Systems.Nature Machine Intelligence7, 6 (2025), 979–980. doi:10.1038/s42256-025-01043-5 In press

  18. [18]

    Johannes Kruse, Kasper Lindskow, Michael Riis Andersen, Ryotaro Shimizu, Julian McAuley, Pierre-Alexandre Mattei, and Jes Frellsen. 2025. Normative Alignment of Recommender Systems via Internal Label Shift. InProceedings of the Nineteenth ACM Conference on Recommender Systems (RecSys ’25). Association for Computing Machinery, New York, NY, USA, 1240–1245....

  19. [19]

    Johannes Kruse, Kasper Lindskow, Saikishore Kalloori, Marco Polignano, Clau- dio Pomo, Abhishek Srivastava, Anshuk Uppal, Michael Riis Andersen, and Jes Frellsen. 2024. EB-NeRD a large-scale dataset for news recommendation. In Proceedings of the Recommender Systems Challenge 2024(Bari, Italy)(RecSysChal- lenge ’24). Association for Computing Machinery, Ne...

  20. [20]

    Johannes Kruse, Kasper Lindskow, Saikishore Kalloori, Marco Polignano, Clau- dio Pomo, Abhishek Srivastava, Anshuk Uppal, Michael Riis Andersen, and Jes Frellsen. 2024. RecSys Challenge 2024: Balancing Accuracy and Editorial Values in News Recommendations. InProceedings of the 18th ACM Conference on Recom- mender Systems(Bari, Italy)(RecSys ’24). Associat...

  21. [21]

    Jian Li, Jieming Zhu, Qiwei Bi, Guohao Cai, Lifeng Shang, Zhenhua Dong, Xin Jiang, and Qun Liu. 2022. MINER: Multi-Interest Matching Network for News Recommendation. InFindings of the Association for Computational Linguistics: ACL 2022, Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Irel...

  22. [22]

    Shiwei Li, Huifeng Guo, Xing Tang, Ruiming Tang, Lu Hou, Ruixuan Li, and Rui Zhang. 2024. Embedding Compression in Recommender Systems: A Survey. ACM Comput. Surv.56, 5, Article 130 (Jan. 2024), 21 pages. doi:10.1145/3637841

  23. [23]

    Defu Lian, Haoyu Wang, Zheng Liu, Jianxun Lian, Enhong Chen, and Xing Xie. 2020. LightRec: A Memory and Search-Efficient Recommender System. In Proceedings of The Web Conference 2020(Taipei, Taiwan)(WWW ’20). Association for Computing Machinery, New York, NY, USA, 695–705. doi:10.1145/3366423. 3380151

  24. [24]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692

  25. [25]

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781 [cs.CL] https://arxiv.org/abs/1301.3781

  26. [26]

    Tao Qi, Fangzhao Wu, Chuhan Wu, and Yongfeng Huang. 2022. News Rec- ommendation with Candidate-aware User Modeling. arXiv:2204.04726 [cs.IR] https://arxiv.org/abs/2204.04726

  27. [27]

    Raza and C

    S. Raza and C. Ding. 2022. News recommender system: a review of recent progress, challenges, and opportunities.Artificial Intelligence Review55 (2022), 749–800. doi:10.1007/s10462-021-10043-x

  28. [28]

    Steffen Rendle, Walid Krichene, Li Zhang, and John Anderson. 2020. Neural Collaborative Filtering vs. Matrix Factorization Revisited. InProceedings of the 14th ACM Conference on Recommender Systems(Virtual Event, Brazil)(RecSys ’20). Association for Computing Machinery, New York, NY, USA, 240–248. doi:10. 1145/3383313.3412488

  29. [29]

    Barry Smyth and Paul McClave. 2001. Similarity vs. Diversity. InProceedings of the 4th International Conference on Case-Based Reasoning: Case-Based Reason- ing Research and Development (ICCBR ’01). Springer-Verlag, Berlin, Heidelberg, 347–361

  30. [30]

    Harald Steck. 2019. Embarrassingly Shallow Autoencoders for Sparse Data. InThe World Wide Web Conference(San Francisco, CA, USA)(WWW ’19). Association for Computing Machinery, New York, NY, USA, 3251–3257. doi:10.1145/3308558. 3313710

  31. [31]

    Robin Verachtert, Olivier Jeunen, and Bart Goethals. 2023. Scheduling on a budget: Avoiding stale recommendations with timely updates.Machine Learning with Applications11 (2023), 100455. doi:10.1016/j.mlwa.2023.100455

  32. [32]

    Sanne Vrijenhoek, Lien Michiels, Johannes Kruse, Alain Starke, {Jordi Viader} Guerrero, and Nava Tintarev. 2023. Report on NORMalize: The First Workshop on the Normative Design and Evaluation of Recommender Systems.CEUR Workshop Proceedings3639 (2023)

  33. [33]

    Sanne Vrijenhoek, Lien Michiels, Johannes Kruse, Alain Starke, Nava Tintarev, and Jordi Viader Guerrero. 2023. NORMalize: The First Workshop on Normative Design and Evaluation of Recommender Systems. InProceedings of the 17th ACM Conference on Recommender Systems(Singapore, Singapore)(RecSys ’23). Association for Computing Machinery, New York, NY, USA, 12...

  34. [34]

    Shuhei Watanabe. 2023. Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles for Better Empirical Performance. arXiv:2304.11127 [cs.LG] https://arxiv.org/abs/2304.11127

  35. [35]

    Chuhan Wu, Fangzhao Wu, Mingxiao An, Jianqiang Huang, Yongfeng Huang, and Xing Xie. 2019. NPA: Neural News Recommendation with Personalized Attention. InProceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining(Anchorage, AK, USA)(KDD ’19). Association for Computing Machinery, New York, NY, USA, 2576–2584. doi:10.114...

  36. [36]

    Chuhan Wu, Fangzhao Wu, Suyu Ge, Tao Qi, Yongfeng Huang, and Xing Xie. 2019. Neural News Recommendation with Multi-Head Self-Attention. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP- IJCNLP). Association for Computational Linguistics...

  37. [37]

    doi:10.18653/v1/D19-1671

  38. [38]

    Chuhan Wu, Fangzhao Wu, Yongfeng Huang, and Xing Xie. 2023. Personalized News Recommendation: Methods and Challenges.ACM Trans. Inf. Syst.41, 1, Article 24 (Jan. 2023), 50 pages. doi:10.1145/3530257

  39. [39]

    Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, and Ming Zhou. 2020. MIND: A Large-scale Dataset for News Recommendation. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). ...

  40. [40]

    I must have clicked on something

    Árni Már Einarsson, Elisabetta Petrucci, Jannie Møller Hartley, Stine Lomborg, and Johannes Kruse. 2025. “I must have clicked on something” – Users ´ Experiences and Evaluations of News Recommender Systems.Journalism Practice0, 0 (2025), 1–20. arXiv:https://doi.org/10.1080/17512786.2025.2572972 doi:10.1080/17512786.2025.2572972