Pith. sign in

REVIEW 3 major objections 6 minor 74 references

Informfully Recommenders -- Reproducibility Framework for Diversity-aware Intra-session Recommendations

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims a single open-source framework can run norm-aware recommender experiments end to end, and that diversity-targeting models and re-rankers reach set targets while keeping accuracy on par with neural baselines.

desk verdict Open-source framework that actually works and fills a real gap in normative recommender systems, but the paper overreaches when it claims re-rankers preserve accuracy: AUC was never computed for re-ranked lists. read the letter →

arxiv 2508.13019 v1 pith:SVQHHUVF submitted 2025-08-18 cs.IR

classification cs.IR
keywords norm-awarerecommendersystemsreproducibilityframeworkdiversity-awarerecommendationintra-sessionre-rankingnewsnormativetargetdistributionusersimulationdiversitymetrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that norm-aware recommender research has lacked a single reproducibility framework spanning every stage of the pipeline, and that the proposed open-source extension fills that gap with an end-to-end setup: dataset augmentation, model training, static and dynamic intra-session re-ranking, normative and traditional diversity metrics, and item visualization for online user studies. Using three news datasets, the paper argues that target-distribution-optimizing models and re-rankers reach the normative target values for sentiment and party diversity while keeping AUC comparable to standard neural baselines. A sympathetic reader should care because this makes societal norms such as democratic diversity testable, comparable, and deployable in an ordinary recommendation pipeline instead of a bespoke research prototype.

What carries the argument

The load-bearing object is the Normative Target Distribution (NTD), a table that maps attribute values to desired percentages and is plugged into both in-processing models and post-processing re-rankers. The paper uses five buckets for party mentions—governing parties 15%, opposition parties 15%, both 15%, others 15%, no mention 40%—together with sentiment buckets whose widths vary by dataset. Greedy-KL and PM-2 use the same NTD to re-rank candidate lists, while the diversity-driven random walk D-RDW uses it to guide graph exploration, and the PLD and EPD filtering algorithms implement participative and deliberative democracy models respectively. The evaluation stage then defines normative target values (NTV) as the best achievable diversity scores under the dataset and NTD, giving every other method a concrete benchmark to approach. This machinery does the argument's work by making diversity an explicit, tunable target rather than a post-hoc metric.

What would settle it

Run the same NTD-optimizing models and re-rankers on a held-out news dataset with a different party and sentiment composition while keeping the same NTD, and check whether sentiment and party Gini and intra-list distance still match the normative target values and AUC stays within the neural baseline range; if the match disappears, the claim is dataset-specific rather than general.

Watch

Extended reading notes

Core claim

At the center of the framework is a normative target distribution (NTD): a user-specified set of desired frequencies for item attributes such as political-party mentions and sentiment. The paper's central discovery is that NTD-optimizing algorithms—two normative filtering models, a diversity-driven random walk, and the Greedy-KL and PM-2 re-rankers—produce recommendation lists whose sentiment and party Gini coefficients and intra-list distances match the normative target values across all three datasets, while their AUC scores stay close to those of neural baselines. This is presented as evidence that societal diversity norms can be operationalized as a concrete distributional target and enforced at the model or re-ranking stage without a major accuracy cost. Category diversity was not similarly improved, which the paper attributes to the NTD not including article category. The framework also contributes a user simulator for intra-session re-ranking and a state-saving mechanism that lets researchers reuse intermediate results across stages.

Load-bearing premise

The claim carries only as much normative weight as the hand-set target distribution does; if the chosen party and sentiment proportions do not correspond to a defensible democratic ideal, then reaching them is not normatively meaningful.

Editorial extensions

If this is right

  • Researchers can compare pre-processing, in-processing, and post-processing interventions within one pipeline on identical data and metrics, so claims about diversity gains can be attributed to a specific stage.
  • NTD-based models and re-rankers give platform developers a lightweight way to enforce editorial norms such as balanced party exposure: on all three datasets the sentiment and party targets were met exactly, with AUC close to neural models.
  • The user simulator makes intra-session dynamics testable offline, so position-biased and attribute-biased browsing can be examined before an online user study.
  • The Save State Manager allows a single candidate list to be reused across re-ranking and evaluation runs, reducing the cost of reproducibility checks.
  • Category diversity remains a gap because the NTD omits it; including category in the target distribution should bring the same gains as for sentiment and party.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An unstated extension of the NTD idea is to define target distributions for other normative dimensions—gender representation, viewpoint balance, topic breadth—and reuse the same framework to test them; the paper only exercises party and sentiment.
  • The fact that category diversity was not improved suggests the NTD mechanism itself, not the optimizers, controls which diversity dimensions are met; adding category to the five buckets and rerunning the same experiments is a natural test.
  • The reported AUC is computed on the full article pool and the paper notes re-ranker AUC was not recomputed, so the accuracy-cost claim for re-rankers would be stronger with rank-aware AUC on the recommended top-20 lists.
  • Because the NTD is hand-set, the same framework could also compare competing normative choices: the divergence between two target distributions' resulting lists could quantify the practical difference between editorial policies.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. Informfully Recommenders extends the Cornac recommender-system framework with a norm-aware, diversity-focused reproducibility pipeline covering preprocessing (data loading and text augmentation for sentiment, political actors, complexity, event clusters, categories), in-processing (neural baselines, normative filtering algorithms PLD/EPD, random walks including D-RDW), post-processing (static re-rankers G-KL, PM-2, MMR and dynamic intra-session re-ranking with a user simulator), and evaluation (RADio and traditional diversity metrics, plus integration with the Informfully visualization platform). The authors demonstrate the pipeline through offline experiments on EB-NeRD, MIND, and NeMig, reporting diversity metrics and AUC for the model stage.

Significance. The primary contribution is a public, modular, open-source framework that, if taken up, could lower the barrier to norm-driven recommender-system research; it is the first such end-to-end framework known to the authors, and the accompanying documentation and configuration files support reproducibility. The framework's integration with Cornac and with the Informfully research platform is a genuine asset. The experimental results are best interpreted as an internal-consistency demonstration that the NTD-optimizing components steer the chosen metrics toward their targets; the accuracy-diversity tradeoff for the re-ranking stage is not quantified, and the point estimates lack variance.

major comments (3)
  1. [Section 4 (Metrics) and Tables 3-5] The claim in Section 5 ("Accuracy and Diversity Tradeoffs") that target-distribution-optimizing models and re-rankers "not only diversify the recommendation list but also present relevant items similar to the state-of-the-art neural models" is unsupported for re-rankers. The AUC column is blank for all re-ranked rows (G-KL, PM-2, MMR, POS, ATT) in all three tables, and the paper explicitly states "we chose not to recompute the AUC for the re-rankers." The accuracy-diversity tradeoff for the post-processing stage, one of the two showcased capabilities, is therefore unquantified. To support the claim, the authors should compute an accuracy metric on re-ranked lists, or explicitly limit the claim to the model stage.
  2. [Section 4 (NTD definition) and Section 5 (Traditional Diversity Metrics)] The demonstration that G-KL, PM-2, D-RDW, PLD, and EPD "consistently reach the NTV" for sentiment and party Gini/ILD is partly circular. The Sent./Party Gini and ILD metrics are computed with respect to the same target distributions (the five NTD buckets, e.g., 15/15/15/15/40 for parties, and the sentiment bins) that these algorithms are explicitly designed to match. High performance on these metrics is therefore by construction, not an independent empirical finding. The paper should either use distribution-free diversity metrics (e.g., Shannon entropy without a fixed target) or frame these results as an internal-consistency check that the optimizers do what they claim, not as evidence for the normative value of the targets.
  3. [Sections 4-5] The experimental comparisons report single point estimates without variance, confidence intervals, or statistical tests. Claims such as "there is a tie in AUC," "neural models substantially outperform the other two families on MIND," and "top spots for NeMig are shared" are based on naked differences between a single run per condition. Given that the paper emphasizes reproducibility, the authors should report at least multiple seeds with standard deviations, or clearly state that the experiments are illustrative only and not intended as a competitive benchmark.
minor comments (6)
  1. [Section 5, Traditional Diversity Metrics] The phrase "To use, this presents evidence for the effectiveness of NTD-based approaches" appears to contain a typo; it should likely read "To us, this presents evidence...".
  2. [Section 3.1, Political Actors and Section 4] The text in Section 3.1 says labels for "Governing Party," "Opposition Party," or "Others/Foreign Parties," while Section 4 defines the buckets as governing party, opposition party, or others (independent and foreign parties). Align the terminology for consistency.
  3. [Section 5, RADio Diversity Metrics] The statement that on EB-NeRD "target distribution-optimizing models achieve the top score in all but one category" is not supported by Table 3: the best Activation value (0.377) belongs to NPA+ATT, a neural model with dynamic re-ranking, not to a target-distribution-optimizing model. Clarify the intended comparison or adjust the sentence.
  4. [Tables 3-5, NTV row] The NTV row values differ across datasets (e.g., Sent. Gini NTV is 0.133 on EB-NeRD and MIND but 0.000 on NeMig). Provide a short derivation or explanation of how NTV is computed, since this is not defined in the text.
  5. [References] Reference [3] and [4] are the same paper (An et al., 2019), and RWE-D is cited as [45] in Tables 3-5 but as [46] in Section 3.2. Unify these citations.
  6. [Section 4, Metrics] The sentence "We use five diversity metrics for measuring divergence of Activation, Category Calibration, Complexity Calibration, Fragmentation, Alternative Voices, and Representation" lists six metrics. Correct the count or the list.

Circularity Check

1 steps flagged · score 6.0 of 10

The NTV demonstration is circular by construction: the NTD-optimizing algorithms and the Normative Target Values are both derived from the same NTD, so the 'finding' that D-RDW, G-KL, and PM-2 reach the target values restates their design objective; the re-ranker AUC claim is a separate evidence gap, not circularity.

  1. self definitional [Section 4 (Re-rankers; NTD definition) and Section 5 (Traditional Diversity Metrics; NTV row)]
    "G-KL and PM-2 re-rankers use the same diversity dimensions and distributions as those defined by our aforementioned target distribution. ... By default, NTD consists of five buckets: 1) governing parties (15%), 2) opposition parties (15%), 3) governing and opposition parties (15%), 4) others (e.g., independent and foreign parties, 15%), and 5) articles with no political party mentions (40%). ... A separate row for Normative Target Values (NTV) is included that shows the best achievable scores given the underlying dataset and NTD. ..."

    D-RDW is defined around the NTD ('combines it with NTD'), and G-KL/PM-2 are explicitly set to 'use the same diversity dimensions and distributions as those defined by our aforementioned target distribution.' The NTV row is defined as 'the best achievable scores given the underlying dataset and NTD,' i.e., the metric values obtained when a list matches that same distribution. Therefore the reported result that these methods reach NTV for sentiment and party Gini/ILD is a direct restatement of the optimization objective, not an independent empirical discovery. The metric and the algorithm are built from the same normative target, so the outcome is forced by construction.

full rationale

The core contribution of Informfully Recommenders, an open-source four-stage normative reproducibility framework extending Cornac, is supported by the architecture, code, and comparison table, and is not circular. The circularity is confined to the experimental demonstration: the headline that NTD-based models and re-rankers reach the normative target values for sentiment and party diversity is self-definitional, because the same NTD fixes both the algorithms' optimization objective and the NTV reference row. The AUC half of the headline is not circular but is currently unsupported for re-rankers: Section 4 states 'we chose not to recompute the AUC for the re-rankers,' yet Section 5 asserts that re-rankers 'present relevant items similar to the state-of-the-art neural models'; that is a measurement gap, not a reduction by construction. The paper also relies on several same-author citations for PLD, EPD, D-RDW, and the Informfully platform, but those are used as prior implementations and do not by themselves force the central framework claim. On balance, the central framework claim retains independent content, but one load-bearing experimental result reduces to the algorithm's own target distribution, so the circularity score is 6 rather than higher.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The central framework claim rests on standard software-engineering assumptions that the components work as described and the code runs, while the experimental demonstration rests on hand-set target distributions and norm operationalization choices. No invented natural entities are introduced; D-RDW is a new algorithmic component rather than a postulated entity.

free parameters (6)
  • NTD party distribution percentages = governing 15%, opposition 15%, both 15%, others 15%, no party mention 40%
    Chosen by hand in Section 4 (Models) to define the target distribution for D-RDW, G-KL, and PM-2; not derived from data or from an external normative standard.
  • Sentiment distribution bins for EB-NeRD and MIND = [-1,-0.5):20%, [-0.5,0):30%, [0,0.5):30%, [0.5,1]:20%
    Chosen by hand in Section 4 to define the sentiment component of the NTD for two datasets.
  • Sentiment distribution bins for NeMig = [-1,0):50%, [0,1]:50%
    Adjusted for NeMig because the paper states 'the data did not allow for a more detailed assessment of sentiment' (Section 4 footnote).
  • Recommendation list size k = 20
    Top-20 recommendations used for all experiments, chosen with reference to prior work [24, 59] in Section 4.
  • Random walk hop count = 3 hops (5 for D-RDW on NeMig)
    Number of hops used for random walk models; D-RDW on NeMig uses 5 hops because the graph is too sparse for 20 items with fewer hops (Section 4 footnote).
  • Re-ranker diversity dimension weights = equal weights for sentiment and political parties
    Section 4 (Re-rankers) states the diversity dimensions are weighted equally, a hand-set choice that affects all re-ranker outcomes.
assumptions (4)
  • domain assumption RADio metrics are valid operationalizations of normative diversity
    The paper relies on Vrijenhoek et al. [59] and democracy theory [28] to treat calibration, fragmentation, activation, representation, and alternative voices as normative diversity measures; this is an unproven value judgment imported from prior work.
  • domain assumption A normative target distribution over selected item attributes (political parties, sentiment) is an appropriate operationalization of diversity norms
    Section 3.2 introduces NTDs as the mechanism for normative filtering, and Section 4 sets the specific target values; the choice of attributes and values is an editorial or policy decision, not derived from evidence.
  • domain assumption The user simulator's click model approximates real user behavior
    Section 3.3 specifies two default behaviors (position bias and preference for previously read categories); Section 6 acknowledges that online user studies are still needed, so this assumption is explicitly unvalidated.
  • standard math Standard metrics (Gini, ILD, AUC, RADio) as implemented in the framework are correctly computed
    Sections 3.4 and 4 take these metric definitions from cited literature; the framework's implementations are not independently verified in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Informfully Recommenders -- Reproducibility Framework for Diversity-aware Intra-session Recommendations." pith.science (2026). https://pith.science/paper/SVQHHUVF

@misc{pith2026250813019,
  author       = {Pith},
  title        = {Pith review of: Informfully Recommenders -- Reproducibility Framework for Diversity-aware Intra-session Recommendations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SVQHHUVF}},
  note         = {Machine review of arXiv:2508.13019}
}
read the original abstract

Norm-aware recommender systems have gained increased attention, especially for diversity optimization. The recommender systems community has well-established experimentation pipelines that support reproducible evaluations by facilitating models' benchmarking and comparisons against state-of-the-art methods. However, to the best of our knowledge, there is currently no reproducibility framework to support thorough norm-driven experimentation at the pre-processing, in-processing, post-processing, and evaluation stages of the recommender pipeline. To address this gap, we present Informfully Recommenders, a first step towards a normative reproducibility framework that focuses on diversity-aware design built on Cornac. Our extension provides an end-to-end solution for implementing and experimenting with normative and general-purpose diverse recommender systems that cover 1) dataset pre-processing, 2) diversity-optimized models, 3) dedicated intrasession item re-ranking, and 4) an extensive set of diversity metrics. We demonstrate the capabilities of our extension through an extensive offline experiment in the news domain.

Figures

Figures reproduced from arXiv: 2508.13019 by the authors.

Figure 1
Figure 1. Informfully Recommenders extension of the existing Cornac pipeline by implementing a diversity-aware four-stage [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

74 extracted references · 55 canonical work pages

  1. [1]

    Gediminas Adomavicius and YoungOk Kwon. 2011. Improving aggregate rec- ommendation diversity using ranking-based techniques. IEEE Transactions on Knowledge and Data Engineering 24, 5 (2011), 896–911

  2. [2]

    Jafar Afzali, Aleksander Mark Drzewiecki, Krisztian Balog, and Shuo Zhang. 2023. Usersimcrs: A user simulation toolkit for evaluating conversational recommender systems. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining . 1160–1163

  3. [3]

    Mingxiao An, Fangzhao Wu, Chuhan Wu, Kun Zhang, Zheng Liu, and Xing Xie. 2019. Neural news recommendation with long-and short-term user rep- resentations. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 336–345

  4. [4]

    Mingxiao An, Fangzhao Wu, Chuhan Wu, Kun Zhang, Zheng Liu, and Xing Xie

  5. [5]

    Vito Walter Anelli, Alejandro Bellogín, Antonio Ferrara, Daniele Malitesta, Fe- lice Antonio Merra, Claudio Pomo, Francesco Maria Donini, and Tommaso Di Noia. 2021. Elliot: A comprehensive and rigorous framework for reproducible recommender systems evaluation. In Proceedings of the 44th international ACM SIGIR conference on research and development in inf...

  6. [6]

    Andreas Argyriou, Miguel González-Fierro, and Le Zhang. 2020. Microsoft recommenders: Best practices for production-ready recommendation systems. In Companion Proceedings of the Web Conference 2020 . 50–51

  7. [7]

    Christine Bauer, Chandni Bagchi, Olusanmi A Hundogan, and Karin van Es. 2024. Where are the values? a systematic literature review on news recommender systems. ACM Transactions on Recommender Systems 2, 3 (2024), 1–40

  8. [8]

    Joeran Beel and Haley Dixon. 2021. The ‘unreasonable’effectiveness of graphical user interfaces for recommender systems. In Adjunct Proceedings of the 29th ACM Conference on User Modeling, Adaptation and Personalization . 22–28

Show all 74 references
  1. [9]

    Abraham Bernstein, Claes De Vreese, Natali Helberger, Wolfgang Schulz, Katha- rina Zweig, Lucien Heitz, Suzanne Tolmeijer, et al . 2021. Diversity in News Recommendation. Dagstuhl Manifestos 9, 1 (2021), 43–61

  2. [10]

    Keith Bradley and Barry Smyth. 2001. Improving Recommendation Diversity. https://api.semanticscholar.org/CorpusID:11075976

  3. [11]

    Jaime Carbonell and Jade Goldstein. 1998. The use of MMR, diversity-based reranking for reordering documents and producing summaries. In Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval. 335–336

  4. [12]

    Pablo Castells, Neil Hurley, and Saul Vargas. 2021. Novelty and diversity in recommender systems. In Recommender systems handbook. Springer, 603–646

  5. [13]

    Chong Chen, Min Zhang, Yongfeng Zhang, Yiqun Liu, and Shaoping Ma. 2020. Efficient neural matrix factorization without sampling for recommendation.ACM Transactions on Information Systems (TOIS) 38, 2 (2020), 1–28

  6. [14]

    Xiong-Hui Chen, Bowei He, Yang Yu, Qingyang Li, Zhiwei Qin, Wenjie Shang, Jieping Ye, and Chen Ma. 2023. Sim2rec: A simulator-based decision-making approach to optimize real-world long-term user engagement in sequential recom- mender systems. In 2023 IEEE 39th International Co...

  7. [15]

    Patrick John Chia, Jacopo Tagliabue, Federico Bianchi, Chloe He, and Brian Ko

  8. [16]

    Fabian Christoffel, Bibek Paudel, Chris Newell, and Abraham Bernstein. 2015. Blockbusters and wallflowers: Accurate, diverse, and scalable recommendations with random walks. In Proceedings of the 9th ACM Conference on Recommender Systems. 163–170

  9. [17]

    Charles LA Clarke, Maheedhar Kolla, Gordon V Cormack, Olga Vechtomova, Azin Ashkan, Stefan Büttcher, and Ian MacKinnon. 2008. Novelty and diversity in information retrieval evaluation. In Proceedings of the 31st annual international ACM SIGIR conference on Research and develop...

  10. [18]

    Elisia L Cohen. 2002. Online journalism as market-driven journalism. Journal of broadcasting & Electronic media 46, 4 (2002), 532–548

  11. [19]

    Van Dang and W Bruce Croft. 2012. Diversity by proportionality: an election- based approach to search result diversification. In Proceedings of the 35th interna- tional ACM SIGIR conference on Research and development in information retrieval . 65–74

  12. [20]

    Michael D Ekstrand. 2020. Lenskit for python: Next-generation software for recommender systems experiments. In Proceedings of the 29th ACM international conference on information & knowledge management . 2999–3006

  13. [21]

    Naieme Hazrati and Francesco Ricci. 2022. Simulating users’ interactions with recommender systems. In Adjunct proceedings of the 30th acm conference on user modeling, adaptation and personalization . 95–98

  14. [22]

    Lucien Heitz. 2023. Classification of Normative Recommender Systems. In Pro- ceedings of the First Workshop on the Normative Design and Evaluation of Recom- mender Systems

  15. [23]

    Lucien Heitz, Julian A Croci, Madhav Sachdeva, and Abraham Bernstein. 2024. Informfully - Research Platform for Reproducible User Studies. In Proceedings of the 18th ACM Conference on Recommender Systems . 660–669

  16. [24]

    Lucien Heitz, Oana Inel, and Sanne Vrijenhoek. 2024. Recommendations for the Recommenders: Reflections on Prioritizing Diversity in the RecSys Challenge. In Proceedings of the Recommender Systems Challenge 2024 . 22–26

  17. [25]

    Lucien Heitz, Juliane A Lischka, Rana Abdullah, Laura Laugwitz, Hendrik Meyer, and Abraham Bernstein. 2023. Deliberative Diversity for News Recommendations: Operationalization and Experimental User Study. In Proceedings of the 17th ACM Conference on Recommender Systems . 813–819

  18. [26]

    Lucien Heitz, Juliane A Lischka, Alena Birrer, Bibek Paudel, Suzanne Tolmeijer, Laura Laugwitz, and Abraham Bernstein. 2022. Benefits of diverse news rec- ommendations for democracy: A user study. Digital Journalism 10, 10 (2022), 1710–1730

  19. [27]

    Lucien Heitz, Nicolas Mattis, Oana Inel, and Wouter van Atteveldt. 2024. IDEA – Informfully Dataset with Enhanced Attributes. In Proceedings of the Second Workshop on the Normative Design and Evaluation of Recommender Systems

  20. [28]

    Natali Helberger. 2019. On the democratic role of news recommenders. Digital Journalism 7, 8 (2019), 993–1012

  21. [29]

    Natali Helberger, Kari Karppinen, and Lucia D’acunto. 2018. Exposure diversity as a design principle for recommender systems. Information, Communication & Society 21, 2 (2018), 191–207

  22. [30]

    Andreea Iana, Mehwish Alam, Alexander Grote, Nevena Nikolajevic, Katharina Ludwig, Philipp Müller, Christof Weinhardt, and Heiko Paulheim. 2023. NeMig-A Bilingual News Collection and Knowledge Graph about Migration. InProceedings of the International Workshop on News Recommend...

  23. [31]

    Eugene Ie, Chih-wei Hsu, Martin Mladenov, Vihan Jain, Sanmit Narvekar, Jing Wang, Rui Wu, and Craig Boutilier. 2019. Recsim: A configurable simulation platform for recommender systems. arXiv preprint arXiv:1909.04847 (2019)

  24. [32]

    Dietmar Jannach and Christine Bauer. 2020. Escaping the McNamara fallacy: Towards more impactful recommender systems research. Ai Magazine 41, 4 (2020), 79–95

  25. [33]

    Dietmar Jannach, Lukas Lerche, Fatih Gedikli, and Geoffray Bonnin. 2013. What recommenders recommend–an analysis of accuracy, popularity, and sales diver- sity effects. In International conference on user modeling, adaptation, and person- alization. Springer, 25–37

  26. [34]

    Johannes Kruse, Kasper Lindskow, Saikishore Kalloori, Marco Polignano, Clau- dio Pomo, Abhishek Srivastava, Anshuk Uppal, Michael Riis Andersen, and Jes Frellsen. 2024. EB-NeRD a large-scale dataset for news recommendation. In Proceedings of the Recommender Systems Challenge 2...

  27. [35]

    Johannes Kruse, Kasper Lindskow, Saikishore Kalloori, Marco Polignano, Clau- dio Pomo, Abhishek Srivastava, Anshuk Uppal, Michael Riis Andersen, and Jes Frellsen. 2024. RecSys Challenge 2024: Balancing Accuracy and Editorial Val- ues in News Recommendations. In Proceedings of ...

  28. [36]

    Matevž Kunaver and Tomaž Požrl. 2017. Diversity in recommender systems–A survey. Knowledge-based systems 123 (2017), 154–162

  29. [37]

    Jiayu Li, Hanyu Li, Zhiyu He, Weizhi Ma, Peijie Sun, Min Zhang, and Shaoping Ma. 2024. ReChorus2. 0: A Modular and Task-Flexible Recommendation Library. In Proceedings of the 18th ACM Conference on Recommender Systems . 454–464

  30. [38]

    Runze Li, Lucien Heitz, Oana Inel, and Abraham Bernstein. 2025. D-RDW: Diversity-Driven Random Walks for News Recommender Systems. InProceedings of the 19th ACM Conference on Recommender Systems

  31. [39]

    Dawen Liang, Rahul G Krishnan, Matthew D Hoffman, and Tony Jebara. 2018. Variational autoencoders for collaborative filtering. In Proceedings of the 2018 world wide web conference . 689–698

  32. [40]

    Felicia Loecherbach, Judith Moeller, Damian Trilling, and Wouter van Atteveldt

  33. [41]

    Pasquale Lops, Marco Polignano, Cataldo Musto, Antonio Silletti, and Giovanni Semeraro. 2023. ClayRS: An end-to-end framework for reproducible knowledge- aware recommender systems. Information Systems 119 (2023), 102273

  34. [42]

    Hongyu Lu, Min Zhang, Weizhi Ma, Ce Wang, Feng Xia, Yiqun Liu, Leyu Lin, and Shaoping Ma. 2019. Effects of user negative experience in mobile news streaming. In Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval. 705–714

  35. [43]

    Lien Michiels, Robin Verachtert, and Bart Goethals. 2022. Recpack: An (other) experimentation toolkit for top-n recommendation using implicit feedback data. In Proceedings of the 16th ACM Conference on Recommender Systems . 648–651

  36. [44]

    Darryl Ong, Quoc-Tuan Truong, and Hady W Lauw. 2024. Cornac-AB: An Open- Source Recommendation Framework with Native A/B Testing Integration. In Companion Proceedings of the ACM Web Conference 2024 . 1027–1030

  37. [45]

    Bibek Paudel and Abraham Bernstein. 2021. Random walks with erasure: Diver- sifying personalized recommendations on social and information networks. In Proceedings of the Web Conference 2021 . 2046–2057

  38. [46]

    Bibek Paudel, Fabian Christoffel, Chris Newell, and Abraham Bernstein. 2016. Updatable, accurate, diverse, and scalable recommendations for interactive ap- plications. ACM Transactions on Interactive Intelligent Systems (TiiS) 7, 1 (2016), 1–34. RecSys ’25, September 22–26, 20...

  39. [47]

    Felix Petersen, Debarghya Mukherjee, Yuekai Sun, and Mikhail Yurochkin. 2021. Post-processing for individual fairness. Advances in Neural Information Processing Systems 34 (2021), 25944–25955

  40. [48]

    Massimo Quadrana, Paolo Cremonesi, and Dietmar Jannach. 2018. Sequence- aware recommender systems. ACM computing surveys (CSUR) 51, 4 (2018), 1–36

  41. [49]

    Matthew Richardson, Ewa Dominowska, and Robert Ragno. 2007. Predicting clicks: estimating the click-through rate for new ads. In Proceedings of the 16th international conference on World Wide Web. 521–530

  42. [50]

    Aghiles Salah, Quoc-Tuan Truong, and Hady W Lauw. 2020. Cornac: A Com- parative Framework for Multimodal Recommender Systems. Journal of Machine Learning Research 21, 95 (2020), 1–5

  43. [51]

    Holli Sargeant, Eliska Pirkova, Matthias C Kettemann, Marlena Wisniak, Martin Scheinin, Emmi Bevensee, Katie Pentney, Lorna Woods, Lucien Heitz, Bojana Kostic, et al. 2022. Spotlight on Artificial Intelligence and Freedom of Expression: A Policy Manual. Organization for Securi...

  44. [52]

    Harald Steck. 2018. Calibrated recommendations. In Proceedings of the 12th ACM conference on recommender systems . 154–162

  45. [53]

    Zhu Sun, Hui Fang, Jie Yang, Xinghua Qu, Hongyang Liu, Di Yu, Yew-Soon Ong, and Jie Zhang. 2022. Daisyrec 2.0: Benchmarking recommendation for rigorous evaluation. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 7 (2022), 8206–8226

  46. [54]

    Panagiotis Symeonidis, Dmitry Chaltsev, Chemseddine Berbague, and Markus Zanker. 2022. Sequence-aware news recommendations by combining intra-with inter-session user information. Information Retrieval Journal 25, 4 (2022), 461– 480

  47. [55]

    Celina Treuillier, Sylvain Castagnos, Evan Dufraisse, and Armelle Brun. 2022. Being Diverse is Not Enough: Rethinking Diversity Evaluation to Meet Chal- lenges of News Recommender Systems. In Adjunct Proceedings of the 30th ACM Conference on User Modeling, Adaptation and Perso...

  48. [56]

    Quoc-Tuan Truong, Aghiles Salah, and Hady Lauw. 2021. Multi-modal recom- mender systems: Hands-on exploration. In Fifteenth ACM Conference on Recom- mender Systems. 834–837

  49. [57]

    Quoc-Tuan Truong, Aghiles Salah, Thanh-Binh Tran, Jingyao Guo, and Hady W Lauw. 2021. Exploring Cross-Modality Utilization in Recommender Systems. IEEE Internet Computing (2021)

  50. [58]

    Saúl Vargas, Linas Baltrunas, Alexandros Karatzoglou, and Pablo Castells. 2014. Coverage, redundancy and size-awareness in genre diversity for recommender systems. In Proceedings of the 8th ACM Conference on Recommender systems . 209–216

  51. [59]

    Sanne Vrijenhoek, Gabriel Bénédict, Mateo Gutierrez Granada, Daan Odijk, and Maarten De Rijke. 2022. RADio–Rank-Aware Divergence Metrics to Measure Normative Diversity in News Recommendations. In Proceedings of the 16th ACM Conference on Recommender Systems . 208–219

  52. [60]

    Sanne Vrijenhoek, Savvina Daniil, Jorden Sandel, and Laura Hollink. 2024. Diver- sity of what? On the different conceptualizations of diversity in recommender systems. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. 573–584

  53. [61]

    Sanne Vrijenhoek, Mesut Kaya, Nadia Metoui, Judith Möller, Daan Odijk, and Natali Helberger. 2021. Recommenders with a mission: assessing diversity in news recommendations. In Proceedings of the 2021 conference on human information interaction and retrieval. 173–183

  54. [62]

    Sanne Vrijenhoek, Lien Michiels, Johannes Kruse, Alain Starke, Nava Tintarev, and Jordi Viader Guerrero. 2023. Normalize: The first workshop on normative design and evaluation of recommender systems. In Proceedings of the 17th ACM Conference on Recommender Systems . 1252–1254

  55. [63]

    Mingyang Wan, Daochen Zha, Ninghao Liu, and Na Zou. 2023. In-processing modeling techniques for machine learning fairness: A survey. ACM Transactions on Knowledge Discovery from Data 17, 3 (2023), 1–27

  56. [64]

    Chuhan Wu, Fangzhao Wu, Mingxiao An, Jianqiang Huang, Yongfeng Huang, and Xing Xie. 2019. NPA: neural news recommendation with personalized attention. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining . 2576–2584

  57. [65]

    Chuhan Wu, Fangzhao Wu, Suyu Ge, Tao Qi, Yongfeng Huang, and Xing Xie. 2019. Neural news recommendation with multi-head self-attention. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natur...

  58. [66]

    Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, et al. 2020. Mind: A large-scale dataset for news recommendation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics...

  59. [67]

    Sirui Yao, Yoni Halpern, Nithum Thain, Xuezhi Wang, Kang Lee, Flavien Prost, Ed H Chi, Jilin Chen, and Alex Beutel. 2021. Measuring recommender system effects with simulated users. arXiv preprint arXiv:2101.04526 (2021)

  60. [68]

    Kesen Zhao, Shuchang Liu, Qingpeng Cai, Xiangyu Zhao, Ziru Liu, Dong Zheng, Peng Jiang, and Kun Gai. 2023. KuaiSim: A comprehensive simulator for recom- mender systems. Advances in Neural Information Processing Systems 36 (2023), 44880–44897

  61. [69]

    Wayne Xin Zhao, Shanlei Mu, Yupeng Hou, Zihan Lin, Yushuo Chen, Xingyu Pan, Kaiyuan Li, Yujie Lu, Hui Wang, Changxin Tian, et al. 2021. Recbole: Towards a unified, comprehensive and efficient framework for recommendation algorithms. In proceedings of the 30th acm international...

  62. [70]

    Jieming Zhu, Quanyu Dai, Liangcai Su, Rong Ma, Jinyang Liu, Guohao Cai, Xi Xiao, and Rui Zhang. 2022. Bars: Towards open benchmarking for recommender systems. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2912–2923

  63. [71]

    Jieming Zhu, Jinyang Liu, Shuai Yang, Qi Zhang, and Xiuqiang He. 2021. Open benchmarking for click-through rate prediction. In Proceedings of the 30th ACM international conference on information & knowledge management . 2759–2769

  64. [2019]

    In Proceedings of the 57th annual meeting of the association for computational linguistics

    Neural news recommendation with long-and short-term user representa- tions. In Proceedings of the 57th annual meeting of the association for computational linguistics. 336–345

  65. [2020]

    Digital Journalism 8, 5 (2020), 605–642

    The unified framework of media diversity: A systematic literature review. Digital Journalism 8, 5 (2020), 605–642

  66. [2022]

    In Companion Proceedings of the Web Conference 2022

    Beyond ndcg: behavioral testing of recommender systems with reclist. In Companion Proceedings of the Web Conference 2022 . 99–104

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.