Pith. sign in

REVIEW 4 major objections 4 minor 62 references

We're Still Doing It (All) Wrong: Recommender Systems, Fifteen Years Later

T0 review · 4 major / 4 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read Recommender systems research is still doing it wrong, fifteen years after the original critique, and the fix is a shift in posture, not new metrics.

desk verdict A clear, well-sourced position paper that restates Amatriain's critique for 2025; the field-level generalizations outrun the evidence, but as a workshop provocation it is a solid piece. read the letter →

arxiv 2509.09414 v1 pith:WE7XFPZT submitted 2025-09-11 cs.IR cs.AI

classification cs.IRcs.AI
keywords recommendersystemsevaluationmethodologyofflinereproducibilityordinalratingsepistemichumilityhuman-centeredcomputingpositionpaper
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the core methodological failures Amatriain called out in 2011 still define recommender systems research: ordinal ratings are treated as interval data, models are optimized on historical interaction logs rather than on user experience, and evaluation relies on fragile offline benchmarks that measure accuracy, not impact. It claims these practices have not been corrected but merely covered over with more complex models, larger datasets, and better tooling. The authors contend that reproducibility frameworks, new metrics, and community workshops have not changed standard practice, and that the field is therefore solving increasingly polished versions of the wrong problem. Their proposal is not technical but epistemological: recommender research should adopt epistemic humility, treat evaluation as a site of inquiry rather than a procedural step, and recognize recommenders as sociotechnical interventions that shape what people see and believe.

What carries the argument

The argument is carried by the recurrence of a single methodological default pattern: data is treated as stable ground truth (ordinal ratings as interval numbers, clicks as revealed preference), models are judged by offline accuracy metrics on a few standard datasets, and evaluation is treated as a procedural step rather than an interpretive act. The paper unpacks this pattern through three lenses—the persistence of the same fallacies in new model families, the cult of evaluation that rewards narrow metric gains, and the distinction between reproducibility and reliability—and contrasts it with a proposed alternative posture: evaluation as a site of inquiry, with disaggregated and distributio

What would settle it

A falsifying observation would be a systematic audit of a random sample of, say, 200 recommender-system papers from 2023–2025 across RecSys, SIGIR, WWW, and KDD showing that a majority report beyond-accuracy metrics, use human evaluation or variance-aware estimates, and evaluate on diverse datasets—or a study showing that standard frameworks (e.g., Elliot, RecBole, Cornac, LensKit) produce consistent algorithm rankings when configured identically.

Watch

Extended reading notes

Core claim

The paper's central claim is that Amatriain's 2011 diagnosis remains accurate: recommender systems research still prioritizes optimization over understanding, evaluation over reflection, and abstraction over accountability. The specific fallacies he identified—treating ordinal ratings as interval data, chasing minuscule accuracy gains, and measuring success against historical data instead of human outcomes—persist in more subtle and systemic forms. The paper supports this with evidence from a systematic review of evaluation practices (32% of evaluation papers use a single metric, over 70% use three or fewer), reproducibility studies showing that standard frameworks produce inconsistent resul

Load-bearing premise

The argument assumes that the cited evidence—a systematic review of evaluation papers, reproducibility studies of standard toolkits, and the authors' own peer-review anecdotes—is representative of the entire recommender systems community, including industry practice and venues beyond the RecSys conference.

Editorial extensions

If this is right

  • If the paper is right, the field's standard publication and evaluation practices are systematically misaligned with the goal of serving users, meaning that many reported improvements are artifacts of fragile benchmarks rather than genuine advances.
  • New frameworks for evaluation (FEVR, CAFE) and reproducibility-oriented initiatives (such as standardized baseline tuning) will not change practice unless the field's reward structures value human-centered and reflective work as much as algorithmic novelty.
  • The widespread adoption of large language models in recommendation may be unjustified: evidence cited in the paper indicates LLMs can memorize common datasets and often perform on par with resource-light baselines in offline evaluations, while carrying much higher environmental costs.
  • The paper's framing of recommenders as sociotechnical interventions implies that research should measure influence and impact on users and society, not just predictive accuracy, which would change which results count as contributions.
  • Participatory and co-design approaches, where users and other stakeholders help define goals and metrics, become a necessary complement to offline evaluation rather than a peripheral specialty.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct testable consequence of the paper's claim is that a large-scale audit of papers from 2023–2025 across major venues would still find that the majority use one or two standard datasets, few beyond-accuracy metrics, and no user studies—if the paper's diagnosis is correct.
  • The paper's argument implicitly predicts that current reproducibility tooling (e.g., standardized evaluation frameworks) will not by itself reduce the variance across implementations for common baselines, since the variance stems from undisclosed design defaults and methodological choices as much as from code bugs.
  • The authors leave implicit that the same critique likely applies to neighboring fields such as information retrieval and general machine learning, where leaderboard-driven evaluation and benchmark saturation are similarly decoupled from real-world user value.
  • An editorial extension: the strongest lever for change may be institutional—changing what conferences and journals reward—rather than developing yet more evaluation frameworks, since the paper itself notes that such frameworks remain underutilized.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This position paper revisits Xavier Amatriain's 2011 critique of recommender systems research and argues that the field is 'still doing it (all) wrong': ordinal ratings are still treated as interval data, evaluation continues to prioritize offline accuracy metrics over user experience, reproducibility efforts have not produced comparability, and new practices (e.g., large language models, resource-intensive architectures) have introduced additional problems. The authors support their argument with a systematic review of evaluation-focused recommender systems research [19], studies of framework-dependent evaluation outcomes [3,29], and examples of community-led reform initiatives. The paper concludes with a call for epistemic humility, human-aligned evaluation, and participatory design. It is explicitly framed as a provocation and reflection rather than a new empirical study.

Significance. If the paper's diagnosis is correct, it identifies a systematic misalignment between the field's dominant evaluation practices and its stated goal of serving users. The paper is a useful synthesis that connects long-standing methodological critiques (ordinal vs. interval scales, offline proxy metrics, dataset concentration) with newer concerns (environmental cost, LLM memorization, ethical fragility). Its strengths include grounding several claims in the systematic survey of evaluation-focused research [19], citing concrete reproducibility evidence [3,4,29], and offering a constructive list of community initiatives. However, the paper's central assertion is broad: it generalizes from a subset of evaluation-focused papers and from the authors' own experiences to 'most of the community' and 'the vast majority of published works.' That generalization is the main correctness risk and needs to be addressed before the headline claim can be fully credited.

major comments (4)
  1. [Section 1 and Section 2] The central claim is stated as 'most of the community still today continues to chase minuscule performance gains' (Section 1) and 'the vast majority of published works' (Section 2). The strongest quantitative support, the systematic review [19], covers evaluation-focused recommender systems research from 2017 to 2022, not all recommender systems papers, and not necessarily current practice nor work published at SIGIR/WWW/KDD or in industry. The reproducibility studies [3,4,29] demonstrate that framework choices and hyperparameters change measured outcomes, but they do not show that the underlying assumptions are 'unchanged or altered in superficial ways' across venues. This is load-bearing for the paper's thesis. Please either provide venue-stratified or longitudinal evidence for the generalization, or explicitly restrict the claim to the evidence base and adjust phrases such as 'most of
  2. [Sections 3 and 4] The statistics '32% of papers use a single metric; over 70% use three or fewer' and 'nearly half of papers use a single dataset; two-thirds rely on datasets used only once' are taken from [19], which is itself a review of evaluation-focused research. The text presents these as field-wide facts. This conflation matters because the paper's conclusion about persistent fallacies depends on the current breadth of the problem. Please clarify (a) the precise population from which these numbers are drawn, (b) the 2017-2022 time window, and (c) any known differences for non-evaluation-focused or more recent work.
  3. [Section 5] The 'New Sins' section makes several empirical generalizations without quantitative support: 'Few papers disclose compute costs or carbon impact' and 'A significant chunk of recommender system research has embraced LLMs as a panacea.' These statements go beyond the cited examples. If they are meant as the authors' assessment from experience, say so; if they are meant as field-wide claims, cite systematic evidence similar to [19]. As written, these load-bearing supporting claims are anecdotal and would be easy for a sympathetic reader to overstate.
  4. [Section 2, reviewer-anecdote paragraph] The passage that relies on 'reviewer comments on submitted manuscripts' and 'the authors have multiple examples of such comments' is presented as evidence that 'machine learning and algorithmic approaches often dominate.' This is an anecdotal basis for a field-level claim. It should either be clearly labeled as personal experience, contextualized with available data (e.g., publication analyses such as [17]), or omitted from the argumentative chain that leads to 'the vast majority of published works.'
minor comments (4)
  1. [Author affiliation] The affiliation line contains a typo: 'Netherldans' should be 'Netherlands.'
  2. [Author footnote] The author footnote contains an apparent OCR/typesetting artifact: '/envel⌢pe-⌢penalansaid@acm.org' should be cleaned up to a normal email address.
  3. [Section 3, Soboroff quote] The quote from Ian Soboroff is presented via a web archive link. It would be helpful to include the full date of the tweet and, if possible, a more conventional citation, since the quote is used as a conceptual anchor.
  4. [Section 7] The concluding claim that 'the publication and reward structures in the field continue to favor narrow performance gains' is plausible but is asserted without a citation. A reference to an analysis of incentives (e.g., the existing citation [59], which is about ML scholarship more broadly) would strengthen this statement; as written it reads as an editorial judgment.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's argument is an interpretive review supported by independently checkable surveys and external reproducibility studies, not a derivation from its own fitted inputs.

full rationale

No circularity found. This is an argumentative review paper rather than a quantitative derivation: its central claim ('Amatriain's critique remains true... Recommender systems research is still doing it wrong') is an interpretive assessment of community practice, not a result derived from the authors' own equations or fitted parameters. The quantitative evidence for the persistence of narrow evaluation practices comes primarily from [19], a systematic literature review of evaluation-focused recommender systems research (2017–2022); that review is an independently checkable empirical study whose stated scope does not presuppose the paper's conclusion, and its co-authorship by one of the present authors does not make the argument circular. The other self-citations (e.g., [13], [23], [39], [43], [46], [49]) serve as pointers to prior work, workshop initiatives, or proposed alternative frameworks; none is used to define a key term in terms of the target conclusion, and none is invoked as an unavoidable 'uniqueness theorem' that forces the authors' choice. No fitted parameter is renamed as a prediction, no known result is merely relabeled, and no ansatz is smuggled in solely through self-citation. The concern that [19] covers only evaluation-focused papers and may not support the broad 'most of the community' generalization is a scope/representativeness correctness risk, not a circularity. The argument is a rhetorical synthesis of external evidence, and it does not reduce to its own inputs by construction.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities. The paper's argument rests on the representativeness of cited survey and reproducibility evidence, a normative premise that recommenders should serve human welfare and stakeholder values, and an asserted causal claim that publication and reward structures perpetuate flawed evaluation. These are assumptions the reader must accept for the central critique to hold.

assumptions (3)
  • domain assumption The evidence in the cited surveys (Bauer et al. [19]; Schmidt et al. [29]) and personal reviewer anecdotes is representative of recommender-systems practice as a whole.
    The paper's diagnosis in Sections 1 and 3 that 'most of the community still today continues to chase minuscule performance gains' and that human-focused research is devalued rests on this representativeness, which is not independently quantified.
  • domain assumption Recommender systems are sociotechnical interventions whose success should be judged by human welfare and stakeholder values, not just predictive accuracy.
    Normative premise stated in Section 7 and used as the basis for the 'better alternative' in Table 1; without it the call for epistemic humility and participatory design has no grounding.
  • domain assumption The persistence of flawed evaluation stems primarily from publication and reward structures rather than from ignorance.
    Section 7 claims 'researchers, developers, and reviewers are aware of these issues. Yet the publication and reward structures in the field continue to favor narrow performance gains.' This causal assumption is asserted, not demonstrated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of We're Still Doing It (All) Wrong: Recommender Systems, Fifteen Years Later." pith.science (2026). https://pith.science/paper/WE7XFPZT

@misc{pith2026250909414,
  author       = {Pith},
  title        = {Pith review of: We're Still Doing It (All) Wrong: Recommender Systems, Fifteen Years Later},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WE7XFPZT}},
  note         = {Machine review of arXiv:2509.09414}
}
read the original abstract

In 2011, Xavier Amatriain sounded the alarm: recommender systems research was "doing it all wrong" [1]. His critique, rooted in statistical misinterpretation and methodological shortcuts, remains as relevant today as it was then. But rather than correcting course, we added new layers of sophistication on top of the same broken foundations. This paper revisits Amatriain's diagnosis and argues that many of the conceptual, epistemological, and infrastructural failures he identified still persist, in more subtle or systemic forms. Drawing on recent work in reproducibility, evaluation methodology, environmental impact, and participatory design, we showcase how the field's accelerating complexity has outpaced its introspection. We highlight ongoing community-led initiatives that attempt to shift the paradigm, including workshops, evaluation frameworks, and calls for value-sensitive and participatory research. At the same time, we contend that meaningful change will require not only new metrics or better tooling, but a fundamental reframing of what recommender systems research is for, who it serves, and how knowledge is produced and validated. Our call is not just for technical reform, but for a recommender systems research agenda grounded in epistemic humility, human impact, and sustainable practice.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

62 extracted references · 11 canonical work pages

  1. [19]

    Bauer, E

    C. Bauer, E. Zangerle, A. Said, Exploring the Landscape of Recommender Systems Evaluation: Prac- tices and Perspectives, ACM Trans. Recomm. Syst. 2 (2024) 11:1–11:31. doi:10.1145/3629170

  2. [29]

    Evaluating the performance-deviation of itemKNN in RecBole and LensKit

    M. Schmidt, J. Nitschke, T. Prinz, Evaluating the performance-deviation of itemKNN in RecBole and LensKit, 2024. doi:10.48550/arXiv.2407.13531.arXiv:2407.13531

  3. [17]

    Smyth, People who liked this also liked

    B. Smyth, People who liked this also liked... a publication analysis of three decades of recommender systems research, ACM Transactions on Recommender Systems (2025)

  4. [1]

    Amatriain, Recommender Systems: We’re doing it (all) wrong, AI, software, tech, and people

    X. Amatriain, Recommender Systems: We’re doing it (all) wrong, AI, software, tech, and people. Not in that order. By X (2011)

  5. [2]

    J. A. Konstan, J. Riedl, Recommender systems: from algorithms to user experience, User modeling and user-adapted interaction 22 (2012) 101–123

  6. [3]

    Akhadam, O

    A. Akhadam, O. Kbibchi, L. Mekouar, Y. Iraqi, A comparative evaluation of recommender systems tools, IEEE Access (2025)

  7. [4]

    Shehzad, D

    F. Shehzad, D. Jannach, Everyone’sa winner! on hyperparameter tuning of recommendation models, in: Proceedings of the 17th ACM Conference on Recommender Systems, 2023, pp. 652–657

  8. [5]

    Quadrana, P

    M. Quadrana, P. Cremonesi, D. Jannach, Sequence-Aware Recommender Systems, ACM Comput. Surv. 51 (2018) 1–36. doi:10.1145/3190616

Show all 62 references
  1. [6]

    M. D. Ekstrand, V. Mahant, Sturgeon and the Cool Kids: Problems with Top-N Recommender Evaluation, in: Proceedings of the 30th Florida Artificial Intelligence Research Society Conference, FLAIRS 30, AAAI Press, 2017. URL: https://aaai.org/papers/639-flairs-2017-15534/

  2. [7]

    M. D. Smucker, H. Chamani, Extending MovieLens-32M to Provide New Evaluation Objectives,

  3. [8]

    Cañamares, P

    R. Cañamares, P. Castells, Should I Follow the Crowd?: A Probabilistic Analysis of the Effectiveness of Popularity in Recommender Systems, in: The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR ’18, ACM, 2018, pp. 415–424. doi...

  4. [9]

    O. Jeunen, Revisiting offline evaluation for implicit-feedback recommender systems, in: Pro- ceedings of the 13th ACM Conference on Recommender Systems, RecSys ’19, Association for Computing Machinery, New York, NY, USA, 2019, pp. 596–600. doi:10.1145/3298689.3347069

  5. [10]

    Saito, T

    Y. Saito, T. Udagawa, H. Kiyohara, K. Mogi, Y. Narita, K. Tateno, Evaluating the Robustness of Off-Policy Evaluation, in: Proceedings of the 15th ACM Conference on Recommender Systems, RecSys ’21, Association for Computing Machinery, New York, NY, USA, 2021, pp. 114–123. doi:1...

  6. [11]

    A. J. B. Chaney, B. M. Stewart, B. E. Engelhardt, How algorithmic confounding in recommendation systems increases homogeneity and decreases utility, in: Proceedings of the 12th ACM Confer- ence on Recommender Systems, RecSys ’18, ACM, 2018, pp. 224–232. doi: 10.1145/3240323. 3240370

  7. [12]

    Kouki, I

    P. Kouki, I. Fountalis, N. Vasiloglou, X. Cui, E. Liberty, K. Al Jadda, From the Lab to Production: A Case Study of Session-Based Recommendations in the Home-Improvement Domain, in: Proceed- ings of the Fourteenth ACM Conference on Recommender Systems, RecSys ’20, ACM, 2020, p...

  8. [13]

    Bauer, A

    C. Bauer, A. Said, E. Zangerle, Evaluation perspectives of recommender systems: Driving research and education (dagstuhl seminar 24211), Dagstuhl Reports 14 (2024) 58–172. doi:10.4230/DagRep. 14.5.58

  9. [14]

    Resnick, N

    P. Resnick, N. Iacovou, M. Suchak, P. Bergstrom, J. Riedl, GroupLens: An Open Architecture for Collaborative Filtering of Netnews, in: CSCW ’94, ACM, New York, NY, USA, 1994, pp. 175–186. doi:10.1145/192844.192905

  10. [15]

    Billsus, M

    D. Billsus, M. J. Pazzani, A hybrid user model for news story classification, in: Proceedings of the Seventh International Conference on User Modeling, volume 407 ofCISM International Centre for Mechanical Sciences, Springer, 1999, pp. 99–108. doi:10.1007/978-3-7091-2490-1_10

  11. [16]

    W. Hill, L. Stead, M. Rosenstein, G. Furnas, Recommending and evaluating choices in a virtual community of use, in: CHI ’95: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, ACM Press/Addison-Wesley Publishing Co., New York, NY, USA, 1995, pp. 194–20...

  12. [18]

    S. M. McNee, J. Riedl, J. A. Konstan, Being accurate is not enough: how accuracy metrics have hurt recommender systems, in: CHI’06 extended abstracts on Human factors in computing systems, 2006, pp. 1097–1101

  13. [20]

    Kaminskas, D

    M. Kaminskas, D. Bridge, Diversity, serendipity, novelty, and coverage: a survey and empirical analysis of beyond-accuracy objectives in recommender systems, ACM Transactions on Interactive Intelligent Systems (TiiS) 7 (2016) 1–42

  14. [21]

    Voorhees, D

    E. Voorhees, D. Harman, Overview of the sixth text retrieval conference (TREC-6), Nist Special Publication Sp (1998) 1–24

  15. [22]

    Barocas, A

    S. Barocas, A. Guo, E. Kamar, J. Krones, M. R. Morris, J. W. Vaughan, W. D. Wadsworth, H. Wallach, Designing disaggregated evaluations of AI systems: Choices, considerations, and tradeoffs, in: Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, Association...

  16. [23]

    M. D. Ekstrand, B. Carterette, F. Diaz, Distributionally-informed recommender system evaluation, ACM Transactions on Recommender Systems 2 (2024) 6:1–27. doi:10.1145/3613455

  17. [24]

    Zangerle, C

    E. Zangerle, C. Bauer, Evaluating Recommender Systems: Survey and Framework, ACM Comput. Surv. 55 (2022) 170:1–170:38. doi:10.1145/3556536

  18. [25]

    Bauer, L

    C. Bauer, L. Chen, N. Ferro, N. Fuhr, Conversational agents: A framework for evaluation (cafe)(dagstuhl perspectives workshop 24352), Dagstuhl Reports 14 (2025) 53–58

  19. [26]

    Z. Sun, D. Yu, H. Fang, J. Yang, X. Qu, J. Zhang, C. Geng, Are We Evaluating Rigorously? Benchmarking Recommendation for Reproducible Evaluation and Fair Comparison, in: RecSys ’20, Association for Computing Machinery, New York, NY, USA, 2020, pp. 23–32. doi:10.1145/ 3383313.3412489

  20. [27]

    Hidasi, Á

    B. Hidasi, Á. T. Czapp, Widespread Flaws in Offline Evaluation of Recommender Systems, in: Proceedings of the 17th ACM Conference on Recommender Systems, RecSys ’23, Association for Computing Machinery, New York, NY, USA, 2023, pp. 848–855. doi:10.1145/3604915.3608839

  21. [28]

    Z. Meng, R. McCreadie, C. Macdonald, I. Ounis, Exploring Data Splitting Strategies for the Evaluation of Recommendation Models, in: Fourteenth ACM Conference on Recommender Systems, Association for Computing Machinery, New York, NY, USA, 2020, pp. 681–686. doi:10. 1145/3383313.3418479

  22. [31]

    Spillo, A

    G. Spillo, A. De Filippo, C. Musto, M. Milano, G. Semeraro, Comparing data reduction strategies for energy-efficient green recommender systems, Journal of Intelligent Information Systems (2025). doi:10.1007/s10844-025-00965-1

  23. [32]

    Spillo, A

    G. Spillo, A. G. Valerio, F. Franchini, A. De Filippo, C. Musto, M. Milano, G. Semeraro, RecSys CarbonAtor: Predicting Carbon Footprint of Recommendation System Models, in: L. Boratto, A. De Filippo, E. Lex, F. Ricci (Eds.), Recommender Systems for Sustainability and Social Go...

  24. [33]

    Di Palma, F

    D. Di Palma, F. A. Merra, M. Sfilio, V. W. Anelli, F. Narducci, T. Di Noia, Do llms memorize recommendation datasets? a preliminary study on movielens-1m, in: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025,...

  25. [34]

    Milano, M

    S. Milano, M. Taddeo, L. Floridi, Recommender systems and their ethical challenges, AI & SOCIETY 35 (2020) 957–967. doi:10.1007/s00146-020-00950-y

  26. [35]

    S. Wang, X. Zhang, Y. Wang, F. Ricci, Trustworthy Recommender Systems, ACM Trans. Intell. Syst. Technol. 15 (2024) 84:1–84:20. doi:10.1145/3627826

  27. [36]

    M. D. Ekstrand, M. C. Willemsen, Behaviorism is not enough: Better recommendations through listening to users, in: Proceedings of the 10th ACM Conference on Recommender Systems, 2016, pp. 221–224. doi:10.1145/2959100.2959179

  28. [37]

    Burke, M

    R. Burke, M. Sylvester, Post-userist recommender systems: A manifesto, 2024. doi: 10.48550/ arXiv.2410.11870.arXiv:2410.11870

  29. [38]

    N. J. Belkin, S. E. Robertson, Some ethical and political implications of theoretical research in information science, in: Proceedings of the ASIS Annual Meeting, 1976. URL: https://www. researchgate.net/publication/255563562

  30. [39]

    A. Said, Recommender Systems for Social Good:The Role of Accountability and Sustainability, in: Workshop on Recommender Systems for Social Good, volume Proceeding of the 2024 workshop on Recommender Systems for Social Good ofRecSoGood, Springer, 2024, pp. 1–4

  31. [40]

    Jannach, A

    D. Jannach, A. Said, M. Tkalcic, M. Zanker, Recommender Systems for Good (RS4Good): Survey of Use Cases and a Call to Action for Research that Matters, ACM Trans. Recomm. Syst. (2025). doi:10.1145/3746648

  32. [41]

    Felfernig, M

    A. Felfernig, M. Wundara, T. N. T. Tran, S. Polat-Erdeniz, S. Lubos, M. El Mansi, D. Garber, V.-M. Le, Recommender systems for sustainability: overview and research issues, Frontiers in big Data 6 (2023) 1284511

  33. [42]

    G. K. Patro, A. Chakraborty, A. Banerjee, N. Ganguly, Towards safety and sustainability: designing local recommendations for post-pandemic world, in: Proceedings of the 14th ACM Conference on Recommender Systems, 2020, pp. 358–367

  34. [43]

    Ekstrand, M

    M. Ekstrand, M. S. Pera, A. Said, AltRecSys: A Workshop on Alternative, Unexpected, and Critical Ideas in Recommendation, in: Proceedings of the 18th ACM Conference on Recommender Systems, RecSys ’24, Association for Computing Machinery, New York, NY, USA, 2024, pp. 1216–1218....

  35. [44]

    Vrijenhoek, L

    S. Vrijenhoek, L. Michiels, J. Kruse, A. Starke, N. Tintarev, J. Viader Guerrero, Normalize: The first workshop on normative design and evaluation of recommender systems, in: Proceedings of the 17th ACM Conference on Recommender Systems, 2023, pp. 1252–1254

  36. [45]

    Boratto, A

    L. Boratto, A. De Filippo, E. Lex, F. Ricci, First international workshop on recommender systems for sustainability and social good (recsogood 2024), in: Proceedings of the 18th ACM Conference on Recommender Systems, 2024, pp. 1239–1241

  37. [46]

    M. D. Ekstrand, A. Razi, A. Sarcevic, M. S. Pera, R. Burke, K. L. Wright, Recommending with, not for: Co-designing recommender systems for social good, ACM Transactions on Recommender Systems Just Accepted (2025). doi:10.1145/3759261

  38. [47]

    Burke, G

    R. Burke, G. Adomavicius, T. Bogers, T. Di Noia, D. Kowald, J. Neidhardt, Özlem Özgöbek, M. S. Pera, N. Tintarev, J. Ziegler, De-centering the (traditional) user: Multistakeholder evaluation of recommender systems, International Journal of Human-Computer Studies 203 (2025) 103...

  39. [48]

    Heitz, J

    L. Heitz, J. A. Croci, M. Sachdeva, A. Bernstein, Informfully - Research Platform for Reproducible User Studies, in: Proceedings of the 18th ACM Conference on Recommender Systems, RecSys ’24, Association for Computing Machinery, New York, NY, USA, 2024, pp. 660–669. doi:10.114...

  40. [49]

    Burke, M

    R. Burke, M. Ekstrand, Conducting Recommender Systems User Studies Using POPROX, in: Ad- junct Proceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization, UMAP Adjunct ’25, Association for Computing Machinery, New York, NY, USA, 2025, pp. 1–2. doi:...

  41. [50]

    Brodt, F

    T. Brodt, F. Hopfgartner, Shedding light on a living lab: The CLEF NEWSREEL open recommenda- tion platform, in: Proceedings of the 5th Information Interaction in Context Symposium, ACM, Regensburg Germany, 2014, pp. 223–226. doi:10.1145/2637002.2637028

  42. [51]

    [Accessed 02-09-2025]

    RecSys Challenge 2024, https://www.recsyschallenge.com/2024/, 2024. [Accessed 02-09-2025]

  43. [52]

    [Accessed 02-09-2025]

    RecSys Challenge 2021, https://www.recsyschallenge.com/2021/, 2021. [Accessed 02-09-2025]

  44. [53]

    Olteanu, S

    A. Olteanu, S. L. Blodgett, A. Balayn, A. Wang, F. Diaz, F. d. P. Calmon, M. Mitchell, M. Ekstrand, R. Binns, S. Barocas, Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI- Informed Conception of Rigor, 2025. doi:10.48550/arXiv.2506.14652. arXiv:2506.14652

  45. [54]

    Bauer, C

    C. Bauer, C. Bagchi, O. A. Hundogan, K. van Es, Where are the values? a systematic literature review on news recommender systems, ACM Transactions on Recommender Systems 2 (2024) 1–40

  46. [55]

    Stray, A

    J. Stray, A. Halevy, P. Assar, D. Hadfield-Menell, C. Boutilier, A. Ashar, C. Bakalar, L. Beattie, M. Ek- strand, C. Leibowicz, et al., Building human values into recommender systems: An interdisciplinary synthesis, ACM Transactions on Recommender Systems 2 (2024) 1–57

  47. [56]

    T. V. Rampisela, T. Ruotsalo, M. Maistro, C. Lioma, Can we trust recommender system fairness evaluation? the role of fairness and relevance, in: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024, pp. 271–281

  48. [57]

    Ungruh, M

    R. Ungruh, M. Al Nahadi, M. S. Pera, Mirror, mirror: Exploring stereotype presence among top-n recommendations that may reach children, ACM Transactions on Recommender Systems (2025)

  49. [58]

    J. L. Herlocker, J. A. Konstan, L. G. Terveen, J. T. Riedl, Evaluating collaborative filtering recom- mender systems, ACM Transactions on Information Systems (TOIS) 22 (2004) 5–53

  50. [59]

    Z. C. Lipton, J. Steinhardt, Troubling Trends in Machine Learning Scholarship: Some ML papers suffer from flaws that could mislead the public and stymie future research., Queue 17 (2019) Pages 80:45–Pages 80:77. doi:10.1145/3317287.3328534

  51. [60]

    Seaver, Algorithms as culture: Some tactics for the ethnography of algorithmic systems, Big Data & Society 4 (2017) 2053951717738104

    N. Seaver, Algorithms as culture: Some tactics for the ethnography of algorithmic systems, Big Data & Society 4 (2017) 2053951717738104. doi:10.1177/2053951717738104

  52. [61]

    Lucherini, M

    E. Lucherini, M. Sun, A. Winecoff, A. Narayanan, T-RECS: A simulation tool to study the societal impact of recommender systems, 2021. URL: http://arxiv.org/abs/2107.08959

  53. [62]

    Mitra, Search and Society: Reimagining Information Access for Radical Futures, 2024

    B. Mitra, Search and Society: Reimagining Information Access for Radical Futures, 2024. doi:10. 48550/arXiv.2403.17901.arXiv:2403.17901

  54. [2025]

    doi:10.48550/arXiv.2504.01863.arXiv:2504.01863

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.