Pith. sign in

REVIEW 2 major objections 5 minor 107 references

Measuring the Business Value of Recommender Systems

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This research commentary argues that recommender systems generate real, sometimes large business value, but the range of measured effects is so wide and the statistics so weak that neither the realistic size of that value nor the…

desk verdict A useful, honest survey of field tests on business value; the core claim about offline accuracy is plausible but rests on thin evidence. read the letter →

arxiv 1908.08328 v3 pith:6643LT6E submitted 2019-08-22 cs.IR cs.AIcs.LG

classification cs.IRcs.AIcs.LG
keywords recommendersystemsbusinessvaluefieldtestsA/Btestingofflineevaluationclick-throughrateconversionsalesimpact
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Recommender systems clearly influence what people click, buy, and watch, but the size of their business payoff is not well understood. Reviewing published field tests from e-commerce, news, streaming, dating, and job portals, the authors find reported effects that range from marginal revenue gains around 0.3% to massive lifts of several hundred percent in gross merchandise value. They argue that popular offline accuracy measures such as RMSE or precision are not reliable stand-ins for business value, because in most studies the most accurate algorithms did not win online field tests. The consequence matters for anyone deciding whether to invest in better algorithms versus simpler presentation changes, since one field test doubled clicks just by changing the position and size of the recommendation widget. The paper's core message is that both the realistic quantification of business effects and the academic performance assessment of recommenders remain open problems.

What carries the argument

The analytical engine of the paper is a measurement taxonomy that sorts business value into five categories: click-through rates, adoption and conversion rates, sales and revenue, effects on sales distributions, and user engagement and behavior. Each category implies a different conclusion about value: click-through rates are easy to measure but can reward clickbait; adoption measures are domain-specific; direct revenue is the most informative but often unavailable; sales-distribution shifts can increase or decrease diversity; and engagement is only a proxy for retention. The taxonomy does the work of showing why reported numbers across studies are not comparable and why no single offline metric has yet been tied reliably to business outcomes.

What would settle it

Collect a preregistered sample of recommender A/B tests that includes unpublished and failed tests; if those tests show effects centered near zero, the review's implied typical payoff is overstated - or, conversely, a large paired offline/online benchmark where offline accuracy consistently predicts online revenue would refute the claim that no such correspondence can be assumed.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is a documented gap: the surveyed industry reports show recommenders can create substantial business value in many ways - higher click-through, conversion, sales, engagement - but the reported magnitude varies enormously and is often measured with weak statistics or indirect proxies. In studies that compare algorithms both offline and online, the majority found that the most accurate offline models did not produce the best online results, and in a few cases a simple repositioning of the recommendation widget outperformed any algorithmic change. The authors therefore conclude that the business value of recommenders is real yet poorly quantified, and that offline accuracy metrics cannot be assumed to predict business value without per-case validation.

Load-bearing premise

The load-bearing premise is that the published industry reports and blog posts are honest and representative of real recommender deployments; if unpublished negative results are common or the publicized numbers are cherry-picked, the reported range of effects would overstate typical business value.

Editorial extensions

If this is right

  • Companies using recommenders can expect positive effects on user behavior in many domains, but the expected magnitude depends heavily on the baseline, the domain, and the chosen measurement.
  • A typical direct revenue lift reported in field tests is between one and five percent, which is still substantial in absolute terms for large businesses.
  • An algorithm that wins on offline accuracy such as RMSE should not be assumed to win online; most published comparisons found no such correspondence, and some found the opposite.
  • User interface and presentation choices can dominate algorithmic improvements, with one field test doubling click-through rate just by changing widget position and size.
  • Nearly all surveyed A/B tests lacked sample-size analyses and detailed statistics, so the published effect sizes may not be reliable without better reporting.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If unpublished or failed A/B tests are systemically underrepresented in the literature, the true average business lift of recommenders is probably below the one-to-five percent direct-revenue range the paper compiles.
  • The paper's taxonomy implies a testable prediction: direct measures like revenue should be more stable across studies than indirect measures like click-through rate, because indirect measures are more sensitive to interface and presentation effects.
  • A concrete next step the paper does not spell out is a public paired benchmark where logged interactions and A/B outcomes come from the same deployment, allowing offline metrics to be calibrated against business value instead of assumed.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper is a research commentary that reviews field studies (A/B tests and other real-world deployments) of recommender systems, with the goal of synthesizing what is known about the business value of recommenders. It organizes the reported effects into five measurement categories: click-through rates, adoption/conversion rates, sales and revenue, effects on sales distributions, and user engagement/behavior. It then discusses challenges of measuring business value, pitfalls of A/B testing, and the degree to which offline accuracy metrics such as RMSE, precision, and recall predict online success or business outcomes. The paper concludes that recommender systems can create substantial but highly variable business value, that direct revenue effects are often in the low single-digit percentage range while indirect engagement effects can be much larger, and that important open questions remain about both the realistic quantification of business effects and the use of offline accuracy metrics as proxies for business value.

Significance. If taken with appropriate caveats, the paper performs a useful service: it compiles a scattered literature of industrial field tests, structures the measurement space, and highlights methodological weaknesses that are often underappreciated when offline benchmark results are interpreted. The authors are commendably explicit about several limitations, including the lack of statistical detail in most surveyed A/B tests (Section 3.3) and the difficulty of extrapolating offline accuracy gains to business impact (Section 3.4.1). The distinction between direct revenue measures and indirect engagement measures is a valuable organizing framework for both practitioners and researchers. However, the review is deliberately non-systematic: there is no search protocol, no inclusion criteria, and no quantitative synthesis. The most consequential claim, that offline accuracy is not a reliable proxy for online/business success, rests on a small and likely publication-biased convenience sample. The paper would benefit from more carefully calibrated language in the abstract and conclusion so that its useful qualitative message is not overstated.

major comments (2)
  1. [Section 3.4.1] The sentence 'in the majority of these attempts, the most accurate offline models did neither lead to the best online success nor to a better accuracy perception' is supported only by a count of eight negative comparisons ([9,19,28,30,40,71,73,84]) against three positive ones ([14,20,50]). No inclusion criteria, search protocol, or effect-size synthesis are provided, and Section 3.3 itself concedes that 'in almost all surveyed cases, an analysis of the required sample size and detailed statistical analyses of the A/B tests were missing.' The authors should revise this passage to say explicitly that the statement is an observation about the specific studies they identified, not about the broader population of recommender deployments, and that publication bias and the heterogeneity of baselines and metrics preclude any general quantitative conclusion. The abstract and Section 5 should be aligned with this weaker, defensible claim rather than suggesting that the unreliability of offline accuracy is an established empirical regularity.
  2. [Sections 2 and 4.1] The review reports many large headline effect sizes, such as a 200% CTR increase at YouTube, a 500% increase in Gross Merchandise Bought at eBay, and Amazon's 35% cross-sales figure, without explicitly stating that these are self-selected, publicly disclosed outcomes and therefore likely an upper-bound sample rather than a representative distribution. Section 4.1's statement that direct revenue increases 'are more often reported to lie between one and five percent' is an informal impression, not a computed summary. The authors should add a visible caveat, perhaps at the end of Section 2.1 and repeated in Section 4.1, that the compiled figures are illustrative and that the absence of unpublished negative results, differences in baselines, and varying test durations mean that no typical or expected business-value figure can be reliably estimated from this material.
minor comments (5)
  1. [Section 3.2] The phrase 'increases in sales between one and five percent are reported on average' is ambiguous; the authors should clarify whether this is an informal range, a median, or a rough summary, and should point to the specific studies from which the range is derived.
  2. [Section 3.3] The sentence 'In almost all surveyed cases, an analysis of the required sample size and detailed statistical analyses of the A/B tests were missing' has a subject-verb agreement issue ('analysis' is singular); consider rewriting as '...was missing.'
  3. [Section 2.2.1] When describing the Google News results in [68], the authors note that improved recommendations 'stole' clicks from other parts of the page; this is an important qualification that deserves a corresponding mention in the Section 3 discussion of CTR as a business measure, not only in the descriptive part.
  4. [Section 2.2.2] The discussion of the LinkedIn skill recommendation field test [8] appropriately notes the confound between the recommendation method and the user interface change, but the same caution is not applied consistently to other field tests in Section 2 where UI placement or presentation may have changed; a brief general remark in Section 3.1 about this confound would be helpful.
  5. [References] Several sources are blog posts, industry white papers, or non-archival reports; given that the paper is a review, the authors should state in a short methodological paragraph what types of sources were considered admissible and how they were located.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the review's conclusions are inductive summaries of external field studies, not derivations from their own inputs.

full rationale

This is a research commentary and literature review with no equations, no fitted parameters, and no predictive model, so there is no fitted-input-called-prediction or self-definitional step to exhibit. The central claim that open questions remain about business-value quantification and offline evaluation is a summary of the surveyed industry reports and field studies in Section 2, and the paper itself flags the evidentiary limitations of those studies in Section 3.3 ('in almost all surveyed cases, an analysis of the required sample size and detailed statistical analyses of the A/B tests were missing'); that is an evidence-quality limitation, not a circularity. The most consequential substantive claim, that offline accuracy is a weak proxy for business success (Section 3.4.1), is supported by a cited set of empirical comparisons ([9,19,28,30,40,71,73,84] against [14,20,50]) that includes one co-authored study ([40]) but is not defined in terms of it; the claim is an inductive generalization over externally reported experiments and is not solely dependent on the co-authored study. The authors do cite their own prior work in several places (e.g., [37,38,43,50,70]), but these citations are used as ordinary references to independent empirical studies, proposals, or taxonomies, not as unexamined premises that force the conclusion. No step reduces, by construction or by a self-citation chain, to its own input.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central conclusions depend only on domain assumptions about the trustworthiness and representativeness of published field studies; no free parameters or invented entities are introduced.

assumptions (2)
  • domain assumption Published field-test reports are accurate and representative of real-world recommender deployments.
    Section 2 relies on company blog posts, conference papers, and case studies from Netflix, YouTube, eBay, and others; Section 3.3 admits that most surveyed A/B tests lack sample-size and statistical analyses.
  • domain assumption Observed business lifts in the reviewed studies are attributable to the recommender system rather than to confounds such as seasonality, UI changes, or simultaneous marketing campaigns.
    Several reviewed studies rely on before/after comparisons or one-week tests; the paper itself notes these confounds in Sections 3.1 and 3.3 but still uses the numbers as evidence for the value of recommenders.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Measuring the Business Value of Recommender Systems." pith.science (2026). https://pith.science/paper/6643LT6E

@misc{pith2026190808328,
  author       = {Pith},
  title        = {Pith review of: Measuring the Business Value of Recommender Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6643LT6E}},
  note         = {Machine review of arXiv:1908.08328}
}
read the original abstract

Recommender Systems are nowadays successfully used by all major web sites (from e-commerce to social media) to filter content and make suggestions in a personalized way. Academic research largely focuses on the value of recommenders for consumers, e.g., in terms of reduced information overload. To what extent and in which ways recommender systems create business value is, however, much less clear, and the literature on the topic is scattered. In this research commentary, we review existing publications on field tests of recommender systems and report which business-related performance measures were used in such real-world deployments. We summarize common challenges of measuring the business value in practice and critically discuss the value of algorithmic improvements and offline experiments as commonly done in academic environments. Overall, our review indicates that various open questions remain both regarding the realistic quantification of the business effects of recommenders and the performance assessment of recommendation algorithms in academia.

Figures

Figures reproduced from arXiv: 1908.08328 by the authors.

Figure 1
Figure 1. Overview of Measurement Approaches increase in CTR over a time-decayed popularity-based baseline. Interestingly, this trivial popularity￾based baseline was among the best methods in their live trial. In [68], Liu et al. also experimented with a content-based, collaborative hybrid for Google News recommendations. One particularity of their method was that it considered “local trends” and thus the recent popularity of… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

107 extracted references · 78 canonical work pages

  1. [1]

    S. F. Abdinnour-Helm, B. S. Chaparro, and S. M. Farmer. Using the end-user computing satisfaction (EUCS) instrument to measure satisfaction with a web site. Decision Sciences, 36(2):341–364, 2005

  2. [2]

    Abdollahpouri, R

    H. Abdollahpouri, R. Burke, and B. Mobasher. Recommender systems as multistakeholder environments. In Proceedings of the 25th Conference on User Modeling, Adaptation and Personalization , UMAP ’17, pages 347–348, 2017

  3. [3]

    Abdollahpouri, G

    H. Abdollahpouri, G. Adomavicius, R. Burke, I. Guy, D. Jannach, T. Kamishima, J. Krasnodebski, and L. Pizzato. Beyond personalization: Research directions in multistakeholder recommendation. ArXiv e-prints, 2019. URL h/t_tps://arxiv.org/abs/1905.01986

  4. [4]

    F. Amat, A. Chandrashekar, T. Jebara, and J. Basilico. Artwork Personalization at Net/f_lix. InProceedings of the 12th ACM Conference on Recommender Systems , RecSys ’18, pages 487–488, 2018

  5. [5]

    Amatriain and J

    X. Amatriain and J. Basilico. Net/f_lix recommendations: Beyond the 5 stars. h/t_tps://medium.com/net/f_lix- techblog/net/f_lix-recommendations-beyond-the-5-stars-part-1-55838468f429, 2012

  6. [6]

    T. G. Armstrong, A. Moffat, W. Webber, and J. Zobel. Improvements that don’t add up: Ad-hoc retrieval results since 1998. In Proceedings of the 18th ACM Conference on Information and Knowledge Management , CIKM ’09, pages 601–610, 2009

  7. [7]

    Bambini, P

    R. Bambini, P. Cremonesi, and R. Turrin. Recommender Systems Handbook, chapter A Recommender System for an IPTV Service Provider: a Real Large-Scale Production Environment, pages 299–331. Springer, 2011. Eds. Francesco Ricci, Lior Rokach, Bracha Shapira, and Paul B. Kantor

  8. [8]

    Bastian, M

    M. Bastian, M. Hayes, W. Vaughan, S. Shah, P. Skomoroch, H. Kim, S. Uryasev, and C. Lloyd. LinkedIn Skills: Large-scale Topic Extraction and Inference. In Proceedings of the 8th ACM Conference on Recommender Systems , RecSys ’14, pages 1–8, 2014

Show all 107 references
  1. [9]

    Beel and S

    J. Beel and S. Langer. A comparison of offline evaluations, online evaluations, and user studies in the context of research-paper recommender systems. In Proceedings of the 22nd International Conference on /T_heory and Practice of Digital Libraries, TPDL ’15, pages 153–168, 2015

  2. [10]

    J. Beel, M. Genzmehr, S. Langer, A. N¨urnberger, and B. Gipp. A comparative analysis of offline and online evaluations and discussion of research paper recommender system evaluation. In Proceedings of the International Workshop on Reproducibility and Replication in Recommender S...

  3. [11]

    J. Beel, S. Langer, and B. Gipp. TF-IDuF: A novel term-weighting scheme for user modeling based on users’ personal document collections. In Proceedings of the iConference 2017 , pages 452–459, 2017

  4. [12]

    A. V. Bodapati. Recommendation systems with purchase data. Journal of Marketing Research , 45(1):77–93, 2008

  5. [13]

    F. J. Z. Borgesius, D. Trilling, J. Moeller, B. Bodo, C. H. de Vreese, and N. Helberger. Should we worry about /f_ilter bubbles? Internet Policy Review, 5(1), 2016. 11h/t_tp://www.clef-newsreel.org/ ACM Transactions on Management Information Systems, Vol. 10, No. 4, Article 1....

  6. [14]

    Y. M. Brovman, M. Jacob, N. Srinivasan, S. Neola, D. Galron, R. Snyder, and P. Wang. Optimizing similar item recommendations in a semi-structured marketplace to maximize conversion. In Proceedings of the 10th ACM Conference on Recommender Systems , RecSys ’16, pages 199–202, 2016

  7. [15]

    Carraro and D

    D. Carraro and D. Bridge. Debiased offline evaluation of recommender systems: A weighted-sampling approach (extended abstract). In Proceedings of the ACM RecSys 2019 Workshop on Reinforcement and Robust Estimators for Recommendation (REVEAL ’19), 2019

  8. [16]

    Chapelle, T

    O. Chapelle, T. Joachims, F. Radlinski, and Y. Yue. Large-scale validation and analysis of interleaved search evaluation. ACM Transactions on Information Systems , 30(1):6:1–6:41, 2012

  9. [17]

    Chen, Y.-C

    P.-Y. Chen, Y.-C. Chou, and R. J. Kauffman. Community-based recommender systems: Analyzing business models from a systems operator’s perspective. In Proceedings of the 42nd Hawaii International Conference on System Sciences , HICSS ’09, pages 1–10, 2009

  10. [18]

    Chen and J

    Y. Chen and J. F. Canny. Recommending ephemeral items at web scale. In Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval , SIGIR ’11, pages 1013–1022, 2011

  11. [19]

    Cremonesi, F

    P. Cremonesi, F. Garzo/t_to, and R. Turrin. Investigating the persuasion potential of recommender systems from a quality perspective: An empirical study. ACM Transactions on Interactive Intelligent Systems , 2(2):11:1–11:41, June 2012

  12. [20]

    Cremonesi, F

    P. Cremonesi, F. Garzo/t_to, and R. Turrin. User-centric vs. system-centric evaluation of recommender systems. In Proceedings of the 14th International Conference on Human-Computer Interaction , INTERACT ’13, pages 334–351, 2013

  13. [21]

    M. F. Dacrema, P. Cremonesi, and D. Jannach. Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches. In Proceedings of the 2019 ACM Conference on Recommender Systems (RecSys ’19), Copenhagen, 2019

  14. [22]

    A. S. Das, M. Datar, A. Garg, and S. Rajaram. Google news personalization: Scalable online collaborative /f_iltering. In Proceedings of the 16th International Conference on World Wide Web , WWW ’07, pages 271–280, 2007

  15. [23]

    Davidson, B

    J. Davidson, B. Liebald, J. Liu, P. Nandy, T. Van Vleet, U. Gargi, S. Gupta, Y. He, M. Lambert, B. Livingston, and D. Sampath. /T_he YouTube Video Recommendation System. In Proceedings of the Fourth ACM Conference on Recommender Systems, RecSys ’10, pages 293–296, 2010

  16. [24]

    A. Deng, Y. Xu, R. Kohavi, and T. Walker. Improving the sensitivity of online controlled experiments by utilizing pre-experiment data. In Proceedings of the Sixth ACM International Conference on Web Search and Data Mining, WSDM ’13, pages 123–132, 2013

  17. [25]

    M. B. Dias, D. Locher, M. Li, W. El-Deredy, and P. J. Lisboa. /T_he value of personalised recommender systems to e-business: A case study. In Proceedings of the 2008 ACM Conference on Recommender Systems , RecSys ’08, pages 291–294, 2008

  18. [26]

    W. J. Doll and G. Torkzadeh. /T_he measurement of end-user computing satisfaction.MIS /Q_uarterly, 12(2):259–274, 1988

  19. [27]

    M. A. Domingues, F. Gouyon, A. M. Jorge, J. P. Leal, J. Vinagre, L. Lemos, and M. Sordo. Combining usage and content in an online recommendation system for music in the long tail. International Journal of Multimedia Information Retrieval, 2(1):3–13, 2013

  20. [28]

    M. D. Ekstrand, F. M. Harper, M. C. Willemsen, and J. A. Konstan. User perception of differences in recommender algorithms. In Proceedings of the 8th ACM Conference on Recommender Systems , RecSys ’14, pages 161–168, 2014

  21. [29]

    Friedrich and M

    G. Friedrich and M. Zanker. A taxonomy for generating explanations in recommender systems. AI Magazine, 32(3): 90–98, Jun. 2011

  22. [30]

    Garcin, B

    F. Garcin, B. Faltings, O. Donatsch, A. Alazzawi, C. Bru/t_tin, and A. Huber. Offline and online evaluation of news recommender systems at swissinfo.ch. In Proceedings of the 8th ACM Conference on Recommender Systems , RecSys ’14, pages 169–176, 2014

  23. [31]

    Gilo/t_te, C

    A. Gilo/t_te, C. Calauz`enes, T. Nedelec, A. Abraham, and S. Doll´e. Offline a/b testing for recommender systems. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining , WSDM ’18, pages 198–206, 2018

  24. [32]

    C. A. Gomez-Uribe and N. Hunt. /T_he Net/f_lix recommender system: Algorithms, business value, and innovation. Transactions on Management Information Systems , 6(4):13:1–13:19, 2015

  25. [33]

    B. Gumm. A/B Testing: /T_he Most Powerful Way to Turn Clicks into Customers, chapter Metrics and the Statistics Behind A/B Testing, pages 180–193. Wiley, 2013. Eds. Dan Siroker and Pete Koomen

  26. [34]

    Ilias, N

    A. Ilias, N. B. M. Suki, M. R. Yasoa, and M. Z. A. Razak. /T_he end-user computing satisfaction (EUCS) on computerized accounting system (CAS): How they perceived? Journal of Internet Banking and Commerce , 13(1):1–18, 2008

  27. [35]

    Jagerman, I

    R. Jagerman, I. Markov, and M. de Rijke. When people change their mind: Off-policy evaluation in non-stationary recommendation environments. In Proceedings of the Twel/f_th ACM International Conference on Web Search and Data Mining, WSDM ’19, pages 447–455, 2019. ACM Transactio...

  28. [36]

    J. Jang, D. Zhao, W. Hong, Y. Park, and M. Y. Yi. Uncovering the underlying factors of smart TV UX over time: A multi-study, mixed-method approach. In Proceedings of the ACM International Conference on Interactive Experiences for TV and Online Video , TVX ’16, pages 3–12, 2016

  29. [37]

    Jannach and G

    D. Jannach and G. Adomavicius. Recommendations with a purpose. In Proceedings of the 10th ACM Conference on Recommender Systems, RecSys ’16, pages 7–10, 2016

  30. [38]

    Jannach and G

    D. Jannach and G. Adomavicius. Price and pro/f_it awareness in recommender systems. InProceedings of the 2017 Workshop on Value-A ware and Multi-Stakeholder Recommendation (V AMS) at RecSys 2017, 2017

  31. [39]

    Jannach and K

    D. Jannach and K. Hegelich. A case study on the effectiveness of recommendations in the mobile internet. In Proceedings of the 10th ACM Conference on Recommender Systems , RecSys ’09, pages 205–208, 2009

  32. [40]

    Jannach and L

    D. Jannach and L. Lerche. Offline performance vs. subjective quality experience: A case study in video game recommendation. In Proceedings of the ACM Symposium on Applied Computing , SAC ’17, pages 1649–1654, 2017

  33. [41]

    Jannach, L

    D. Jannach, L. Lerche, and M. Jugovac. Adaptation and evaluation of recommendations for short-term shopping goals. In Proceedings of the 9th ACM Conference on Recommender Systems , RecSys ’15, pages 211–218, 2015

  34. [42]

    Jannach, L

    D. Jannach, L. Lerche, and M. Jugovac. Item familiarity as a possible confounding factor in user-centric recommender systems evaluation. i-com Journal of Interactive Media , 14(1):29–39, 2015

  35. [43]

    Jannach, L

    D. Jannach, L. Lerche, I. Kamehkhosh, and M. Jugovac. What recommenders recommend: an analysis of rec- ommendation biases and possible countermeasures. User Modeling and User-Adapted Interaction , 25(5):427–491, 2015

  36. [44]

    Jannach, P

    D. Jannach, P. Resnick, A. Tuzhilin, and M. Zanker. Recommender systems — Beyond matrix completion. Communi- cations of the ACM, 59(11):94–102, 2016

  37. [45]

    Jannach, M

    D. Jannach, M. Ludewig, and L. Lerche. Session-based item recommendation in e-commerce: On short-term intents, reminders, trends, and discounts. User-Modeling and User-Adapted Interaction, 27(3–5):351–392, 2017

  38. [46]

    Jannach, O

    D. Jannach, O. S. Shalom, and J. A. Konstan. Towards more impactful recommender systems research. In Proceedings of the ACM RecSys 2019 Workshop on the Impact of Recommender Systems (ImpactRS’ 19) , 2019

  39. [47]

    Jiang, S

    R. Jiang, S. Chiappa, T. La/t_timore, A. Gy¨orgy, and P. Kohli. Degenerate feedback loops in recommender systems. In Proceedings of the 2019 Conference on Arti/f_icial Intelligence, Ethics, and Society (AIES ’19), pages 383–390, 2019

  40. [48]

    Joachims, A

    T. Joachims, A. Swaminathan, and T. Schnabel. Unbiased learning-to-rank with biased feedback. In Proceedings of the 17th International Joint Conference on Arti/f_icial Intelligence, IJCAI ’17, pages 781–789, 2017

  41. [49]

    Jugovac, D

    M. Jugovac, D. Jannach, and L. Lerche. Efficient optimization of multiple recommendation quality factors according to individual user tendencies. Expert Systems With Applications, 81:321–331, 2017

  42. [50]

    Kamehkhosh and D

    I. Kamehkhosh and D. Jannach. User Perception of Next-Track Music Recommendations. In Proceedings of the 25th Conference on User Modeling, Adaptation and Personalization , UMAP ’17, pages 113–121, 2017

  43. [51]

    Kamehkhosh, D

    I. Kamehkhosh, D. Jannach, and G. Bonnin. How automated recommendations affect the playlist creation behavior of users. In Proceedings of the Workshop on Intelligent Music Interfaces for Listening and Creation (MILC) at IUI 2018 , 2018

  44. [52]

    Kaminskas and D

    M. Kaminskas and D. Bridge. Diversity, serendipity, novelty, and coverage: A survey and empirical analysis of beyond-accuracy objectives in recommender systems. ACM Transactions on Interactive Intelligent Systems , 7(1): 2:1–2:42, 2016

  45. [53]

    Karahoca, E

    A. Karahoca, E. Bayraktar, E. Tatoglu, and D. Karahoca. Information system design for a hospital emergency department: A usability analysis of so/f_tware prototypes.Journal of Biomedical Informatics , 43(2):224 – 232, 2010

  46. [54]

    Katukuri, T

    J. Katukuri, T. K¨onik, R. Mukherjee, and S. Kolay. Recommending similar items in large-scale online marketplaces. In IEEE International Conference on Big Data 2014 , pages 868–876, 2014

  47. [55]

    Katukuri, T

    J. Katukuri, T. Konik, R. Mukherjee, and S. Kolay. Post-purchase recommendations in large-scale online marketplaces. In Proceedings of the 2015 IEEE International Conference on Big Data , Big Data ’15, pages 1299–1305, 2015

  48. [56]

    Kirshenbaum, G

    E. Kirshenbaum, G. Forman, and M. Dugan. A live comparison of methods for personalized article recommendation at forbes.com. In Machine Learning and Knowledge Discovery in Databases , pages 51–66, 2012

  49. [57]

    Kiseleva, M

    J. Kiseleva, M. J. M¨uller, L. Bernardi, C. Davis, I. Kovacek, M. S. Einarsen, J. Kamps, A. Tuzhilin, and D. Hiemstra. Where to Go on Your Next Trip? Optimizing Travel Destinations Based on User Preferences. In Proceedings of the 38th International ACM SIGIR Conference on Rese...

  50. [58]

    K¨ocher, M

    S. K¨ocher, M. Jugovac, D. Jannach, and H. Holzm¨uller. New hidden persuaders: An investigation of a/t_tribute-level anchoring effects of product recommendations. Journal of Retailing, 95:24–41, 2019

  51. [59]

    Kohavi, A

    R. Kohavi, A. Deng, B. Frasca, R. Longbotham, T. Walker, and Y. Xu. Trustworthy online controlled experiments: Five puzzling outcomes explained. In Proceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , KDD ’12, pages 786–794, 2012

  52. [60]

    J. A. Konstan and J. Riedl. Recommender systems: from algorithms to user experience.User Modeling and User-Adapted Interaction, 22(1):101–123, 2012

  53. [61]

    Lawrence, G

    R. Lawrence, G. Almasi, V. Kotlyar, M. Viveros, and S. Duri. Personalization of supermarket product recommendations. Data Mining and Knowledge Discovery , 5(1):11–32, 2001. ACM Transactions on Management Information Systems, Vol. 10, No. 4, Article 1. Publication date: Decembe...

  54. [62]

    Lee and K

    D. Lee and K. Hosanagar. Impact of recommender systems on sales volume and diversity. In Proceedings of the 2014 International Conference on Information Systems , ICIS ’14, 2014

  55. [63]

    Lee and K

    D. Lee and K. Hosanagar. How Do Recommender Systems Affect Sales Diversity? A Cross-Category Investigation via Randomized Field Experiment. Information Systems Research, 30(1):239–259, 2019

  56. [64]

    Lerche, D

    L. Lerche, D. Jannach, and M. Ludewig. On the value of reminders within e-commerce recommendations. In Proceedings of the 2016 ACM Conference on User Modeling, Adaptation, and Personalization , UMAP ’16, pages 27–35, 2016

  57. [65]

    L. Li, W. Chu, J. Langford, and X. Wang. Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms. In Proceedings of the Fourth ACM International Conference on Web Search and Data Mining, WSDM ’11, pages 297–306, 2011

  58. [66]

    J. Lin. /T_he neural hype and comparisons against weak baselines. SIGIR Forum, 52(2):40–51, Jan. 2019

  59. [67]

    Z. C. Lipton and J. Steinhardt. Troubling Trends in Machine Learning Scholarship. ArXiv e-prints, 2018. URL h/t_tps://arxiv.org/abs/1807.03341

  60. [68]

    J. Liu, P. Dolan, and E. R. Pedersen. Personalized news recommendation based on click behavior. In Proceedings of the 15th International Conference on Intelligent User Interfaces , IUI ’10, pages 31–40, 2010

  61. [69]

    Ludewig and D

    M. Ludewig and D. Jannach. Evaluation of session-based recommendation algorithms.User-Modeling and User-Adapted Interaction, 28(4–5):331–390, 2018

  62. [70]

    Ludewig and D

    M. Ludewig and D. Jannach. User-centric evaluation of session-based recommendations for an automated radio station. In Proceedings of the 2019 ACM Conference on Recommender Systems (RecSys ’19) , 2019

  63. [71]

    Maksai, F

    A. Maksai, F. Garcin, and B. Faltings. Predicting online performance of news recommender systems through richer evaluation metrics. In Proceedings of the 9th ACM Conference on Recommender Systems , RecSys ’15, pages 179–186, 2015

  64. [72]

    McKinney, Y

    V. McKinney, Y. Kanghyun, and F. M. Zahedi. /T_he measurement of web-customer satisfaction: An expectation and discon/f_irmation approach.Information Systems Research, 13(3):296 – 315, 2002

  65. [73]

    S. M. McNee, I. Albert, D. Cosley, P. Gopalkrishnan, S. K. Lam, A. M. Rashid, J. A. Konstan, and J. Riedl. On the recommending of citations for research papers. In Proceedings of the 2002 ACM Conference on Computer Supported Cooperative Work, CSCW ’02, pages 116–125, 2002

  66. [74]

    S. M. McNee, J. Riedl, and J. A. Konstan. Being accurate is not enough: How accuracy metrics have hurt recommender systems. In CHI ’06 Extended Abstracts on Human Factors in Computing Systems , CHI EA ’06, pages 1097–1101, 2006

  67. [75]

    A. More, L. Baltrunas, N. Vlassis, and J. Basilico. Recap: Designing a more efficient estimator for off-policy evaluation in bandits with large action spaces. In Proceedings of the ACM RecSys 2019 Workshop on Reinforcement and Robust Estimators for Recommendation (REVEAL ’19), 2019

  68. [76]

    Nilashi, D

    M. Nilashi, D. Jannach, O. bin Ibrahim, M. D. Esfahani, and H. Ahmadi. Recommendation quality, transparency, and website quality for trust-building in recommendation agents. Electronic Commerce Research and Applications , 19: 70–84, 2016

  69. [77]

    Nunes and D

    I. Nunes and D. Jannach. A systematic review and taxonomy of explanations in decision support and recommender systems. User-Modeling and User-Adapted Interaction, 27(3–5):393–444, 2017

  70. [78]

    L. Qin, S. Chen, and X. Zhu. Contextual combinatorial bandit and its application on diversi/f_ied online recommendation. In Proceedings of the 2014 SIAM International Conference on Data Mining (SDM ’14) , pages 461–469, 2014

  71. [79]

    /Q_uadrana, P

    M. /Q_uadrana, P. Cremonesi, and D. Jannach. Sequence-aware recommender systems.ACM Computing Surveys, 54: 1–36, 2018

  72. [80]

    Radlinski, R

    F. Radlinski, R. Kleinberg, and T. Joachims. Learning diverse rankings with multi-armed bandits. In Proceedings of the 25th International Conference on Machine Learning , ICML ’08, pages 784–791, 2008

  73. [81]

    Rahman and J

    M. Rahman and J. C. Oh. Graph bandit for diverse user coverage in online recommendation. Applied Intelligence, 48 (8):1979–1995, 2018

  74. [82]

    Rendle, C

    S. Rendle, C. Freudenthaler, Z. Gantner, and L. Schmidt-/T_hieme. BPR: Bayesian personalized ranking from implicit feedback. In Proceedings of the Twenty-Fi/f_th Conference on Uncertainty in Arti/f_icial Intelligence, UAI ’09, pages 452–461, 2009

  75. [83]

    Rodriguez, C

    M. Rodriguez, C. Posse, and E. Zhang. Multiple objective optimization in recommender systems. In Proceedings of the Sixth ACM Conference on Recommender Systems , RecSys ’12, pages 11–18, 2012

  76. [84]

    Rosse/t_ti, F

    M. Rosse/t_ti, F. Stella, and M. Zanker. Contrasting offline and online results when evaluating recommendation algorithms. In Proceedings of the 10th ACM Conference on Recommender Systems , RecSys ’16, pages 31–34, 2016

  77. [85]

    A. Said, D. Tikk, Y. Shi, M. Larson, K. Stumpf, and P. Cremonesi. Recommender Systems Evaluation: A 3D Benchmark. In Proceedings of the 2012 Workshop on Recommendation Utility Evaluation: Beyond RMSE (RUE) at RecSys ’12 , pages 21/f_i/question_exclam–23, 2012

  78. [86]

    Shani, D

    G. Shani, D. Heckerman, and R. I. Brafman. An MDP-Based Recommender System. Journal of Machine Learning Research, 6:1265–1295, 2005. ACM Transactions on Management Information Systems, Vol. 10, No. 4, Article 1. Publication date: December 2019. 1:22 Dietmar Jannach and Michael Jugovac

  79. [87]

    Shin and W.-Y

    D.-H. Shin and W.-Y. Kim. Applying the technology acceptance model and /f_low theory to cyworld user behavior: Implication of the web2.0 user acceptance. CyberPsychology & Behavior, 11(3):378–382, 2008

  80. [88]

    Smyth, P

    B. Smyth, P. Co/t_ter, and S. Oman. Enabling intelligent content discovery on the mobile internet. InProceedings of the Twenty-Second AAAI Conference on Arti/f_icial Intelligence, AAAI ’07, pages 1744–1751, 2007

  81. [89]

    Spertus, M

    E. Spertus, M. Sahami, and O. Buyukkokten. Evaluating similarity measures: A large-scale study in the orkut social network. In Proceedings of the Eleventh ACM SIGKDD International Conference on Knowledge Discovery in Data Mining , KDD ’05, pages 678–684, 2005

  82. [90]

    H. Steck. Training and testing of recommender systems on data missing not at random. In Proceedings of the 16th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , KDD ’10, pages 713–722, 2010

  83. [91]

    Szpektor, Y

    I. Szpektor, Y. Maarek, and D. Pelleg. When relevance is not enough: Promoting diversity and freshness in personalized question recommendation. In Proceedings of the 22nd International Conference on World Wide Web , WWW ’13, pages 1249–1260, 2013

  84. [92]

    K. Y. Tam and S. Y. Ho. Web personalization as a persuasion strategy: An elaboration likelihood model perspective. Information Systems Research, 16(3):271–291, 2005

  85. [93]

    Tintarev and J

    N. Tintarev and J. Masthoff. Evaluating the effectiveness of explanations for recommender systems. User Modeling and User-Adapted Interaction, 22(4):399–439, 2012

  86. [94]

    Vargas and P

    S. Vargas and P. Castells. Rank and relevance in novelty and diversity metrics for recommender systems. InProceedings of the Fi/f_th ACM Conference on Recommender Systems, RecSys ’11, pages 109–116, 2011

  87. [95]

    K. Wagstaff. Machine learning that ma/t_ters. InProceedings of the Twenty-Ninth International Conference on Machine Learning, ICML ’12, pages 529–536, 2012

  88. [96]

    Wobcke, A

    W. Wobcke, A. Krzywicki, Y. Sok, X. Cai, M. Bain, P. Compton, and A. Mahidadia. A deployed people-to-people recommender system in online dating. AI Magazine, 36(3):5–18, 2015

  89. [97]

    Xiao and I

    B. Xiao and I. Benbasat. E-commerce product recommendation agents: Use, characteristics, and impact.MIS /Q_uarterly, 31(1):137–209, 2007

  90. [98]

    Y. Xu, Z. Li, A. Gupta, A. Bugdayci, and A. Bhasin. Modeling professional similarity by mining professional career trajectories. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’14, pages 1945–1954, 2014

  91. [99]

    J. Yi, Y. Chen, J. Li, S. Se/t_t, and T. W. Yan. Predictive model performance: Offline and online evaluations. InProceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’13, pages 1294–1302, 2013

  92. [100]

    K.-H. Yoo, U. Gretzel, and M. Zanker. Persuasive Recommender Systems: Conceptual Background and Implications . Springer, 2012

  93. [101]

    K.-H. Yoo, U. Gretzel, and M. Zanker. Persuasive Recommender Systems – Conceptual Background and Implications . Springer, 2013

  94. [102]

    Zanker, M

    M. Zanker, M. Bricman, S. Gordea, D. Jannach, and M. Jessenitschnig. Persuasive online-selling in quality and taste domains. In Proceedings of the 7th International Conference on E-Commerce and Web Technologies , EC-Web ’06, pages 51–60, 2006

  95. [103]

    Zanker, M

    M. Zanker, M. Fuchs, W. H¨opken, M. Tuta, and N. M¨uller. Evaluating recommender systems in tourism — A case study from Austria. In P. O’Connor, W. H¨opken, and U. Gretzel, editors, Proceedings of the 2008 International Conference on Information and Communication Technologies ...

  96. [104]

    Zhang and N

    M. Zhang and N. Hurley. Avoiding monotony: Improving the diversity of recommendation lists. In Proceedings of the 2008 ACM Conference on Recommender Systems , RecSys ’08, pages 123–130, 2008

  97. [105]

    Zheng, D

    H. Zheng, D. Wang, Q. Zhang, H. Li, and T. Yang. Do clicks measure recommendation relevancy?: An empirical user study. In Proceedings of the Fourth ACM Conference on Recommender Systems , RecSys ’10, pages 249–252, 2010

  98. [106]

    T. Zhou, Z. Kuscsik, J.-G. Liu, M. Medo, J. R. Wakeling, and Y.-C. Zhang. Solving the apparent diversity-accuracy dilemma of recommender systems. Proceedings of the National Academy of Sciences , 107(10):4511–4515, 2010

  99. [107]

    Ziegler, S

    C.-N. Ziegler, S. M. McNee, J. A. Konstan, and G. Lausen. Improving recommendation lists through topic diversi/f_ication. In Proceedings of the 14th International Conference on World Wide Web , WWW ’05, pages 22–32, 2005. ACM Transactions on Management Information Systems, Vol...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.