{"id":"f34d5997-9205-4997-abfc-4632fbd00893","arxiv_id":"1908.08328","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of recommender deployments shows that business value is measured inconsistently and that offline accuracy gains frequently do not translate into online business impact.","lead":"This paper reviews real-world field tests of recommender systems and finds that reported business effects range from fractions of a percent to several-fold gains depending on the measure and baseline. It argues that click-through rates and offline accuracy metrics often fail to predict actual business value.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Offline-accuracy claim rests on a small convenience sample, not a systematic synthesis","rationale":"The central claim is deliberately modest, and much of the survey's qualitative synthesis is sound. My concern targets the second half of the claim, namely the assessment of recommendation algorithms in academia. Section 3.4.1's assertion that offline accuracy does not predict online success is based on an informal tally of a handful of cases rather than a systematic review; the authors do not report a search strategy, inclusion criteria, or a protocol for extracting 'online success.' Section 3.3 concedes that the underlying field tests themselves usually lack power analyses, so the base evidence is weak. The paper partially hedges by saying the correspondence 'cannot be assumed in general,' which is safe, but the abstract's claim about 'various open questions' in performance assessment and the implications section lean on the stronger reading. A fair stress-test should therefore check whether the claimed negative result holds under a systematic synthesis. I agree with the reader's concern about non-systematic selection, though I locate the most load-bearing version in the offline-proxy argument rather than in the representativeness of industry reports alone. This does not invalidate the paper; it is a reason to keep the caveats in place, so the existing CONDITIONAL verdict stands unchanged.","tokens_in":24189,"tokens_out":7619,"duration_ms":78542,"concrete_test":"Conduct a pre-registered systematic review with explicit inclusion criteria for studies reporting both an offline accuracy metric and an online business or user-perceived quality metric in the same setting. Extract the direction and magnitude of association for each study. If a majority of well-powered, pre-registered comparisons show a positive association, the survey's 'not reliable proxies' conclusion should be downgraded to 'unproven in general'; if the association is null or mixed across settings, the claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's most consequential claim, that offline accuracy metrics are unreliable proxies for business value, is supported in Section 3.4.1 by a narrative count over roughly eight comparisons ([9,19,28,30,40,71,73,84]) against three positive cases ([14,20,50]). No search protocol, inclusion criteria, or effect-size synthesis is provided, and Section 3.3 concedes that 'in almost all surveyed cases, an analysis of the required sample size and detailed statistical analyses of the A/B tests were missing.' The statement that 'the most accurate offline models did neither lead to the best online success nor to a better accuracy perception' is therefore a heuristic over a heterogeneous, likely publication-biased sample rather than an established empirical regularity. The paper's own hedge, that the correspondence 'cannot be assumed in general,' is defensible, but the abstract and the discussion lean on a stronger reading of this thin evidence. If the sample is unrepresentative, the 'performance assessment of recommendation algorithms in academia' part of the central claim is overstated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a research commentary that reviews field studies (A/B tests and other real-world deployments) of recommender systems, with the goal of synthesizing what is known about the business value of recommenders. It organizes the reported effects into five measurement categories: click-through rates, adoption/conversion rates, sales and revenue, effects on sales distributions, and user engagement/behavior. It then discusses challenges of measuring business value, pitfalls of A/B testing, and the degree to which offline accuracy metrics such as RMSE, precision, and recall predict online success or business outcomes. The paper concludes that recommender systems can create substantial but highly variable business value, that direct revenue effects are often in the low single-digit percentage range while indirect engagement effects can be much larger, and that important open questions remain about both the realistic quantification of business effects and the use of offline accuracy metrics as proxies for business value.","tokens_in":24353,"tokens_out":4506,"duration_ms":50505,"significance":"If taken with appropriate caveats, the paper performs a useful service: it compiles a scattered literature of industrial field tests, structures the measurement space, and highlights methodological weaknesses that are often underappreciated when offline benchmark results are interpreted. The authors are commendably explicit about several limitations, including the lack of statistical detail in most surveyed A/B tests (Section 3.3) and the difficulty of extrapolating offline accuracy gains to business impact (Section 3.4.1). The distinction between direct revenue measures and indirect engagement measures is a valuable organizing framework for both practitioners and researchers. However, the review is deliberately non-systematic: there is no search protocol, no inclusion criteria, and no quantitative synthesis. The most consequential claim, that offline accuracy is not a reliable proxy for online/business success, rests on a small and likely publication-biased convenience sample. The paper would benefit from more carefully calibrated language in the abstract and conclusion so that its useful qualitative message is not overstated.","major_comments":[{"comment":"The sentence 'in the majority of these attempts, the most accurate offline models did neither lead to the best online success nor to a better accuracy perception' is supported only by a count of eight negative comparisons ([9,19,28,30,40,71,73,84]) against three positive ones ([14,20,50]). No inclusion criteria, search protocol, or effect-size synthesis are provided, and Section 3.3 itself concedes that 'in almost all surveyed cases, an analysis of the required sample size and detailed statistical analyses of the A/B tests were missing.' The authors should revise this passage to say explicitly that the statement is an observation about the specific studies they identified, not about the broader population of recommender deployments, and that publication bias and the heterogeneity of baselines and metrics preclude any general quantitative conclusion. The abstract and Section 5 should be aligned with this weaker, defensible claim rather than suggesting that the unreliability of offline accuracy is an established empirical regularity.","section":"Section 3.4.1"},{"comment":"The review reports many large headline effect sizes, such as a 200% CTR increase at YouTube, a 500% increase in Gross Merchandise Bought at eBay, and Amazon's 35% cross-sales figure, without explicitly stating that these are self-selected, publicly disclosed outcomes and therefore likely an upper-bound sample rather than a representative distribution. Section 4.1's statement that direct revenue increases 'are more often reported to lie between one and five percent' is an informal impression, not a computed summary. The authors should add a visible caveat, perhaps at the end of Section 2.1 and repeated in Section 4.1, that the compiled figures are illustrative and that the absence of unpublished negative results, differences in baselines, and varying test durations mean that no typical or expected business-value figure can be reliably estimated from this material.","section":"Sections 2 and 4.1"}],"minor_comments":[{"comment":"The phrase 'increases in sales between one and five percent are reported on average' is ambiguous; the authors should clarify whether this is an informal range, a median, or a rough summary, and should point to the specific studies from which the range is derived.","section":"Section 3.2"},{"comment":"The sentence 'In almost all surveyed cases, an analysis of the required sample size and detailed statistical analyses of the A/B tests were missing' has a subject-verb agreement issue ('analysis' is singular); consider rewriting as '...was missing.'","section":"Section 3.3"},{"comment":"When describing the Google News results in [68], the authors note that improved recommendations 'stole' clicks from other parts of the page; this is an important qualification that deserves a corresponding mention in the Section 3 discussion of CTR as a business measure, not only in the descriptive part.","section":"Section 2.2.1"},{"comment":"The discussion of the LinkedIn skill recommendation field test [8] appropriately notes the confound between the recommendation method and the user interface change, but the same caution is not applied consistently to other field tests in Section 2 where UI placement or presentation may have changed; a brief general remark in Section 3.1 about this confound would be helpful.","section":"Section 2.2.2"},{"comment":"Several sources are blog posts, industry white papers, or non-archival reports; given that the paper is a review, the authors should state in a short methodological paragraph what types of sources were considered admissible and how they were located.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a commentary rather than a systematic review, and the editors may wish to make sure the final version makes that scope explicit in the abstract. The core qualitative conclusions are plausible and well aligned with the cited literature, but the strongest claim about offline accuracy as an unreliable proxy needs more careful framing before publication. With the requested re-scoping and added caveats, the paper would be a useful contribution to the TMIS audience."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a literature review and research commentary on measuring the business value of recommender systems. The new thing here is the synthesis: the authors collect a broad set of field-test reports from industry and academia, organize the measured outcomes into five families (CTR, adoption/conversion, sales/revenue, sales-distribution effects, engagement), and use that map to argue two things: business value is real but poorly quantified, and offline accuracy metrics are not a safe proxy for it. That is a useful service, and the paper does it honestly. It repeatedly flags the statistical weaknesses of the underlying A/B tests, notes that reported effects vary wildly depending on baseline and domain, and reminds readers that UI changes can matter more than algorithm swaps. The taxonomy itself is tidy and likely to be cited.\n\nThe main soft spot, as you might expect from a review, is the evidence base for the strongest claim. The assertion that the most accurate offline models usually fail to win online is supported by a narrative count of roughly eight comparisons against three positive cases. There is no systematic search protocol, no inclusion criteria, no effect-size synthesis. The paper even concedes in Section 3.3 that almost all surveyed field tests lacked proper sample-size analysis. That is enough to make the claim a reasonable heuristic, but not an established empirical regularity. The authors do hedge in Section 3.4.1—\"cannot be assumed in general\"—but the abstract and discussion lean on the stronger reading. So the stress-test note is fair: the offline-accuracy claim is softer than it first appears.\n\nA smaller quibble: the paper says it could not identify any satisfaction or UX surveys in deployed systems. That may be true of the papers they searched, but absence in this selected sample is weak evidence for absence in industry. The authors acknowledge companies are unlikely to publish such results, so the observation is more a gap in the public record than a finding about practice.\n\nOverall, this is a solid, readable survey with a proportionate caveat about its own foundations. It does not introduce new data or methods, but it doesn't need to. Practitioners and information-systems researchers will get a clear map of business metrics and a sensible warning about offline evaluation. It deserves a serious referee, and with a bit of methodological transparency in the selection process and a softer framing of the offline-accuracy claim, it would be a nice contribution.\n\nMy recommendation: send it out for review, not desk reject.","headline":"A useful, honest survey of field tests on business value; the core claim about offline accuracy is plausible but rests on thin evidence.","tokens_in":24813,"tokens_out":2453,"would_cite":true,"duration_ms":25290,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This research commentary argues that recommender systems generate real, sometimes large business value, but the range of measured effects is so wide and the statistics so weak that neither the realistic size of that value nor the…","keywords":["recommender systems","business value","field tests","A/B testing","offline evaluation","click-through rate","conversion rate","sales impact"],"falsifier":"Collect a preregistered sample of recommender A/B tests that includes unpublished and failed tests; if those tests show effects centered near zero, the review's implied typical payoff is overstated - or, conversely, a large paired offline/online benchmark where offline accuracy consistently predicts online revenue would refute the claim that no such correspondence can be assumed.","tokens_in":1556,"feed_emoji":"📈","tokens_out":2907,"duration_ms":89904,"temperature":0.7,"pith_summary":"Recommender systems clearly influence what people click, buy, and watch, but the size of their business payoff is not well understood. Reviewing published field tests from e-commerce, news, streaming, dating, and job portals, the authors find reported effects that range from marginal revenue gains around 0.3% to massive lifts of several hundred percent in gross merchandise value. They argue that popular offline accuracy measures such as RMSE or precision are not reliable stand-ins for business value, because in most studies the most accurate algorithms did not win online field tests. The consequence matters for anyone deciding whether to invest in better algorithms versus simpler presentation changes, since one field test doubled clicks just by changing the position and size of the recommendation widget. The paper's core message is that both the realistic quantification of business effects and the academic performance assessment of recommenders remain open problems.","feed_headline":"Recommender payoffs range from 0.3% to 500% in field tests","feed_subtitle":"A review of real deployments finds the payoff is real, poorly measured, and often unrelated to offline accuracy.","key_machinery":"The analytical engine of the paper is a measurement taxonomy that sorts business value into five categories: click-through rates, adoption and conversion rates, sales and revenue, effects on sales distributions, and user engagement and behavior. Each category implies a different conclusion about value: click-through rates are easy to measure but can reward clickbait; adoption measures are domain-specific; direct revenue is the most informative but often unavailable; sales-distribution shifts can increase or decrease diversity; and engagement is only a proxy for retention. The taxonomy does the work of showing why reported numbers across studies are not comparable and why no single offline metric has yet been tied reliably to business outcomes.","core_discovery":"On the paper's own terms, the central discovery is a documented gap: the surveyed industry reports show recommenders can create substantial business value in many ways - higher click-through, conversion, sales, engagement - but the reported magnitude varies enormously and is often measured with weak statistics or indirect proxies. In studies that compare algorithms both offline and online, the majority found that the most accurate offline models did not produce the best online results, and in a few cases a simple repositioning of the recommendation widget outperformed any algorithmic change. The authors therefore conclude that the business value of recommenders is real yet poorly quantified, and that offline accuracy metrics cannot be assumed to predict business value without per-case validation.","pith_inferences":["If unpublished or failed A/B tests are systemically underrepresented in the literature, the true average business lift of recommenders is probably below the one-to-five percent direct-revenue range the paper compiles.","The paper's taxonomy implies a testable prediction: direct measures like revenue should be more stable across studies than indirect measures like click-through rate, because indirect measures are more sensitive to interface and presentation effects.","A concrete next step the paper does not spell out is a public paired benchmark where logged interactions and A/B outcomes come from the same deployment, allowing offline metrics to be calibrated against business value instead of assumed."],"forward_implications":["Companies using recommenders can expect positive effects on user behavior in many domains, but the expected magnitude depends heavily on the baseline, the domain, and the chosen measurement.","A typical direct revenue lift reported in field tests is between one and five percent, which is still substantial in absolute terms for large businesses.","An algorithm that wins on offline accuracy such as RMSE should not be assumed to win online; most published comparisons found no such correspondence, and some found the opposite.","User interface and presentation choices can dominate algorithmic improvements, with one field test doubling click-through rate just by changing widget position and size.","Nearly all surveyed A/B tests lacked sample-size analyses and detailed statistics, so the published effect sizes may not be reliable without better reporting."],"supporting_citations":[{"why":"A video streaming service's own account of its recommender architecture and business value; it is the source for the claim that offline experiments are not highly predictive of A/B outcomes and that personalization is worth over one billion dollars per year.","marker":"[32]"},{"why":"Google News personalization field test; it contributes the 38% click lift over a popularity baseline and the observation that no personalized algorithm clearly won.","marker":"[22]"},{"why":"YouTube's system description; it supplies the 60% of home-screen clicks and the over 200% click-through gain over a most-viewed baseline.","marker":"[23]"},{"why":"eBay field test of ephemeral-item recommendations; it reports a near-500% increase in gross merchandise bought, the largest effect in the review.","marker":"[18]"},{"why":"swissinfo.ch news recommender study; it compares offline and online evaluation and shows a UI change doubled click-through rate, supporting the claim that presentation can outweigh algorithms.","marker":"[30]"},{"why":"Randomized field experiment on an online retailer; it shows recommenders reduce aggregate sales diversity, grounding the sales-distribution discussion.","marker":"[63]"},{"why":"Comparison of offline, online, and user-study evaluations of research-paper recommenders; it underpins the finding that the most accurate offline models did not achieve the best online results.","marker":"[9]"},{"why":"Survey of online controlled experiments at a major search engine company; it supplies the A/B test pitfalls, including a bug where degraded search quality temporarily raised query counts and warnings about sample size.","marker":"[59]"},{"why":"Field experiment on DVD sales; it shows purchase-based collaborative filtering lifted sales 35% versus no recommendations, while view-based strategies did not.","marker":"[62]"}],"fun_headline_variants":["Accuracy ≠ profit for recommender systems","Recommenders profit, but offline metrics mislead","Field tests: recommender ROI real, stats weak","Business value of recommenders: proven, unmeasured","Recommender payoffs vary wildly, offline tests fail"],"cache_read_input_tokens":27136,"weakest_assumption_plain":"The load-bearing premise is that the published industry reports and blog posts are honest and representative of real recommender deployments; if unpublished negative results are common or the publicized numbers are cherry-picked, the reported range of effects would overstate typical business value.","fun_headline_variants_meta":{"raw":{"variants":["Accuracy ≠ profit for recommender systems","Recommenders profit, but offline metrics mislead","Field tests: recommender ROI real, stats weak","Business value of recommenders: proven, unmeasured","Recommender payoffs vary wildly, offline tests fail"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000131,"raw_usage":{"total_tokens":1068,"prompt_tokens":825,"completion_tokens":243,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":171}},"tokens_in":441,"tokens_out":243,"duration_ms":3140,"temperature":1.0,"reasoning_tokens":171,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:42:13.866324+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a preregistered sample of recommender A/B tests that includes unpublished and failed tests; if those tests show effects centered near zero, the review's implied typical payoff is overstated - or, conversely, a large paired offline/online benchmark where offline accuracy consistently predicts online revenue would refute the claim that no such correspondence can be assumed.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A video streaming service's own account of its recommender architecture and business value; it is the source for the claim that offline experiments are not highly predictive of A/B outcomes and that personalization is worth over one billion dollars per year."},{"cited_title":"Garcin, B","cited_arxiv_id":null,"evidence_quote":"swissinfo.ch news recommender study; it compares offline and online evaluation and shows a UI change doubled click-through rate, supporting the claim that presentation can outweigh algorithms."},{"cited_title":"Lee and K","cited_arxiv_id":null,"evidence_quote":"Randomized field experiment on an online retailer; it shows recommenders reduce aggregate sales diversity, grounding the sales-distribution discussion."},{"cited_title":"Kohavi, A","cited_arxiv_id":null,"evidence_quote":"Survey of online controlled experiments at a major search engine company; it supplies the A/B test pitfalls, including a bug where degraded search quality temporarily raised query counts and warnings about sample size."},{"cited_title":"Lee and K","cited_arxiv_id":null,"evidence_quote":"Field experiment on DVD sales; it shows purchase-based collaborative filtering lifted sales 35% versus no recommendations, while view-based strategies did not."}],"review_version":1}