{"id":"4bb1dfad-1ce4-4a50-9a6b-a9033c159b65","arxiv_id":"2509.06002","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of A/B-validated industrial recommender systems, split into transaction-oriented and content-oriented categories, with a discussion of the academia-industry gap.","lead":"This survey reviews 228 publications on recommender systems that were validated with real-world A/B tests, organizing them into transaction-oriented and content-oriented categories. It contrasts these industrial systems with academic research, highlighting how constraints like latency, cost, and multi-objective business goals change the design of recommenders.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"A/B-testing filter may bias the survey corpus toward large platforms, undercutting the claimed generality of the transaction/content taxonomy.","rationale":"The reader's weakest assumption is exactly the load-bearing concern. The survey's central contribution is a conceptual classification and a synthesis of industrial practice. For that classification to be credible as a map of real-world recommender systems, the corpus must represent industrial practice broadly. The A/B-testing filter directly determines which papers enter the corpus and which domains are covered; the paper itself admits that several real-world domains are excluded because of this criterion. This makes the filter both restrictive and self-confirming: it selects for larger platforms that can run and publish online experiments, then uses the resulting corpus to draw conclusions about 'industry' versus 'academia.' The concern is not that the taxonomy is wrong or internally inconsistent—it is plausible and the narrative is coherent—but that the evidence presented does not yet establish its generality. A sensitivity analysis that relaxes the A/B requirement would reveal whether the excluded domains and non-A/B methods would alter the taxonomy or its scope. Until then, conditional acceptance with a request for corpus transparency and sensitivity analysis is the right stance. The paper's extensive reference list and the self-flagged limitation in §1.3 are useful, but they do not resolve the selection-bias risk. No change to the reader's CONDITIONAL verdict is needed, because the same condition—releasing the corpus and reconciling the selection rule—would address this concern.","tokens_in":32959,"tokens_out":5985,"duration_ms":71336,"concrete_test":"Release the full list of the 228 papers and the counts after each filter in §1.2, then re-run the classification on the superset obtained by dropping the online-A/B-testing requirement while keeping the industry-authorship and 2020–2024 constraints. If the superset contains substantial numbers of papers in the excluded domains (POI, games, live streaming, social, recruitment) or methods validated only by offline/large-scale simulation, the corpus is not representative and the taxonomy's claimed generality is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a new classification—Transaction-Oriented vs Content-Oriented RecSys—grounded in item characteristics and recommendation objectives captures real-world industrial recommender systems. The evidence base for this claim is the 228-paper corpus selected in §1.2, which requires (a) at least one industry author and (b) validation through online A/B testing. This filter is not neutral. It selects for organizations large and mature enough to run and publish online experiments, and it systematically excludes industrial recommender systems in domains the survey itself names in §1.3—POI, games, live streaming, social community, and job opportunities—because too few A/B-validated papers exist there. The paper thus cannot support the claim that the taxonomy spans 'real-world recommender systems' at large; at best it describes A/B-validated practice in e-commerce, video, news, and audio at large platforms. Moreover, §5.4 acknowledges that much industrial cost/performance work is 'rarely publicly available,' so the A/B requirement may be as much a publication-pipeline artifact as a property of industrial practice. The taxonomy may be plausible, but the corpus does not establish its generality.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys industrial recommender systems research published 2020–2024, selecting papers from major conferences with at least one industry author and online A/B validation (272 papers, 228 after excluding advertising). It introduces a taxonomy with two main classes—Transaction-Oriented RecSys (e-commerce, travel, food delivery, insurance) and Content-Oriented RecSys (video, news, audio)—reviews solution themes per class, contrasts academic and industrial evaluation and constraints, and proposes future research directions centered on user decision-making, theory-guided multi-objective optimization, and realistic problem definitions.","tokens_in":33239,"tokens_out":6226,"duration_ms":72466,"significance":"The survey addresses a real gap: much academic RecSys research is evaluated offline, and there is little organized synthesis of industrial practice. Its strengths are a recent, clearly delimited corpus, a simple and teachable taxonomy, and an unusually explicit set of caveats about coverage (POI, games, live streaming, social, jobs are excluded) and about the scarcity of public cost/performance work. If the corpus is representative and the taxonomy is reliably applied, the paper offers a useful map for academics entering industrial research. The main contribution is conceptual and descriptive rather than empirical; its value depends on corpus transparency and on the strength of the selection rules.","major_comments":[{"comment":"The corpus selection is under-specified, and this is load-bearing for the survey's claims. The paper reports that the filters produced 272 papers, and 228 after excluding advertising, but gives no list of the included papers, no per-domain or per-venue counts, and no PRISMA-style flow diagram. Since every section's synthesis is built on this corpus, readers cannot verify that the taxonomy was applied consistently, or assess sampling bias. Please provide a complete list or searchable appendix of the 228 papers, the number excluded at each step (no industry author, no A/B validation, advertising), and a breakdown across the two proposed classes and the three content subdomains.","section":"§1.2"},{"comment":"The A/B-testing filter biases the corpus toward large platforms and undercuts the unqualified 'real-world' framing. Requiring online A/B validation as 'a necessary component of applied recommender systems' selects for organizations with experimentation infrastructure. The paper itself acknowledges in §1.3 that POI, games, live streaming, social community, and job opportunities are excluded because too few A/B-validated papers exist, and §5.4 notes that much cost/performance work is 'rarely publicly available.' Thus the 228-paper corpus cannot support the abstract's claim that the taxonomy captures 'real-world recommender systems' at large; it describes A/B-validated practice in e-commerce, video, news, and audio at relatively large platforms. Please narrow the generality claims or include and analyze non-A/B industrial papers to test for selection bias.","section":"§1.2 and §1.3"},{"comment":"The 'new classification' is not operationalized as claimed. The two definitions are stated solely in terms of the system's primary objective (prompting transactional actions vs. facilitating consumption), but the abstract says the classification is 'grounded in item characteristics and recommendation objectives.' No item-characteristic dimensions are defined, and the later assignment rule—'the sole criterion ... is the business context' in which a method was A/B-tested—does not address hybrids (e.g., video platforms with purchases or subscriptions, e-commerce recommenders optimizing engagement). Without explicit coding rules, or at least a description of how borderline cases were resolved, the dichotomy risks being a post-hoc labeling scheme. Specify the item-characteristic axes, define boundary cases, and state whether the categorization was done independently.","section":"§1.3, Definitions 1 and 2"}],"minor_comments":[{"comment":"The phrase 'grounded in item characteristics and recommendation objectives' is stronger than what Definitions 1 and 2 support; align the abstract with the actual operationalization.","section":"Abstract"},{"comment":"The count progression is confusing: 272 papers are mentioned after the A/B filter, then 228 after excluding advertising. Clarify which number corresponds to which filtering stage, and report how many papers were excluded for advertising.","section":"§1.2"},{"comment":"The text says the real-world systems are classified into 'two main categories,' but Figure 1 includes an 'Other Recommendation Systems' box. Acknowledge explicitly that the dichotomy is not exhaustive and describe how the 'other' domains relate to the two classes.","section":"§1.3/Figure 1"},{"comment":"Reference [51] contains duplicated author names (e.g., Gong-Duo Zhang, Lihong Gu, Zhiqiang Zhang appear twice); clean up the reference list. Also, several arXiv preprints from 2025 are cited for a survey covering 2020–2024; please confirm they are part of the corpus or mark them as outlook material.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for cs.IR and the self-citations are appropriate. The main risk is corpus representativeness; if the authors supply the missing list and analysis, I would be willing to reconsider. No concerns about novelty disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. This is a genuinely useful survey of industrial recommender practice, organized around a transaction-oriented vs content-oriented split that I think will stick. The authors did something real: they read the A/B-validated papers from 2020–2024 and extracted the engineering constraints—latency, cost, real-time modeling, multi-objective trade-offs—that academic benchmarks usually omit. The two category definitions are clean, and the video/news/audio subsections in Section 4 are concrete and well-cited. They also own a real limitation up front: POI, games, live streaming, social, and jobs get excluded because too few A/B-validated papers exist there. Good.\n\nThe soft spots are mostly about the packaging. The 'systematic review' claim is undercut by a corpus that is not actually listed. 228 papers, no appendix, no per-filter counts. That makes it impossible to check selection bias or replicate the taxonomy assignments. Also, the stated 2020–2024 window leaks: several references (e.g., PinRec, OneRec, 360Brew) are 2025 preprints. Either expand the window or trim the text. And the A/B-testing filter is a real bias: it selects for large platforms with mature experimentation infra. The survey is best read as 'what large platforms have published' rather than 'industrial RecSys at large.' The stress-test note makes this point; I think it's right but not fatal, because the authors flag the excluded domains and the taxonomy doesn't depend on those areas being included.\n\nOne more thing: the forward-looking section on user decision-making (hesitation, tolerance) leans on the authors' own prior work, and the evidence is, by their own account, preliminary. That's fine as an opinion piece within a survey, but it should not be read as a corpus finding.\n\nBottom line: for a reader who wants a map of production RecSys constraints, this is a solid entry point. It deserves peer review provided the authors release the paper list and reconcile the timeline. I would not cite it in my own work unless I were doing a survey myself, but I would send it to a student who asks 'what do people actually deal with in industry?'","headline":"Useful survey of industrial RecSys practice with a plausible transaction/content taxonomy; needs corpus transparency before I'd call it systematic.","tokens_in":33671,"tokens_out":2287,"would_cite":false,"duration_ms":25909,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Industrial recommender systems split into two goal-driven classes—transaction-oriented and content-oriented—and the split exposes constraints that academic offline research mostly ignores.","keywords":["industrial recommender systems","transaction-oriented recommendation","content-oriented recommendation","A/B testing","multi-objective optimization","real-time recommendation","cold-start","generative recommendation"],"falsifier":"Collect all industry-authored recommender-system papers at the same six venues from 2020 to 2024 that report deployment but no online A/B testing; if a large share of those papers exhibit the same latency, cost, and multi-objective constraints as the A/B-validated set, the survey's selection criterion misses a substantial part of industrial practice.","tokens_in":32873,"feed_emoji":"🎯","tokens_out":3343,"duration_ms":40931,"temperature":0.7,"pith_summary":"The survey claims that real-world recommender systems are best understood through a dichotomy: transaction-oriented systems (e-commerce, travel, food delivery) that optimize for conversion, revenue, or GMV, and content-oriented systems (video, news, audio) that optimize for engagement, dwell time, and satisfaction. Based on 228 papers from 2020–2024 with industry authorship and online A/B validation, the authors argue that this split reveals challenges academia usually misses: serving latency, system cost, real-time interest shifts, multi-objective trade-offs, and long-term user value. If the classification holds, it gives academic researchers a map of what actually matters in production and points toward under-explored research directions such as user decision-making and theory-guided multi-objective optimization.","feed_headline":"228 A/B-tested systems reveal two tribes of recommender","feed_subtitle":"Transaction and content recommenders face different constraints—latency, cost, and real-time shifts that offline research misses.","key_machinery":"The central object is the dual classification of recommender systems by business objective: Transaction-Oriented RecSys versus Content-Oriented RecSys. This dichotomy, paired with a strict inclusion filter (industry authorship plus online A/B testing), organizes the entire survey and drives the claim that distinct item characteristics produce distinct industrial challenges. The multi-stage production pipeline—recall, coarse ranking, fine ranking, re-ranking—serves as the recurring industrial mechanism that balances effectiveness with latency and cost.","core_discovery":"On its own terms, the paper establishes a new organizing framework for industrial recommender systems rather than a new algorithm or dataset. It defines Transaction-Oriented RecSys as systems whose primary goal is prompting transactional actions—conversion, revenue, purchase likelihood—and Content-Oriented RecSys as systems whose primary goal is facilitating consumption and engagement, measured by dwell time, clicks, or satisfaction. It then reviews 228 A/B-validated industrial papers from major venues (2020–2024) under this split, showing that the two classes face different data characteristics, real-time requirements, evaluation metrics, and cost constraints. The survey further claims that","pith_inferences":["The transaction/content split could plausibly extend to other A/B-validated domains not covered by the corpus, such as job marketplaces, point-of-interest services, and live streaming, once enough production studies accumulate.","Relaxing the A/B-testing inclusion filter might substantially change the map: companies that cannot run large-scale online experiments still operate production systems, so the survey's industrial picture is biased toward firms with mature experimentation infrastructure.","The survey's call for offline–online consistency suggests a concrete test: building shared simulation environments benchmarked against known A/B outcomes could let academic groups validate deployment-relevant claims without industrial infrastructure.","User decision-making states—hesitation before clicks, tolerance after consumption—are named as underused signals; modeling them could turn psychological constructs into measurable features for long-term retention optimization."],"forward_implications":["If the dichotomy is accurate, academic results measured only by precision, recall, or NDCG cannot be assumed to transfer to production, because the optimization target differs by system type.","Cost and latency become first-class objectives: methods that ignore embedding-layer resource consumption, inference time, or serving budgets will be impractical even if offline metrics improve.","Real-time interest modeling—responding to triggers, scrolling, inventory, and price changes within milliseconds—becomes a core research problem rather than an engineering footnote.","Multi-objective optimization is the norm in industry; balancing CTR, CVR, GMV, dwell time, and long-term retention requires frameworks beyond single-metric tuning.","Generative and foundation-model recommender systems are emerging as a unified paradigm, with evidence that scaling laws hold for recommendation models up to trillion-parameter scale."],"supporting_citations":[{"why":"Shows that the embedding layer dominates parameter and resource consumption, and that automatically searching embedding sizes improves efficiency—load-bearing evidence for the cost-optimization challenge.","marker":"[59]"},{"why":"Introduces trigger-induced instant interest modeling in a production CTR system, supporting the real-time preference capture challenge.","marker":"[120]"},{"why":"Models cold-start recommendation as a reinforcement-learning problem for lifetime value, supporting the long-term versus short-term objective trade-off.","marker":"[53]"},{"why":"Uses large language models to inject knowledge for cold-start and multi-objective recommendation on a music platform, supporting the LLM-based knowledge augmentation direction.","marker":"[149]"},{"why":"Describes timeliness, editorial oversight, and topic cycling in a deployed news recommender, grounding the news-specific challenges.","marker":"[65]"},{"why":"Presents a real-time bandit system for million-scale recommendations with millisecond latency, evidence for the real-time response challenge in content-oriented systems.","marker":"[166]"},{"why":"Debiases watch-time prediction via causal intervention on video duration, supporting the debiasing challenge in video recommendation.","marker":"[170]"},{"why":"Validates a trillion-parameter generative recommender and observes no performance saturation, underpinning the survey's claim that scaling laws apply to recommendation.","marker":"[169]"}],"fun_headline_variants":["Two tribes of recommender: transaction vs content","228 industrial systems split into two recommender tribes","Industrial recsys: the transaction-content divide","Why offline recommender research fails real-world systems","From 228 A/B tests: two recommender families emerge"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The survey counts a paper as industrial only if at least one author is from industry and the method was validated through online A/B testing; if A/B testing is not a necessary condition for capturing production practice, the reviewed corpus and the academic–industrial contrast rest on an incomplete sample.","fun_headline_variants_meta":{"raw":{"variants":["Two tribes of recommender: transaction vs content","228 industrial systems split into two recommender tribes","Industrial recsys: the transaction-content divide","Why offline recommender research fails real-world systems","From 228 A/B tests: two recommender families emerge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000599,"raw_usage":{"total_tokens":2627,"prompt_tokens":724,"completion_tokens":1903,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":1844}},"tokens_in":468,"tokens_out":1903,"duration_ms":16647,"temperature":1.0,"reasoning_tokens":1844,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T04:37:52.353043+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect all industry-authored recommender-system papers at the same six venues from 2020 to 2024 that report deployment but no online A/B testing; if a large share of those papers exhibit the same latency, cost, and multi-objective constraints as the A/B-validated set, the survey's selection criterion misses a substantial part of industrial practice.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Models cold-start recommendation as a reinforcement-learning problem for lifetime value, supporting the long-term versus short-term objective trade-off."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Uses large language models to inject knowledge for cold-start and multi-objective recommendation on a music platform, supporting the LLM-based knowledge augmentation direction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Presents a real-time bandit system for million-scale recommendations with millisecond latency, evidence for the real-time response challenge in content-oriented systems."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Validates a trillion-parameter generative recommender and observes no performance saturation, underpinning the survey's claim that scaling laws apply to recommendation."}],"review_version":1}