{"id":"f93b3659-621b-47f1-9ccf-408d0c030ae8","arxiv_id":"2505.03828","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 2023-2025 survey of sentiment-aware e-commerce recommenders, arguing that review-text sentiment improves accuracy and explainability over rating-only models.","lead":"This paper reviews recent research on recommendation systems that use product review text, not just star ratings, to understand what customers like and dislike. It groups the methods into four families and argues that adding sentiment analysis makes online product suggestions more accurate and easier to explain.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central accuracy/explainability claim rests on a small, partially mis-cited evidence set: Table 1's entries do not trace to the stated references, and Section 2's search is not reproducible.","rationale":"The reader's weakest assumption was that the small selected set is representative of the 2023–2025 literature; this stress test agrees and sharpens the concern with verifiable citation failures in Table 1. I found no logical inconsistency in the argument itself: if the cited studies were correctly selected and correctly reported, the conclusion that sentiment can improve accuracy and explainability would be plausible and broadly consistent with the field. The problem is that the current manuscript makes that conclusion hard to audit. The reference list contains at least three mismatches in or near the central evidence table, and the search procedure is described but not documented, so the claimed trends rest on an unreproducible selection. I do not treat these as evidence of bad faith; they are correctable reporting errors. The existing CONDITIONAL verdict already requires exactly this kind of correction, so I do not move the verdict. One small correction to the reader's own review: Table 1 contains six rows, not five; the substantive point about the small, non-representative sample still stands.","tokens_in":11444,"tokens_out":8374,"duration_ms":81829,"concrete_test":"Independently re-run the Section 2 search on IEEE Xplore, ACM DL, Scopus, and arXiv with explicit queries reconstructed from the text, and record PRISMA-style counts (retrieved, deduplicated, screened, eligible). Then compare the eligible 2023–2025 set with the six Table 1 entries, and resolve each Table 1 citation to its actual paper: verify that RAKCR points to Cui et al., that Chat-Rec points to Gao et al.'s Chat-Rec paper rather than [20], and that the LLM-explanation citation points to Chen et al. rather than [15]. If the eligible set is much larger than six papers, or if any entry cannot be matched to the cited reference, the survey's evidence base is not representative or reliable enough for the claimed general trend.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the cited 2023–2025 studies are correctly attributed and that the six models in Table 1 are a representative sample. Section 2 describes a systematic search of four databases but reports no query strings, no screening counts, and no list of included studies, so representativeness cannot be checked. More importantly, the printed references do not consistently match the text. RAKCR is labeled [1] in Table 1, but reference [1] is Cambria et al.'s XAI-and-LLMs survey, while the actual RAKCR paper is reference [2]. Chat-Rec is labeled [20], but [20] is Gao et al.'s 'Is ChatGPT a good causal reasoner?', not the Chat-Rec paper; Section 2.2 also cites [1] for Chat-Rec. The 'Chen et al. (2023)' explanation work is cited to [15], which is Said (2025). These mismatches sit in the main evidence table, so a reader cannot verify the 'notable findings' the survey repeats. Additionally, no effect sizes or consistent rating-only baselines are reported; the claimed accuracy and explainability benefits are asserted from the original papers rather than synthesized. The hypothesis may be true, but this manuscript does not currently support it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a narrative literature review of sentiment-aware recommendation systems in e-commerce from an NLP perspective, covering work from 2023 to early 2025. It organizes the literature into four methodological categories—deep learning classifiers, transformer/LLM-based methods, graph neural networks, and conversational recommender systems—and summarizes representative models in Table 1 and datasets in Table 2. The abstract and introduction claim that integrating sentiment analysis into recommenders improves prediction accuracy and explainability through detailed opinion extraction. The review also discusses open challenges (noise, aspect alignment, dynamic preferences, cold start, scalability, evaluation, fairness) and proposes future research directions.","tokens_in":11640,"tokens_out":3278,"duration_ms":32892,"significance":"The topic is timely and practically important, and the paper provides a clearly structured entry point for NLP researchers working on e-commerce recommendation. Its strengths include the proposed four-category taxonomy, the concrete discussion of datasets and evaluation metrics, and the explicit enumeration of open challenges and future directions. However, the central claims about accuracy and explainability improvements rest on a small, partially mis-cited evidence set (Table 1) and a search methodology that is not reproducible. Once these verification issues are addressed, the review could serve as a useful synthesis; in its current form, its significance is limited by the lack of a defensible evidence base.","major_comments":[{"comment":"The citation labels in Table 1 do not consistently match the reference list, which undermines the verifiability of the paper's central evidence. For example, RAKCR is labeled [1], but reference [1] is Cambria et al.'s XAI/LLM survey; the actual RAKCR paper is reference [2]. Chat-Rec is labeled [20], but reference [20] is Gao et al.'s 'Is ChatGPT a good causal reasoner?' paper, not Chat-Rec; Section 2.2 also attributes Chat-Rec to [1]. Additionally, Section 2.3 attributes an explanation generation work to 'Chen et al. (2023)' via reference [15], which is Said (2025). Because Table 1 is the main support for the claimed trends and benefits, each row must trace to the correct reference or be removed. The authors should re-verify every citation in the manuscript, not only in Table 1.","section":"Table 1, Section 2"},{"comment":"The described systematic search is not reproducible and does not support the claim that the selected models are representative of 2023–2025 research. The methodology is reported in one paragraph with no full query strings, no screening counts (retrieved, deduplicated, screened, included), and no list of included studies. Without these elements, a reader cannot assess whether the five models in Table 1 are a selective convenience sample or a systematic synthesis. Please provide the exact search queries, the search date, PRISMA-style screening numbers, and either the full list of included studies or a link to a supplementary file.","section":"Section 2 (Literature Search Methodology)"},{"comment":"The paper's central claim—that sentiment integration enhances accuracy and explainability—is asserted in the abstract and introduction but not synthesized from quantitative evidence in the review. Section 3 discusses metrics and reproducibility in general terms but reports no effect sizes, no confidence intervals, and no consistent rating-only baselines for the models in Table 1. The 'Notable Findings' column of Table 1 contains qualitative statements such as 'Outperforms vanilla GCN on rating prediction' and 'higher precision/recall than using ratings alone' without the underlying numbers or directions of improvement. To substantiate the claimed benefits, the authors should either add a summary table of reported metrics (e.g., RMSE, MAE, Precision@N, NDCG) for each cited model relative to its baseline, or explicitly frame the claims as 'reported in the cited studies' without implying that this review independently establishes them.","section":"Section 3 (Results – Data & Metrics) and Abstract"}],"minor_comments":[{"comment":"The abstract describes the review as 'comprehensive,' but the scope (2023–early 2025, five representative models) is narrow. Please qualify the scope language to avoid overstatement.","section":"Section 1 and Abstract"},{"comment":"Table 1 contains a typo in the RAKCR row: 'improving personalization s' should be 'improving personalization.' Please proofread the table text.","section":"Table 1"},{"comment":"The 'Standardized Evaluation and Reproducibility' paragraph lists many tools (MLflow, Weights & Biases, Docker, Hugging Face) in a way that reads as a general reproducibility checklist rather than a synthesis specific to sentiment-aware recommenders. Consider condensing this to recommendations that directly relate to the surveyed literature.","section":"Section 3"},{"comment":"Figures 1 and 2 are referenced in the text, but the figure images are not included in the manuscript text. Ensure that final figure files are provided and that the captions are self-contained.","section":"Figure captions"},{"comment":"Several references lack author names or are incomplete, e.g., [11] 'Hotel-Review Datasets' and [12] 'Dataset list' have no author; [14] is a repository URL. For consistency, please complete all reference entries with authors where available.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The citation mismatches in Table 1 and Section 2.2 are a serious quality concern for a review paper because they directly affect the credibility of the evidence. The literature search description is also far below the standard expected for a systematic or even structured review. These issues are fixable within the manuscript's scope, so I recommend major revision rather than rejection. The single-author format is not a concern. The paper's fit with the journal's scope (cs.IR / e-commerce NLP) is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nYou can skip this paper unless you want a quick tour of the 2023–25 sentiment-aware recommendation literature. It is a review, so there is no new model, dataset, or derivation. What it does well is organize recent work into four buckets—deep learning, transformers, graph networks, conversational systems—and give an accessible plain-language summary. The dataset overview (Amazon, Yelp, TripAdvisor, MovieLens) is handy for a newcomer, and the challenges section (noise, sarcasm, aspect alignment, cold start) names the real practical issues. If a practitioner asked me for a starting point, I'd point them here.\n\nBut the paper's load-bearing evidence is shakier than the prose. Table 1, the main summary of representative models, has citation mismatches that a reviewer would catch immediately. RAKCR is labeled [1], but [1] is Cambria et al.'s XAI/LLM survey; the actual RAKCR paper is [2]. Chat-Rec is labeled [20], but [20] is Gao et al.'s ChatGPT causal-reasoning paper, not the Chat-Rec work—and Section 2.2 cites [1] for Chat-Rec, which is also wrong. The 'Chen et al. (2023)' explanation work is attributed to [15], which is actually Said (2025). These are not cosmetic errors; they sit in the evidence table that supports the survey's central claim that sentiment improves accuracy and explainability. A reader cannot verify the 'notable findings' without guessing which paper is actually being described.\n\nThe search methodology is also underreported. One paragraph says four databases were searched, but no query strings, no screening counts, and no list of included studies are given. With only five models in Table 1, the claim of a comprehensive review of 2023–25 work is not established. And there is no synthesis of effect sizes or consistent baselines—the accuracy/explainability benefits are asserted from the original papers rather than compared across studies. The hypothesis may well be true; this manuscript just doesn't demonstrate it.\n\nSo: the paper has value as an orientation aid, but not as a reliable reference. The citation errors are fixable, but they are exactly the kind of thing that undermines a reader's trust in a review. I would send it to peer review with a request for major revision, mainly to correct the reference mapping and report a reproducible search process. I would not cite it in its current form.","headline":"A conventional survey that reads as a useful orientation to 2023–25 sentiment-aware recommenders, but its main evidence table has citation mismatches and the search methodology is too thin to support the 'comprehensive' label.","tokens_in":12163,"tokens_out":3406,"would_cite":false,"duration_ms":31099,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review claims that adding sentiment analysis of review text to e-commerce recommenders improves accuracy, personalization, and explainability.","keywords":["sentiment-aware recommendation","e-commerce","natural language processing","review-based recommender","transformer models","graph neural networks","conversational recommender","explainable recommendations"],"falsifier":"A systematic replication of the review's search, with full screening counts, that runs a sentiment-aware model against an otherwise identical rating-only baseline on the Amazon, Yelp, and TripAdvisor datasets; if review-text sentiment adds no consistent gain in RMSE, NDCG, or explanation quality across those datasets, the central claim fails.","tokens_in":11229,"feed_emoji":"🛒","tokens_out":5792,"duration_ms":58352,"temperature":0.7,"pith_summary":"Online stores mostly recommend by star ratings, leaving the opinions in review text unused. This review argues that e-commerce recommenders should mine those written reviews for sentiment and feed that signal into the recommendation pipeline. Covering 2023 to early 2025, it groups current work into four families: deep-learning encoders, transformer- and LLM-based models, graph neural networks that propagate sentiment, and conversational systems that adapt to feelings expressed in dialogue. If the reviewed evidence holds, review text is not a side channel but a load-bearing input that improves rating prediction, makes suggestions easier to explain, and helps cold-start situations where little interaction data exists.","feed_headline":"Survey: mining review sentiment sharpens e-commerce recommendations","feed_subtitle":"A 2023-2025 review finds that adding review-text sentiment to ratings improves accuracy and explainability.","key_machinery":"The central mechanism is the sentiment-extraction-to-scoring pipeline: raw reviews are passed through an NLP sentiment module that produces sentiment embeddings, sentiment scores, or aspect-level opinion vectors, and those signals are then merged with collaborative-filtering, graph, or dialogue state before ranking. This machinery does the work in every category of the review. In graph methods, sentiment becomes edge weights or relation types; in conversational systems, it becomes an entity-level emotional score that filters or re-weights candidates; in transformer methods, it is the encoded context that the scoring layer consumes. The implicit identity behind the whole review is that a user's opinion vector over item aspects is a better predictor of future preference than the scalar rating alone.","core_discovery":"The paper's central claim is that treating review text as a source of sentiment, not as a bag of words, lets recommender systems predict ratings and rankings more accurately while also producing recommendations that can be explained by the opinions that drove them. It organizes the 2023-2025 literature into four approach families: deep-learning classifiers that attach sentiment embeddings to user-item interactions; transformer-based methods that extract context-sensitive sentiment; graph neural networks that propagate sentiment signals through user-item-entity graphs; and conversational recommenders that re-weight or filter entities by the sentiment expressed in dialogue. The paper's demonstration is architectural: sentiment flows from raw reviews through an NLP extraction module into the scoring engine, and the resulting scores align with what users explicitly praised or criticized. It reports representative results such as BERT-based sentiment features improving precision and recall over ratings-only baselines, transformer encoders cutting RMSE and MAE relative to non-transformer deep networks, and a conversational system that avoids recommending items the user declared dislike for.","pith_inferences":["A direct but unstated corollary is that the four-way taxonomy can serve as a design menu: choose conversational sentiment monitoring when dialogue exists, graph propagation when entity relations are rich, and transformer encoders when review text is long and nuanced.","The review treats sentiment as a signal that can be plugged into existing recommender backbones; the harder claim that sentiment causes, rather than merely correlates with, better recommendations is left open, and the paper itself flags causal inference as future work.","Because only five models are summarized in detail, a focused replication that runs one sentiment-aware model against the same rating-only baseline across Amazon, Yelp, and TripAdvisor would test whether the claimed gains are robust to domain and dataset scale."],"forward_implications":["Adding review-text sentiment to rating-based collaborative filtering should lower rating-prediction error and improve ranking metrics such as precision, recall, and NDCG.","Recommendation explanations can cite which opinions drove the suggestion, for instance by noting that a user praised battery life, which increases transparency and user trust.","Conversational recommenders that filter entities by the user's expressed sentiment avoid recommending items the user explicitly disliked and can reach good recommendations in fewer dialogue turns.","Transformer- and LLM-based sentiment encoders reduce manual feature engineering and help cold-start cases where only text, such as item descriptions or early reviews, is available.","Because textual sentiment can carry and amplify societal biases, fairness-aware training and debiasing of sentiment analysis are necessary before these systems are widely deployed."],"supporting_citations":[{"why":"Supplies RAKCR, the knowledge-graph convolutional model whose sentiment-weighted edges are evidence for sentiment propagation improving personalization.","marker":"[2]"},{"why":"Supplies the BERT-hybrid result that review sentiment features improve precision and recall over ratings alone on Yelp data.","marker":"[3]"},{"why":"Supplies the comparative evaluation where transformer encoders beat non-transformer deep networks on Amazon Electronics in RMSE and MAE.","marker":"[4]"},{"why":"Supplies the deep-learning review-encoder line, including attention and aspect-topic models, that the survey builds its first trend on.","marker":"[5]"},{"why":"Supplies a 2025 sentiment-CNN plus collaborative-filtering system whose accuracy and F1 gains on Amazon data support the hybrid approach.","marker":"[6]"},{"why":"Supplies the large-scale Amazon Reviews 2023 dataset that motivates text-augmented modeling at e-commerce scale.","marker":"[7]"},{"why":"Supplies SECR, the conversational recommender whose entity-level sentiment filtering avoids disliked items and improves dialogue satisfaction.","marker":"[8]"},{"why":"Supplies SENGR, an early GNN recommender that propagates sentiment signals and supports the graph-based category.","marker":"[18]"}],"fun_headline_variants":["Sentiment in reviews boosts e-commerce recommender accuracy","Mining review sentiment sharpens recommendation precision","Survey: sentiment-aware recommenders beat ratings-only baselines","Review sentiment powers explainable, accurate recommendations","NLP sentiment integration improves recommender systems, 2023-2025 review"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's conclusions rest on the five selected 2023-2025 papers being representative of the wider sentiment-aware recommendation literature.","fun_headline_variants_meta":{"raw":{"variants":["Sentiment in reviews boosts e-commerce recommender accuracy","Mining review sentiment sharpens recommendation precision","Survey: sentiment-aware recommenders beat ratings-only baselines","Review sentiment powers explainable, accurate recommendations","NLP sentiment integration improves recommender systems, 2023-2025 review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000282,"raw_usage":{"total_tokens":1658,"prompt_tokens":923,"completion_tokens":735,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":656}},"tokens_in":539,"tokens_out":735,"duration_ms":7206,"temperature":1.0,"reasoning_tokens":656,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:06:11.993250+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic replication of the review's search, with full screening counts, that runs a sentiment-aware model against an otherwise identical rating-only baseline on the Amazon, Yelp, and TripAdvisor datasets; if review-text sentiment adds no consistent gain in RMSE, NDCG, or explanation quality across those datasets, the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies RAKCR, the knowledge-graph convolutional model whose sentiment-weighted edges are evidence for sentiment propagation improving personalization."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the comparative evaluation where transformer encoders beat non-transformer deep networks on Amazon Electronics in RMSE and MAE."},{"cited_title":"and Yeom, S","cited_arxiv_id":null,"evidence_quote":"Supplies the deep-learning review-encoder line, including attention and aspect-topic models, that the survey builds its first trend on."},{"cited_title":"good”, “bad","cited_arxiv_id":null,"evidence_quote":"Supplies a 2025 sentiment-CNN plus collaborative-filtering system whose accuracy and F1 gains on Amazon data support the hybrid approach."},{"cited_title":"(2023) Large-Scale Amazon Reviews dataset, collected in 2023 by McAuley Lab, Amazon Reviews’23","cited_arxiv_id":null,"evidence_quote":"Supplies the large-scale Amazon Reviews 2023 dataset that motivates text-augmented modeling at e-commerce scale."},{"cited_title":"Builds a movie knowledge graph (MAKG) and filters entities by sentiment scores (using a sentiment lexicon and prompt- based analysis) to inform recommendations","cited_arxiv_id":null,"evidence_quote":"Supplies SECR, the conversational recommender whose entity-level sentiment filtering avoids disliked items and improves dialogue satisfaction."},{"cited_title":"One of the first GNN -based recommenders to include sentiment","cited_arxiv_id":null,"evidence_quote":"Supplies SENGR, an early GNN recommender that propagates sentiment signals and supports the graph-based category."}],"review_version":1}