{"id":"f4fd95ad-0637-4957-bee7-12edb23d0e59","arxiv_id":"2507.21057","paper_version":1,"verdict":"REJECT","confidence":"LOW","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Service reviews show stronger, more varied emotions than product reviews, and the gap is larger for Eastern than Western consumers, while gender has little moderating effect.","lead":"This paper compares emotions in online reviews of products versus services and tests whether culture and gender change the pattern. It reports that service reviews carry stronger and more varied emotions, especially for Eastern consumers, while gender makes little difference.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated dataset category labels and culture coding undermine the central claim: ASOS TrustPilot, described as mixed product/service, is labeled Product, and no rule maps reviewers to Eastern/Western.","rationale":"The reader's weakest_assumption correctly identifies unvalidated dataset labels and culture coding as the foundation of the analysis. My independent read finds the same, plus additional internal evidence: Section 3.2 admits the ASOS TrustPilot dataset mixes products and services, and Table 3's df=2 for a binary culture variable is internally inconsistent. These issues are load-bearing because they affect the independent variable (Category) and moderator (Culture) directly; if labels are wrong, the significant F-tests in Tables 2 and 3 cannot be interpreted as product-versus-service differences. The study would need a data release with validated labels and a transparent culture mapping before the central claim could be assessed. Since this reinforces the reader's REJECT verdict rather than overturning it, no verdict adjustment is needed.","tokens_in":21082,"tokens_out":3104,"duration_ms":32156,"concrete_test":"Request the culture-assignment code and raw per-review category labels, then re-run the Table 2 and Table 3 MANOVAs with ASOS TrustPilot reclassified as 'Service' and with the documented culture mapping. If the Sentiment Score F (currently 259.398) or the Category by Culture F (currently 89.660) changes materially or loses significance, the central claim is an artifact of unvalidated category and culture coding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that services elicit stronger or higher sentiment than products, moderated by Eastern versus Western culture, rests entirely on Table 1's category assignments and the binary culture split. Two problems make this foundation unreliable. First, Section 3.2 describes the ASOS TrustPilot dataset as capturing 'customer experiences with ASOS products and services, focusing on fit, style preferences, and delivery', yet Table 1 labels it 'Product'. TrustPilot is a platform for reviewing companies' overall service, so many of its reviews concern delivery and returns, which are service quality. Mislabeling a mixed dataset as pure Product biases the category contrast. Second, no rule is given for assigning individual reviews to Eastern versus Western culture. Datasets include FashionNova, Amazon (global), Celsius Network, and others; the manuscript never states which nationality or location field was used or how Hofstede dimensions were mapped to a binary. Without that mapping the Category by Culture interaction in Table 3 is not reproducible. Additionally, Table 3 reports df=2 for Category by Culture, but the paper describes only two cultures; if Culture has only two levels, the degrees of freedom are impossible, indicating either an unreported third culture level or a misreported analysis. Either way, the moderation result cannot be trusted as presented. Since every hypothesis test uses these labels, the headline findings are not robust to plausible relabeling.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper examines whether product reviews differ from service reviews in sentiment and emotion (valence, arousal, dominance, basic and advanced emotions), and whether culture and gender moderate these differences. Using six publicly available datasets labeled as product or service, the author applies four pre-trained NLP models to score reviews, then tests hypotheses with MANOVA and a PROCESS moderation analysis. Results show significant category effects on sentiment score, valence, arousal, dominance, and advanced emotion categories, but not on sentiment category or basic emotion categories; culture moderates several outcomes and gender does not. The paper concludes that services evoke stronger and more complex emotional responses and that this gap is especially pronounced for Eastern consumers, extending Maslow's hierarchy and Hofstede's framework.","tokens_in":21364,"tokens_out":7032,"duration_ms":64225,"significance":"The study's ambition is timely: integrating VAD dimensions and fine-grained emotion categories into a comparison of product vs. service reviews, with cultural moderation, addresses a real gap in the e-commerce sentiment literature. The use of publicly available datasets and pre-trained transformer models is reproducible in principle, and the explicit H1-H3 structure makes the empirical claims clear. However, the empirical foundation is not yet reliable: the dataset category labels are inconsistent with the text, culture is not operationalized, and one of the key tables reports impossible degrees of freedom. If the reported effects survive corrected labeling and transparent coding, the findings would be useful for emotion-aware and culture-aware review analytics.","major_comments":[{"comment":"The product/service label for ASOS TrustPilot is contradicted by the paper's own description: Section 3.2 states the ASOS TrustPilot Customer Review Dataset 'captures customer experiences with ASOS products and services, focusing on fit, style preferences, and delivery,' yet Table 1 classifies it as 'Product.' TrustPilot is a platform on which reviewers rate the overall service quality of companies, so labeling it a pure product dataset is unjustified. Because every hypothesis test in Section 4 uses Category as the independent variable, this mislabeling can bias the category contrasts, the cultural interactions, and the gender interactions. Please validate the category assignment (e.g., manual coding of a random sample or an explicit rule based on the review text) and rerun the analyses, or provide evidence that the ASOS reviews are predominantly product-focused.","section":"Section 3.2, Table 1"},{"comment":"The manuscript never specifies how an individual review was assigned to 'Eastern' or 'Western' culture. The datasets in Table 1 include Amazon Reviews (which are global) and Celsius Network, and no nationality or location field is described for these; the paper also does not state how Hofstede's dimensions were mapped onto the binary culture variable. Without a coding rule, the Category × Culture interactions in Table 3 and the plots in Figures 3-5 are not reproducible, and the cultural moderation effect cannot be distinguished from dataset-specific confounds. Please provide the exact mapping (e.g., reviewer self-reported country, IP-based location, or platform-specific fields) and report the number of reviews per culture and per category.","section":"Section 3.2 and Section 3.4"},{"comment":"All Category × Culture rows in Table 3 report df = 2. With two levels of Category (Product, Service) and two levels of Culture (Eastern, Western), the interaction degrees of freedom should be (2−1)×(2−1) = 1. The reported df = 2 implies an unreported third level of Culture (or of Category), or a mis-specified model. Since the text describes only a binary Eastern/Western split, the moderation results cannot be interpreted as reported. Please clarify the actual number of culture levels and correct the table; if a three-level culture variable was used, the hypotheses and the interpretation of the cultural moderation must be revised accordingly.","section":"Table 3"},{"comment":"The support for H1 is overstated. Table 2 shows that Category has no significant effect on Sentiment Category (p = .614) or Basic Emotion Category (p = .621), and Table 3 shows non-significant Category × Culture interactions for Arousal (p = .799) and Advanced Emotion Categories (p = .401). Yet the abstract claims 'clear differences in emotional expression and sentiment between the two' and Section 4.1 concludes that 'services tend to elicit a higher proportion of complex emotions compared to products.' Additionally, Table 4 appears to mark Basic Emotion Category as 'Yes' for H1 despite the non-significant result in Table 2. The conclusions should be restricted to the dependent variables that actually show significant effects, and the discrepancy between Tables 2 and 4 should be corrected.","section":"Abstract, Section 4.1, Table 2, Table 4"},{"comment":"No sample sizes, group means, standard deviations, effect sizes, or confidence intervals are reported. Section 3.4 asserts that 'the dataset, consisting of a relatively equal representation of product and service reviews, male and female consumers, as well as Western and Eastern reviews, enabled balanced comparisons,' but no counts are provided. Without N and effect sizes, the non-significant gender moderation results (Table 3) are uninterpretable—they could reflect a genuinely absent effect or a lack of statistical power—and the magnitude of the claimed cultural effects cannot be evaluated. Please report per-cell sample sizes, means, and either partial η² or Cohen's d with confidence intervals.","section":"Section 3.4 and Section 4"}],"minor_comments":[{"comment":"The heading is numbered '4..3.1.' (typo).","section":"Section 4.3.1"},{"comment":"The text contains 'Hofstede’s’s cultural dimensions' and 'the circumplex model of the effect' (should be 'affect').","section":"Section 2.1"},{"comment":"Figure 2 is difficult to read with 29 emotion categories plotted against an unlabeled x-axis; consider ordering by the product-service difference and adding data labels.","section":"Figure 2"},{"comment":"The four Hugging Face models are named but no repository URLs or version identifiers are given; adding these would aid reproducibility.","section":"Section 3.3"},{"comment":"Table 4 is badly formatted (e.g., 'Basic Emotion CCategory,' missing column separators) and inconsistent with Table 2; it should be regenerated.","section":"Table 4"},{"comment":"The reference list contains incomplete or malformed entries (e.g., Cortis & Davis 2021; Sudirjo et al. 2023; Truong 2024) and duplicate 'Beyond culture' entries (Hall, 1976a/1976b).","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has substantial methodological gaps and appears to be an early draft; as submitted it is not close to being publishable. The author's prior work is cited repeatedly in ways that seem tangential (Truong 2016, 2020, 2022, 2024), including in the literature review and practical implications; this is worth monitoring but not decisive. The publisher should be aware that the empirical claims are not reliable without reanalysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's headline finding—services elicit stronger sentiment than products, especially for Eastern consumers—rests on two unvalidated coding decisions. The ASOS TrustPilot dataset is labeled 'Product' in Table 1 despite Section 3.2 describing it as capturing both products and services, and on a platform like TrustPilot most reviews concern service attributes like delivery and returns. No rule is given for mapping reviewer nationality or location to the Eastern/Western split, so the central moderating variable is not reproducible. The moderation table reports df=2 for Category*Culture even though the paper describes only two cultures; that is either an unreported third level or a misreported analysis. These are load-bearing flaws, not cosmetic ones.\n\nCredit where it is due: the study applies a reasonably broad emotion toolkit (VAD, VADER, GoEmotions) to six datasets, and the basic question—whether product and service reviews differ emotionally—is worth asking. The gender moderation analysis is a sensible addition, and the complex emotion figure is a nice visual. The paper also reports enough statistical detail to see what was tested, which is more than many applied papers do.\n\nThe problems, however, run deep. Dataset labels come from the author's judgment with no validation. The culture coding is opaque. No sample sizes, effect sizes, or confidence intervals are reported anywhere. Significance stars are inconsistent: p=.041 gets the same triple asterisk as p<.001. H1 is only partially supported—category is not significant for sentiment category or basic emotion category—yet the abstract and summary claim broad differences. The frequent self-citations are not disqualifying, but they do not strengthen the evidence.\n\nThis is a practice-oriented marketing paper. A reader wanting a quick descriptive look at review language might skim it, but the analytic foundation is too shaky for the results to be cited as evidence. The topic is real and the tools are current, so with data, code, and corrected labels it could be reassessed.\n\nIf this came across my desk, I would not desk-reject it outright—the flaws are specific and potentially fixable. But I would send it to a referee with instructions to focus on dataset labeling, culture coding, and the reporting inconsistencies. As is, it does not support its claims.","headline":"A plausible but under-supported empirical claim about product vs service review sentiment and culture; the dataset labeling and culture coding need to be fixed before the findings can be trusted.","tokens_in":21843,"tokens_out":2968,"would_cite":false,"duration_ms":30357,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Service reviews carry stronger and more complex emotions than product reviews, and the gap is widest for Eastern consumers.","keywords":["online reviews","sentiment analysis","product versus service","cultural moderation","emotional expression","valence-arousal-dominance","e-commerce","consumer emotions"],"falsifier":"Rerun the same analysis after relabeling the dataset that came from a service-review platform as a service dataset instead of a product dataset; if the service-over-product sentiment gap and the East–West interaction shrink, flatten, or flip, the central claim depends on the dataset labels rather than on the nature of products and services.","tokens_in":20914,"feed_emoji":"💬","tokens_out":9069,"duration_ms":83276,"temperature":0.7,"pith_summary":"This paper argues that reviews of services and products are emotionally different in a systematic way, and that the customer's cultural background changes how that difference shows up. Analyzing six review datasets with machine-learning emotion classifiers, it finds that service reviews score higher in sentiment, arousal, and complex emotion categories such as love, excitement, and disappointment, while product reviews cluster around relief, pride, and caring. Culture moderates the pattern: Eastern reviewers show a much larger rise in sentiment when moving from products to services, while Western reviewers stay flatter or decline. Gender shows no meaningful moderating effect. If the pattern holds, it would mean that product and service reviews are distinct emotional genres, and that review analytics and marketing strategies should be split by category and culture rather than treated as one uniform signal.","feed_headline":"Service reviews carry more emotion than product reviews","feed_subtitle":"The service-product sentiment gap is widest for Eastern customers; gender barely matters.","key_machinery":"The machinery is a layered emotion-measurement stack applied to each review. Sentiment is scored by a multilingual sentiment model and assigned to positive/negative/neutral buckets by a rule-based sentiment analyzer; the valence–arousal–dominance (VAD) model maps each review onto three continuous emotional dimensions; a six-category basic-emotion classifier and a 27-category complex-emotion classifier supply discrete emotional labels. The statistical core is a multivariate analysis of variance with category as the independent variable and culture and gender as moderators, followed by a moderation regression with the same interaction terms. The conceptual pivot is the product/service dichotomy viewed through a hierarchy of needs: products satisfy basic functional needs while services satisfy relational and self-fulfillment needs, and this difference is what predicts the emotion gap and its cultural moderation.","core_discovery":"The central discovery is that the product/service distinction predicts the emotional register of a review. In the paper's data, services elicit higher sentiment scores, higher arousal, higher dominance scores, and a greater share of complex emotion categories—love, excitement, amusement, optimism, and admiration, but also disgust, disappointment, disapproval, sadness, confusion, and embarrassment—whereas products elicit more relief, pride, and caring. The category effect on sentiment, sentiment category, valence, and dominance is itself moderated by culture: Eastern consumers show a pronounced increase in sentiment when moving from products to services, while Western consumers show little change or a decline. The paper interprets these patterns through a hierarchy-of-needs lens, arguing that services engage higher-order social and self-fulfillment needs and therefore produce stronger and more varied emotions, with individualism–collectivism shaping which needs dominate expression. Gender does not moderate the effect.","pith_inferences":["A natural next test would be to score bundled offerings—products sold with installation, subscription, or customer-support components—and check whether their emotion profiles fall between pure products and pure services.","The East–West moderation could partly reflect platform-specific writing norms rather than deep cultural values; comparing identical products sold through local and international platforms would separate the two.","Using continuous individualism scores per reviewer country instead of a binary East–West split would sharpen or weaken the moderation claim, since the binary coding discards within-group cultural variance.","Because service reviews carry more negative complex emotions like disgust and disappointment, polarity-only dashboards will miss the most actionable warnings hidden in service feedback."],"forward_implications":["Treating review sentiment as a single pool will systematically overstate the positivity of service reviews relative to product reviews in any ranking or dashboard.","Service businesses should design around emotional and relational touchpoints, since service reviews carry a wider and more intense emotional spectrum.","In markets with many Eastern consumers, service-heavy campaigns should emphasize experiential and relational value because the sentiment premium of services over products is larger there.","Gender-based personalization of review interpretation is unlikely to pay off; the paper finds no gender moderation of the category effect.","Sentiment classification into positive/neutral/negative labels misses much of the category difference, so analytics that rely only on polarity will understate the service effect."],"supporting_citations":[{"why":"Supplies the hierarchy-of-needs framing that predicts products satisfy basic needs while services engage higher emotional needs.","marker":"(McLeod, 2007)"},{"why":"Distinguishes goods from services on intangibility and variability, motivating the product/service comparison.","marker":"(Vargo & Lusch, 2004)"},{"why":"Provides the 27-category emotion taxonomy used to measure complex emotion categories in reviews.","marker":"(Cowen & Keltner, 2017)"},{"why":"Supplies the labeled emotion dataset that the complex-emotion classifier is built on.","marker":"(Demszky et al., 2020)"},{"why":"Offers empirical precedent that service reviews carry higher emotional intensity, which the study extends.","marker":"(Li et al., 2020)"},{"why":"Provides the individualism–collectivism dimension used to code Eastern versus Western culture in the moderation analysis.","marker":"(Hofstede, 2011)"},{"why":"Supplies the rule-based sentiment analyzer used to assign positive, negative, and neutral sentiment categories.","marker":"(Hutto & Gilbert, 2014)"},{"why":"Provides the circumplex model underlying the valence, arousal, and dominance scores.","marker":"(Russell, 1980)"}],"fun_headline_variants":["Service reviews pack more emotional punch than product reviews","Culture shapes the emotion gap between product and service reviews","Service reviews stir stronger feelings, but culture rewires the effect","Emotion gap widens for Eastern customers across product vs service reviews","Services trigger richer emotions than products, culture flips the intensity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis takes the product/service label attached to each of the six datasets as correct even though one dataset was drawn from a service-review platform and labeled as product reviews, and it collapses reviewer nationality into a binary Eastern/Western split; if either of those assumptions is wrong, the category comparisons and the cultural moderation built on them do not hold.","fun_headline_variants_meta":{"raw":{"variants":["Service reviews pack more emotional punch than product reviews","Culture shapes the emotion gap between product and service reviews","Service reviews stir stronger feelings, but culture rewires the effect","Emotion gap widens for Eastern customers across product vs service reviews","Services trigger richer emotions than products, culture flips the intensity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1270,"prompt_tokens":996,"completion_tokens":274,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":192}},"tokens_in":612,"tokens_out":274,"duration_ms":3375,"temperature":1.0,"reasoning_tokens":192,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:37:40.301989+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the same analysis after relabeling the dataset that came from a service-review platform as a service dataset instead of a product dataset; if the service-over-product sentiment gap and the East–West interaction shrink, flatten, or flip, the central claim depends on the dataset labels rather than on the nature of products and services.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the rule-based sentiment analyzer used to assign positive, negative, and neutral sentiment categories."}],"review_version":1}