{"id":"9822403f-45d5-40b2-b68d-52166459ffc0","arxiv_id":"2507.04751","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new benchmark and prompt framework for generating and automatically evaluating product summaries that blend customer reviews with product metadata, with the best evaluator reaching 0.74 average Spearman correlation with human judgments.","lead":"The paper proposes a new AI task, multi-source opinion summarization, where a language model combines customer reviews with product specifications, descriptions, and ratings into a single short summary. It introduces a benchmark and LLM-based evaluation prompts for this task, and reports that users preferred these combined summaries to review-only ones in 87% of comparisons.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 87% user-preference claim is not yet supported because the comparator is unnamed and every evaluation question asks about the exact metadata M-OS adds.","rationale":"The reader's verdict is conditional, and I agree; I would keep it conditional. The most load-bearing point is the user study, because the abstract and contribution list use the 87% figure as direct evidence that M-OS 'enhances user engagement.' I considered the 0.74 correlation claim as an alternative locus of concern: it is computed on 50 products with no confidence intervals, and the improvement over OP-I-GPT-4o is small (0.74 vs. 0.70), so a confidence interval could easily include no improvement. However, Table 6 reports seven dimension-level correlations with several starred p < 0.05 entries, the average matches the abstract, and the claim is at least tied to a concrete dataset and prompts. The user study, by contrast, has no identifiable baseline and uses outcome questions that are definitionally biased toward M-OS. On its own, the missing baseline might be fixable by a disclosure; combined with the leading questions, it makes the reported 87% uninterpretable as evidence about user engagement. I therefore agree with the reader's weakest assumption but extend it: not only is the baseline unnamed, the outcome instrument is constructed so that M-OS wins by definition. The proposed test would disentangle 'M-OS is preferred because it contains more information' from 'M-OS is a better summary method.' Minor issues — 23,256 products in Table 2 vs. 25,000 in Section 4.1, appendix mislabeling, unreleased proprietary data — are real but secondary. The paper deserves conditional acceptance with a required re-analysis or disclosure of the user study, not rejection: the generation and evaluation framework is clearly specified, and the human annotation protocol with two-round adjudication is a genuine strength.","tokens_in":17893,"tokens_out":6366,"duration_ms":71299,"concrete_test":"Run a pre-registered follow-up user study with a matched comparator: the same generator (Qwen2.5-72B-Instruct), same temperature/top-p, same review input, and the only change being whether M-OS-GEN-PROMPT includes the metadata fields (description, features, specifications, rating). Add neutral outcome questions — overall usefulness, readability, trust, willingness to rely on the summary — alongside the current specification-centric questions, and analyze the data with a mixed-effects logistic regression or cluster-robust standard errors at the participant level. Report the baseline prompt verbatim in the appendix. If the preference gap shrinks to near chance on neutral questions, or if the specification-centric questions account for most of the gap, the 87% figure should be reinterpreted as a property of the added metadata, not of M-OS as a superior summary format.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central engagement claim — '87% of participants preferred M-OS over opinion summaries' — rests on a user study whose comparator is never specified. Section 7 says only that participants compared 'M-OS vs. traditional opinion-summary method', and Appendix C gives no model, prompt, or generation settings for the opinion summaries, despite giving full implementation details for M-OS generation (Appendix E.1). Without a matched baseline, the comparison is uncontrolled: if the traditional summaries were produced by a weaker LLM or a review-only prompt that deliberately omits metadata, the 87% figure measures information asymmetry rather than the value of M-OS as a method. The outcome instrument compounds this: all five evaluation questions in Appendix C ask about 'product specifications', 'technical details', and 'reduced need to look up additional product information' — precisely the features M-OS adds by construction. A preference study built on those questions is almost guaranteed to favor M-OS. The reported chi-square test (chi-squared = 3126.83, df = 1, p < .001) also treats all 6,000 judgments as independent despite being nested in 300 participants x 4 product pairs, so the p-value overstates significance. This does not undermine the separate 0.74 Spearman correlation, but it does mean the headline engagement claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Multi-Source Opinion Summarization (M-OS), in which LLMs generate summaries that integrate customer reviews with product metadata such as descriptions, key features, specifications, and ratings. It introduces M-OS-DATA, a proprietary product-metadata dataset; M-OS-EVAL, an evaluation benchmark with 7 dimensions and two rounds of expert annotation; and two evaluation prompt frameworks, OMNI-PROMPT and SPECTRA-PROMPTS. The authors benchmark 14 LLMs as summary generators and 5 LLMs as automatic evaluators, reporting an average Spearman correlation of 0.74 for OMNI-GPT-4o against human judgment, and a user study in which 87% of participants preferred M-OS summaries over traditional opinion summaries. The central claims are that M-OS improves user engagement and that the proposed prompts outperform previous evaluation methodologies.","tokens_in":18149,"tokens_out":6007,"duration_ms":61927,"significance":"If the results are supported, the paper would make a useful contribution: it provides a new task formulation, a benchmark with expert annotation, a large product-metadata resource, and evidence that reference-free LLM evaluation can align with human judgments for multi-source summaries. A particular strength is that the headline correlation is measured against human annotations rather than against model-generated gold labels, and the annotation procedure includes a second adjudication round that raises inter-rater agreement from 0.70 to 0.86. The generation prompts were not fitted to the preference outcome, which reduces one common circularity concern. However, the user-engagement claim and the claim of superiority over previous methodologies are not yet established at the level of evidence the paper asserts, for the reasons detailed below.","major_comments":[{"comment":"The user-study comparator is never specified. Section 7 states only that participants compared 'M-OS vs. traditional opinion-summary method', and Appendix C gives no model, prompt, or generation settings for the traditional summaries, although Appendix E.1 gives full implementation details for M-OS generation. Without a matched baseline, the 87% preference figure cannot be attributed to M-OS as a method: if the comparator was produced by a weaker model or a prompt that intentionally omitted metadata, the result measures information asymmetry rather than the value of M-OS. The outcome instrument compounds this problem, because all five questions in Appendix C D.2 ask about product specifications, technical details, and reduced need to look up additional product information, which are exactly the features M-OS adds by construction. The paper should identify the comparator, provide its generation configuration, and either use evaluation questions that do not presuppose the value of metadata or analyze the metadata-specific and non-metadata-specific questions separately.","section":"Section 7 and Appendix C"},{"comment":"The chi-square test treats all 6,000 preference judgments as independent, but the judgments are nested in 300 participants, each contributing 20 judgments across 4 product pairs and 5 questions. Judgments from the same participant are likely correlated, so the reported chi-square value of 3126.83 with df=1 and p<.001 substantially overstates the statistical significance, and the derived Cramer's V=0.72 inherits the same problem. The authors should report a participant-level analysis, such as a mixed-effects logistic regression with random intercepts for participants and products, or per-participant preference rates with confidence intervals. This is load-bearing because the 87% engagement claim in the abstract rests on this significance test.","section":"Appendix C D.3"},{"comment":"The headline average Spearman correlation of 0.74 is computed as an average of per-product correlations, each over only 14 summaries, and no confidence intervals or measures of dispersion are reported. The improvement over the OP-I-PROMPT baseline with GPT-4o (0.70 in Table 6) is small, and the asterisks in Table 6 mark only individual dimensions, not the average across dimensions. The claim in the abstract that M-OS-PROMPTS 'surpass the performance of previous methodologies' is therefore not statistically supported. The authors should report bootstrap or across-product confidence intervals for the average correlation, and a significance test for the difference between OMNI-PROMPT and the baselines, ideally accounting for the non-independence of the 14 summaries per product.","section":"Section 6.2, Table 6, Eq. (2)"},{"comment":"Equation (1) motivates a scoring function by sampling n outputs (n≈100) per input to estimate the score distribution p(s_k), but Appendix E.2 states that a temperature of 0.0 was used for evaluation. At temperature 0.0, repeated sampling is deterministic, so the 100 samples would be identical and p(s_k) would be degenerate, making the sampling-based justification for Eq. (1) inapplicable. The authors should clarify whether temperature was actually nonzero during the 100 evaluations, or revise the description of the scoring procedure to match the implementation. This affects the validity of all LLM-evaluator scores that feed into Tables 4 and 6.","section":"Section 3.3 and Appendix E.2"},{"comment":"The comparison baselines OP-PROMPTS and OP-I-PROMPT come from Siledar et al. (2024a), which shares several co-authors with this submission. Since the abstract claims to surpass 'previous methodologies' and the paper calls these baselines 'state-of-the-art', the comparison is at present only against the authors' own prior prompting variants. The authors should either add an independent baseline from another group, or temper the claim to 'surpass the previously proposed OP-PROMPTS and OP-I-PROMPT baselines'.","section":"Section 6.2 and Related Work"}],"minor_comments":[{"comment":"The abstract reports '4,900 summary annotations' for M-OS-EVAL, while Section 4.2 states 14,700 total ratings (3 raters × 50 products × 14 summaries × 7 dimensions). Please clarify whether 4,900 refers to summary-level scores after averaging raters, and state this consistently.","section":"Abstract and Section 4.2"},{"comment":"The list of models includes 'close-sourced' (typo for 'closed-source'), says the API model is GPT-4, while the experiments use GPT-4o, and lists Mistral-Nemo-Instruct-2407 and Qwen2.5-14B-Instruct, which do not appear in Table 5. Please reconcile the model list with the models actually reported.","section":"Appendix D"},{"comment":"The text contains duplicated phrases: 'specificityand specificity' and 'assessment guidelines guidelines'. These should be corrected.","section":"Section 3.2"},{"comment":"In the definition of sentiment consistency, 'the M-OSt common sentiment' appears to be a typo for 'the most common sentiment'.","section":"Appendix A"},{"comment":"Figure 4 is referenced as 'Figure ??' in the text; the cross-reference should be fixed.","section":"Appendix F"},{"comment":"The tables report Spearman and Kendall correlations but give no confidence intervals or measure of variability across products, and the footnote says '*' indicates p<0.05 without describing the significance test. At least one sentence should state how significance was computed.","section":"Tables 4 and 6"}],"recommendation":"major_revision","confidential_remarks":"The baseline-independence concern in major comment 5 is worth checking editorially: OP-PROMPTS and OP-I-PROMPT are attributed to Siledar et al. (2024a), whose author list overlaps substantially with this submission. If the journal treats that paper as prior work by the same group, the 'surpasses previous methodologies' claim should be framed as a comparison with the authors' own earlier prompts. Also, M-OS-DATA is described as proprietary and the collaboration details are withheld for anonymity; the paper should include a clear statement about data availability and whether the evaluation benchmark subset can be released."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The benchmark is the real contribution here; the 87% user-preference headline is not. M-OS-EVAL is a genuinely useful resource: 14,700 human ratings across 7 dimensions for 50 products, each with 14 LLM-generated multi-source summaries, and the annotation quality after the reconciliation round (Krippendorff's alpha 0.86) is respectable. The OMNI-PROMPT and SPECTRA-PROMPTS families extend the prior prompt-based evaluation work by Siledar et al. to include full product metadata, and the reported 0.74 average Spearman correlation for OMNI-GPT-4o against human judgments is plausible evidence that reference-free LLM evaluation can track human quality ratings on this task. That's a real result.\n\nThe soft spots are mostly in the user study. The comparator for the 87% preference figure is never named—no model, no prompt, no generation settings—while full details are given for the M-OS side. Worse, all five survey questions ask about specifications, technical details, and reduced need to look up product information, which is exactly what M-OS adds by construction. Those questions would favor any summary that includes metadata, so the preference number mostly measures information asymmetry, not summary quality. The chi-square test also treats 6,000 judgments as independent when they are nested in 300 participants, so the p-value overstates significance. If the authors name the baseline and re-analyze with a mixed model or per-participant proportions, the engagement claim could be salvaged, but as written it is not supported.\n\nThe correlation analysis has lesser issues: Spearman over 14 summaries per product is coarse, and no confidence intervals are given. The baselines being compared are the authors' own prior prompts, which is a limitation but not a fatal one, and the data are proprietary and not released, which limits reproducibility. These are fixable.\n\nThis paper deserves a serious referee. The benchmark and evaluation framework are useful to anyone working on opinion summarization or LLM-as-judge. I would send it to review with a request to fix the user study and release whatever data can be released. The 0.74 correlation stands as a reasonable empirical result; the 87% claim needs to be re-run or downgraded.","headline":"M-OS-EVAL is a solid benchmark; the 87% user preference claim is unsupported as reported.","tokens_in":18716,"tokens_out":3164,"would_cite":false,"duration_ms":30135,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multi-source opinion summaries, generated by LLMs from product metadata plus reviews, are preferred by 87% of users, and a structured prompt reaches $\\rho = 0.74$ agreement with human evaluators.","keywords":["multi-source opinion summarization","LLM-based evaluation","prompt engineering","reference-free evaluation","opinion summarization benchmark","e-commerce product metadata","user study"],"falsifier":"Re-run the user study with comparison summaries produced by the paper's strongest generator prompted on reviews alone; if the preference margin narrows to near chance, the 87% figure reflects a weak baseline rather than the value of multi-source information.","tokens_in":17689,"feed_emoji":"🛒","tokens_out":13372,"duration_ms":118584,"temperature":0.7,"pith_summary":"This paper argues that opinion summarization in e-commerce should go beyond reviews by folding in product descriptions, key features, specifications, and average ratings, producing what the authors call a multi-source opinion summary (M-OS). To support this, it introduces a proprietary product-metadata dataset of roughly 25,000 products and an evaluation benchmark of 4,900 human ratings across seven dimensions: fluency, coherence, relevance, faithfulness, aspect coverage, sentiment consistency, and specificity. The paper's central claims are that large language models can both generate such summaries and evaluate them without reference summaries when guided by structured prompts, and that users prefer the result: on average 87% of 300 study participants chose M-OS over a review-only summary. The best evaluator prompt reaches an average Spearman correlation of $\\rho = 0.74$ with human judgment, which the authors report as surpassing previous prompt-based evaluation methods.","feed_headline":"87% of users prefer multi-source summaries over review-only ones","feed_subtitle":"Blending product specs and ratings into summaries wins 87% of users; LLM evaluators match human raters at 0.74.","key_machinery":"The central mechanism is the M-OS-PROMPTS family. A generation prompt (M-OS-GEN-PROMPT) instructs the LLM to balance objective product data (title, description, key features, specifications, ratings) with subjective customer reviews, while the evaluation prompts (M-OS-EVAL-PROMPTS) share a four-part architecture: a system message, a task description, evaluation criteria, and an evaluation step. Two variants are introduced: OMNI-PROMPT, a single modular template whose Metric component can be swapped to assess any dimension, and SPECTRA-PROMPTS, seven dimension-specific prompts. Scores come from a weighted scoring function that estimates each candidate score's probability from roughly 100 samples per summary, effectively turning discrete 1–5 scores into a mean. The paper argues that the structured, step-by-step format—explicit percentage ranges for each score level and a demand for justification—reduces the score inflation observed in prior baselines and is what carries the $\\rho = 0.74$ average correlation.","core_discovery":"On its own terms, the paper establishes M-OS as a new task definition: a summary must cover both subjective opinions from reviews and objective product attributes from metadata, and it should serve purchasing decisions. It reports that the strongest generator among 14 benchmarked LLMs is Qwen2.5-72B-Instruct, with an average annotator rating of 4.186 across the seven dimensions, edging out GPT-4o at 4.169. For evaluation, the OMNI-PROMPT paired with GPT-4o achieves the highest average Spearman correlation ($\\rho = 0.74$) with human ratings, outperforming both the metric-dependent SPECTRA-PROMPTS and the earlier prompt baselines. The user study finds 86.6% overall preference for M-OS across 6,000 judgments, with per-criterion preference between 85.7% and 88.4%, and the chi-square statistic $\\chi^2 = 3126.83$ ($df = 1$, $p < .001$) rejects equal preference.","pith_inferences":["A natural follow-up would be to ablate the metadata sources (specifications, ratings, descriptions) to test which one drives user preference; the paper reports overall preferences but no per-source contribution analysis.","The $\\rho = 0.74$ correlation could degrade on out-of-distribution product categories or non-English reviews; a stress test along those lines would reveal whether the prompt framework generalizes beyond the curated e-commerce data.","The seven evaluation dimensions could transfer to other multi-document summarization tasks, such as legal or medical briefs, where faithfulness and aspect coverage carry similar weight, though the criteria would need domain-specific rewriting.","Because the evaluation benchmark draws on only 50 products, scaling it to several hundred would show whether the OMNI-PROMPT advantage is stable or concentrated in particular product categories."],"forward_implications":["E-commerce platforms could replace separate metadata pages and review sections with a single generated summary per product, since the user study indicates M-OS reduces the need for additional lookups.","Reference-free LLM evaluation at $\\rho = 0.74$ could substitute for expensive human annotation when comparing future multi-source summarization models along the same seven dimensions.","Open-source models are viable for both generation and evaluation: Qwen2.5-72B-Instruct outperformed GPT-4o in generation, and Llama-3.1-70B-Instruct and Mistral-7B-Instruct-v0.2 approached GPT-4o in evaluation, which matters for deployments that cannot use proprietary APIs.","Dimension-specific prompts (SPECTRA-PROMPTS) and the modular prompt (OMNI-PROMPT) each beat their respective baselines, indicating that structured evaluation instructions with explicit score ranges reduce score inflation.","The M-OS-EVAL benchmark, with 4,900 human ratings across seven dimensions, provides a reusable testbed for future multi-source opinion summarization systems."],"supporting_citations":[{"why":"Supplies the probability-weighted scoring function used to convert LLM scores into means and establishes the reference-free evaluation paradigm.","marker":"Liu et al., 2023a"},{"why":"Provides the expert-rater, two-round annotation design and meta-evaluation methodology that M-OS-EVAL follows.","marker":"Fabbri et al., 2021"},{"why":"Provides the earlier prompt-based evaluation baselines that the new evaluation prompts are compared against and outperform.","marker":"Siledar et al., 2024a"},{"why":"Defines the prior multi-source opinion summarization baseline that M-OS extends with fuller product metadata.","marker":"Siledar et al., 2024b"},{"why":"Defines the summary-level correlation formula used to compute Spearman and Kendall agreement with human ratings.","marker":"Bhandari et al., 2020"},{"why":"Documents the weak correlation of ROUGE and BERTScore with human judgment, motivating the shift to LLM-based reference-free evaluation.","marker":"Shen and Wan, 2023"}],"fun_headline_variants":["LLMs as architects and critics boost summary preference to 87%","Multi-source summaries win 87% of user votes over review-only","LLM evaluators match human judgment with 0.74 correlation","Fact-enriched summaries beat reviews: 87% user preference"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The user-preference result relies on the 'traditional opinion-summary method' used as the comparison being a fair, representative baseline, but the paper never identifies which model or prompt produced those summaries.","fun_headline_variants_meta":{"raw":{"variants":["LLMs as architects and critics boost summary preference to 87%","Multi-source summaries win 87% of user votes over review-only","LLM evaluators match human judgment with 0.74 correlation","Fact-enriched summaries beat reviews: 87% user preference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000759,"raw_usage":{"total_tokens":3391,"prompt_tokens":984,"completion_tokens":2407,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":2345}},"tokens_in":600,"tokens_out":2407,"duration_ms":15568,"temperature":1.0,"reasoning_tokens":2345,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:40:25.323936+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the user study with comparison summaries produced by the paper's strongest generator prompted on reviews alone; if the preference margin narrows to near chance, the 87% figure reflects a weak baseline rather than the value of multi-source information.","supporting_citations":[],"review_version":1}