{"id":"f36849a3-3794-4bb8-814f-325e3555f2e7","arxiv_id":"2411.13485","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Three prompt methods using gpt-4o-mini generate synthetic product desirability reviews with high internal sentiment alignment (Pearson 0.93 to 0.97), but no external human validation is performed.","lead":"This paper tests three ways to use a low-cost AI model, gpt-4o-mini, to generate synthetic product reviews for Product Desirability Toolkit testing. It reports that the synthetic reviews track requested sentiment scores closely and cost very little, but it does not yet compare them against real human reviews.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 'sentiment alignment' is evaluated by the same LLM that generated the reviews; for Supply-Word the target is also model-generated, so the high correlations may reflect self-consistency, not human-valid sentiment.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing issue: model-generated sentiment scores are being treated as a proxy for human sentiment. I agree with that identification. My stress-test pass confirms the concern is not merely hypothetical: Section III-B.1 describes how all three methods were scored using the 'Complete' prompt with gpt-4o-mini, and for Supply-Word the target score itself is also generated by gpt-4o-mini. Thus the reported Pearson correlations are internal-consistency measures, and the abstract's wording ('Results demonstrated high sentiment alignment...') does not carry the caveat that no human ground truth was involved. The paper itself scopes the result as an initial examination and lists human comparison as future work, which is honest but does not remove the gap between the abstract's claim and the evidence. I also considered the textual-diversity results, but those are secondary to the central claim and the paper already flags the confound of unequal dataset sizes. I see no reason to change the reader's CONDITIONAL verdict: the internal measurements are plausible and reproducible, but external validation with human annotators is a necessary condition before the claimed alignment can be taken at face value. The concrete test I propose would settle whether the concern actually lands.","tokens_in":13854,"tokens_out":3440,"duration_ms":37134,"concrete_test":"Score a stratified sample of, say, 100 reviews per method (or the full 3000) with independent human annotators using a continuous 0-1 sentiment scale, then compute Pearson correlation and mean absolute difference between the original target scores and the human ratings. If human correlations fall materially below 0.93-0.97, or MAD rises well above the reported 0.09-0.13, the reported alignment is an artifact of self-scoring and the central claim needs qualification. Ideally, also run the same generation and scoring pipeline on the two existing human PDT datasets cited in Future Work (refs. [26] and [30]) to benchmark synthetic and human distributions directly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption is that gpt-4o-mini's 'Complete'-prompt sentiment scores (Section III-B.1, Table II) are a valid proxy for the sentiment a human would assign. The headline Pearson correlations in Table IV (0.93, 0.96, 0.97) are computed between the target score and this same model's rating. For Word+Review and Review+Word, the target is externally supplied, but the evaluator is the same model that wrote the review and selected the word, so the correlation can reflect the model's consistency in following its own prompt rather than a human-meaningful property. For Supply-Word the target itself is also produced by gpt-4o-mini from the word, then used to generate the review, then scored again by gpt-4o-mini; the 0.97 correlation is therefore largely a self-consistency check. The Supply-Word-4o comparison does not break this circularity: both are LLM scorers with no anchor to human judgments. The paper's own future-work section concedes that direct comparison with human PDT datasets remains to be done. If the self-scored evaluation is optimistic, the central claim of 'high sentiment alignment' and the recommendation that Review+Word and Supply-Word are 'more reliable for applications requiring precise sentiment reflection' are not supported by the reported evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes three LLM-based methods for synthesizing Product Desirability Toolkit (PDT) datasets with gpt-4o-mini: Word+Review, Review+Word, and Supply-Word. Each method generates 1000 hypothetical software product reviews paired with PDT words and a sentiment score. The datasets are assessed in terms of sentiment alignment between target and evaluated scores, textual diversity (via compression ratio, part-of-speech compression, ROUGE-L homogenization, and n-gram diversity), and generation cost. The authors report high Pearson correlations (0.93–0.97) between target and evaluated scores, claim that Review+Word and Supply-Word show the best alignment, and conclude that LLM-generated synthetic PDT data is a scalable, cost-effective alternative when real data are scarce. The datasets are publicly released on Zenodo.","tokens_in":14076,"tokens_out":3885,"duration_ms":41689,"significance":"If the reported alignment reflects human-meaningful sentiment, the paper would offer a practical, low-cost pipeline for generating PDT-style data, addressing the acknowledged scarcity of large PDT datasets. The work is transparent in its methodology, publishes the generated datasets, provides detailed cost and token accounting, and uses standard diversity metrics. The main limitation is that the evaluation is not independent: the same model family that generates the reviews also scores them, so the reported correlations largely measure self-consistency rather than external fidelity. The paper's future-work section explicitly acknowledges that direct comparison with human PDT datasets remains to be done, and until such validation is provided, the central claim of high sentiment alignment is not fully supported.","major_comments":[{"comment":"The Pearson correlations that support the central claim (0.93, 0.96, 0.97 in Table IV) are computed between the target score and a score produced by gpt-4o-mini using the 'Complete' prompt. For Supply-Word, the target score itself is also generated by gpt-4o-mini (Section III-A), so the correlation is essentially a self-consistency check of the same model family. For Word+Review and Review+Word, the target is externally supplied, but the evaluator is still the same model that wrote the review and selected the word; high correlations can therefore reflect the model's ability to follow its own scoring rubric rather than agreement with human sentiment. The paper's future-work section (Section V) concedes that direct comparison with human PDT datasets remains to be done. In light of this, the abstract's claim of 'high sentiment alignment' and the statement in Section IV-B that Review+Word and Supply-Word are 'more reliable for applications requiring precise sentiment reflection' are not supported by the current evidence. I recommend either tempering these claims or adding an external validation using human annotations or existing human PDT datasets.","section":"Section III-B.1 and Table IV"}],"minor_comments":[{"comment":"The phrase 'to to produce' contains a duplicated 'to'; it should read 'to produce'.","section":"Section I, RQ1"},{"comment":"'challenges inherit in data collection' should be 'challenges inherent in data collection'.","section":"Section II"},{"comment":"In the discussion of Word+Review and Review+Word, 'gpt-4o-min' is a typo for 'gpt-4o-mini', and 'resent' should be 're-sent'.","section":"Section III-A"},{"comment":"The 'Complete' prompt reads 'where is 0.00 is a completely negative sentiment'; the 'is' after 'where' should be removed.","section":"Table II"},{"comment":"The finding that 79.3% of Supply-Word reviews start with 'I recently' while the diversity metrics (CR, CR-POS, HS, NDS) indicate the greatest diversity is counterintuitive. The paper acknowledges dataset-size differences as a caveat, but it would help to explain why such high repetition does not penalize the diversity scores.","section":"Section IV-C and Table V"},{"comment":"The tStat column is reported without defining the test, the null hypothesis, or the expected range; the text says 'out of range tStat values warrant further exploration' but does not clarify what makes them out of range. Please add a definition and interpretation.","section":"Section IV-B, Table IV"},{"comment":"The paragraph about OpenAI's lack of persistent session connections appears in Future Work but is really an infrastructure/cost detail; relocating it to Section IV-D would improve readability.","section":"Section V, cost paragraph"}],"recommendation":"major_revision","confidential_remarks":"The paper's main contribution is a reproducible synthetic-data-generation pipeline with open datasets and careful cost reporting. The load-bearing issue is the self-referential evaluation; without human validation or comparison to existing human PDT datasets, the headline correlations are not evidence of real-world sentiment fidelity. The authors themselves indicate the missing comparison in the future-work section, so the manuscript's own text supports the need for major revision. I would encourage the editor to ask for either (a) an external validation study or (b) a substantially revised presentation that reframes the results as demonstrating internal consistency only, with claims about reliability explicitly deferred."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper gives you three concrete prompt recipes for generating synthetic Product Desirability Toolkit reviews with gpt-4o-mini, and the datasets, prompts, and cost figures are all public and reproducible. The headline Pearson correlations (0.93–0.97) are real but they measure the model agreeing with itself, not agreement with human sentiment. The authors acknowledge this in the future-work section, so the paper is an honest initial examination rather than a finished claim.\n\nWhat is genuinely new: applying LLM-based synthetic data generation to the PDT format is not in the cited literature, and having 1000-row synthetic PDT datasets in the open is useful for people developing sentiment-scoring algorithms for UX work. I also give credit for the cost transparency — $0.06 to $0.11 per 1000 rows, with timing estimates — and for the explicit discussion of word-coverage failures and positive-sentiment bias.\n\nThe soft spots, in order of size. First, the evaluation is circular: gpt-4o-mini writes the reviews and then scores them with the 'Complete' prompt. For Supply-Word the target score itself is generated by the same model, so the 0.97 correlation is largely a same-model consistency check. The Supply-Word-4o comparison does not fix this because it is still LLM scoring, just a different one. Second, no human benchmark or external dataset is used; the paper's own future-work section says direct comparison with human PDT datasets remains to be done. That is the missing piece that would turn an internal-consistency result into a practical recommendation. Third, minor issues: no confidence intervals on the Pearson values, tStat values are out of range, and the diversity comparison is confounded by dataset size differences. None of these are damning; the paper is appropriately scoped.\n\nWho should read this: UX researchers who want to bootstrap PDT algorithm development with synthetic data, and people working on LLM-as-data-generator evaluation methodology. The central claim about 'reliable for applications requiring precise sentiment reflection' is not supported yet, but the recipes are worth having.\n\nRecommendation: send it to peer review. A serious referee can push for a human-labeled comparison set and for the authors to soften the reliability claim. That is a tractable revision, not a fundamental flaw.","headline":"Useful, honest synthetic-PDT recipes with public data; the headline alignment numbers are self-consistency, not human fidelity.","tokens_in":14631,"tokens_out":2205,"would_cite":true,"duration_ms":22029,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that gpt-4o-mini can synthesize Product Desirability Toolkit word-and-review pairs whose sentiment matches the requested score, with correlations from 0.93 to 0.97.","keywords":["synthetic data generation","large language models","Product Desirability Toolkit","sentiment analysis","gpt-4o-mini","user-centered design","text diversity","data augmentation"],"falsifier":"Have human raters score the same 3000 generated word/review pairs on the 0.00-1.00 sentiment scale used by the 'Complete' prompt and compute the correlation between human scores and the target scores; if that correlation is far below the reported 0.93-0.97, the alignment result is an artifact of self-scoring.","tokens_in":13640,"feed_emoji":"🤖","tokens_out":9888,"duration_ms":92415,"temperature":0.7,"pith_summary":"The paper argues that a small, inexpensive language model can produce synthetic Product Desirability Toolkit (PDT) datasets—the word-plus-explanation pairs used to gauge whether users find a product desirable—at a scale real user testing rarely reaches. Using gpt-4o-mini and three prompt templates, it generated 1000 hypothetical software reviews per method and found sentiment scores tracking the intended scores with Pearson correlations from 0.93 to 0.97. Costs came to roughly $0.06 to $0.11 per thousand rows, which would make million-row datasets feasible for about $60 to $110. The credibility of this result rests on whether the model's own sentiment scores are a good proxy for human judgment, since the same model did most of the scoring; the authors flag direct comparison with human PDT datasets as future work.","feed_headline":"Match target sentiment: LLM product reviews hit 0.93-0.97","feed_subtitle":"For 6 to 11 cents per 1,000 reviews, synthetic user-experience data becomes scalable.","key_machinery":"The central object is the Product Desirability Toolkit (PDT), a card-based method where users describe an experience by choosing words from a fixed 118-word reaction-card list and optionally explaining each choice. The machinery is a generate-and-score loop built on gpt-4o-mini: three prompt templates (Word+Review, Review+Word, Supply-Word) create 1000 word/review pairs each, and a 'Complete' scoring prompt assigns each pair a sentiment score between 0.00 and 1.00. Alignment is then measured by the Pearson correlation and absolute differences between those evaluated scores and the target scores, while diversity is measured by compression ratio, part-of-speech compression, homogenization, and n-gram diversity.","core_discovery":"The paper's central finding is that gpt-4o-mini can synthesize Product Desirability Toolkit word/review pairs whose evaluated sentiment closely tracks the intended target sentiment: Pearson correlations were 0.93 for Word+Review, 0.96 for Review+Word, and 0.97 for Supply-Word. The authors single out Review+Word and Supply-Word as more reliable for applications needing precise sentiment reflection. Supply-Word attained full coverage of the 118 PDT reaction cards and the most diverse text by compression-ratio, part-of-speech, homogenization, and n-gram measures, but it was the most verbose and most expensive method at about $0.11 per 1000 rows and showed the strongest positive bias. They attribute part of that positivity to the PDT word list's inherent 60/20/20 positive/negative/neutral makeup and note that differences between target and evaluated scores could arise in either the generation or the scoring phase; a gpt-4o check scoring of Supply-Word produced nearly identical patterns, suggesting generation, not scoring, drives most of the bias.","pith_inferences":["If the self-scoring caveat is resolved by human evaluation, the same generate-and-score loop would make it practical to benchmark LLM-based sentiment quantifiers against million-row PDT datasets, something the field currently lacks.","Because the prompts only name a product, the methods transfer to any product category or to other adjective-card toolkits, not just software.","The 79.3% of Supply-Word reviews starting with 'I recently' shows that standard diversity scores can look healthy while surface phrasing is repetitive; adding a naturalness or template-detection metric could change which method looks best.","A direct test of the bias question would be to weight target scores more heavily in the mid-range and see whether evaluated scores spread out, a possibility the authors mention but do not test."],"forward_implications":["Synthetic PDT datasets of 1000 rows can be produced for between $0.06 and $0.11 depending on method, so million-row datasets would cost roughly $60 to $110 in API charges.","Review+Word and Supply-Word, with Pearson correlations of 0.96 and 0.97, are the methods to prefer when the goal is precise reflection of a target sentiment.","Supply-Word gives the best coverage of the 118 PDT reaction cards and the most diverse vocabulary, at the price of higher cost and more positive bias.","Word+Review is the cheapest and fastest (about 1.5 seconds per review) but has the weakest alignment and sometimes invents words outside the PDT list; client-side validation can catch that.","All three methods carry a positive-sentiment bias, so synthetic data of this kind is recommended for internal research, not for real-world product, marketing, or public-relations decisions."],"supporting_citations":[{"why":"Introduces the 118-word product reaction card set that defines the target word list.","marker":"[22]"},{"why":"Introduces the Product Desirability Toolkit and its documented 60/20/20 sentiment split, used to explain the positive bias.","marker":"[21]"},{"why":"Supplies the LLM-based sentiment-scoring approach on which the 'Complete' prompt is based.","marker":"[37]"},{"why":"Introduces the gpt-4o-mini model used for generation and scoring.","marker":"[41]"},{"why":"Supplies the diversity metrics used to compare the three generated datasets.","marker":"[43]"},{"why":"Releases the three synthetic PDT datasets the paper analyzes.","marker":"[44]"},{"why":"Documents the token prices used to estimate generation costs.","marker":"[45]"},{"why":"Shows the small size of real PDT datasets (n=56) that motivates synthetic data.","marker":"[26]"}],"fun_headline_variants":["LLMs match sentiment in synthetic product reviews","Synthetic reviews hit 0.93-0.97 sentiment match","Low-cost LLM builds product desirability datasets","GPT-4o-mini synthesizes reviews with 0.97 sentiment correlation","Synthetic product reviews: 0.93-0.97 sentiment at low cost"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument's load-bearing premise is that gpt-4o-mini's sentiment scores are a trustworthy stand-in for human judgment, since without that, the reported correlations only show the model agreeing with itself.","fun_headline_variants_meta":{"raw":{"variants":["LLMs match sentiment in synthetic product reviews","Synthetic reviews hit 0.93-0.97 sentiment match","Low-cost LLM builds product desirability datasets","GPT-4o-mini synthesizes reviews with 0.97 sentiment correlation","Synthetic product reviews: 0.93-0.97 sentiment at low cost"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000492,"raw_usage":{"total_tokens":2407,"prompt_tokens":925,"completion_tokens":1482,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":1391}},"tokens_in":541,"tokens_out":1482,"duration_ms":12003,"temperature":1.0,"reasoning_tokens":1391,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:21:14.198794+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human raters score the same 3000 generated word/review pairs on the 0.00-1.00 sentiment scale used by the 'Complete' prompt and compute the correlation between human scores and the target scores; if that correlation is far below the reported 0.93-0.97, the alignment result is an artifact of self-scoring.","supporting_citations":[{"cited_title":"Benedek and T","cited_arxiv_id":null,"evidence_quote":"Introduces the 118-word product reaction card set that defines the target word list."},{"cited_title":"Using LLMs to establish implicit user sentiment of software desirability,","cited_arxiv_id":null,"evidence_quote":"Supplies the LLM-based sentiment-scoring approach on which the 'Complete' prompt is based."},{"cited_title":"[Online]","cited_arxiv_id":null,"evidence_quote":"Introduces the gpt-4o-mini model used for generation and scoring."},{"cited_title":"Hastings, S","cited_arxiv_id":null,"evidence_quote":"Releases the three synthetic PDT datasets the paper analyzes."},{"cited_title":"[Online]","cited_arxiv_id":null,"evidence_quote":"Documents the token prices used to estimate generation costs."},{"cited_title":"CARMA: Assessing usability through a non-biased online survey technique,","cited_arxiv_id":null,"evidence_quote":"Shows the small size of real PDT datasets (n=56) that motivates synthetic data."}],"review_version":1}