REVIEW 4 major objections 7 minor 12 references
This paper claims that AI-generated product ideas outperform student ideas in the upper tail of quality: GPT-4 ideas are seven times more likely to rank among the top 10% of pooled ideas, despite being perceived as less novel and more simil
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 05:48 UTC pith:BS4YZI5E
load-bearing objection Solid empirical study on LLM ideation, but the 7:1 top-decile edge is partly confounded by GPT-4's writing fluency, so treat it as an upper bound. the 4 major comments →
Using Large Language Models for Idea Generation in Innovation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the discovery is that GPT-4-generated product ideas dominate student-generated ideas in the upper tail of quality: among the 400 ideas pooled, 35 of the top 40 (87.5%) came from GPT-4, a 7:1 ratio that a chi-square test rejects equal representation. Average purchase intent is also higher for AI ideas (0.48 vs 0.40, with few-shot 0.49). The same data show costs: human raters find AI ideas less novel (0.36 vs 0.41) and text-embedding similarity is higher among AI ideas (average cosine 0.42 vs 0.22), with a capture-recapture estimate of roughly 700-950 discoverable AI ideas in this space. The authors interpret the combination as evidence that AI ideation shifts the val
What carries the argument
The design's load-bearing piece is the idea-as-lottery-ticket model: innovation performance is governed by the upper tail of an unknown quality distribution, so the paper compares the 90th percentile rather than the mean. Quality is measured by average purchase-intent probability from a five-box survey, novelty by human ratings, and diversity by cosine similarity between text embeddings with a 0.8 overlap threshold. A capture-recapture model translates overlap counts into an estimate of the total discoverable idea space. These measures together let the paper separate the average effect (Study 1), the diversity cost (Study 2), and the top-decile advantage (Study 3).
Load-bearing premise
The entire comparison rests on treating average purchase-intent ratings by survey respondents as the value of an idea; the authors note that GPT-4's fluent phrasing, not the idea itself, may be driving those ratings.
What would settle it
Take the 200 human ideas, have GPT-4 rewrite each in a uniform LLM-style pitch, and re-run the purchase-intent survey. If the average gap and the 7:1 top-decile ratio shrink to near zero, the apparent AI advantage is a presentation effect. Alternatively, generate a fresh set of ideas with explicit novelty-seeking prompts; if they still dominate the top decile, the quality advantage is substantive.
If this is right
- If the 7:1 top-decile advantage holds, an organization can get more exceptional candidate ideas per idea generated by using an LLM rather than a comparable number of student or novice ideators.
- Because GPT-4 can produce the same number of ideas much faster, the true advantage in a fixed time budget is larger than 7:1; the paper calls its estimate conservative.
- Few-shot prompting raises average purchase intent slightly but increases idea overlap, so the choice of prompt is a quality-versus-diversity tradeoff.
- Lower perceived novelty does not strongly correlate with purchase intent, so the novelty deficit may not reduce the commercial value of the best ideas.
- The shift from generation to evaluation emerges as a practical consequence: if high-quality drafts are cheap and abundant, resources should move to screening and refinement.
Where Pith is reading between the lines
- The purchase-intent measurement may reward GPT-4's fluent 'pitch' rather than the underlying concept; if so, the 7:1 ratio partly measures writing style, a possibility the paper flags.
- The capture-recapture estimate implies a finite idea space of roughly 700-1,000 high-quality AI ideas for this product category; a testable extension is whether explicit diversity-forcing prompts enlarge that space.
- If the results generalize beyond college products under $50, the comparative advantage may be largest in well-documented consumer domains and smaller in specialized fields where training data are thin.
- A cheap falsifying check: ask GPT-4 to rewrite the human ideas in its own presentation style; if the purchase-intent gap disappears, the substance advantage vanishes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript compares product-idea generation by university students (N=200) with GPT-4 ideas generated via zero-shot and few-shot prompting (N=100 each). Purchase-intent ratings from online respondents are used as the quality measure, text embeddings measure idea similarity, and human raters assess novelty. The main claims are that AI ideas have higher average purchase intent (0.48 vs 0.40), are somewhat less novel and more mutually similar, and account for 35 of the top 40 pooled ideas (a 7:1 advantage). The authors argue that this 7:1 ratio is conservative because it excludes AI's higher generation speed.
Significance. This is one of the first large-sample studies of LLM-based ideation in a product-development context. It uses a standard market-research quality metric, a documented experimental protocol, and a range of robustness checks (alternative weights, CLMMs, quantile regression, alternative similarity thresholds). If the purchase-intent measure is accepted as reflecting idea substance, the top-decile result is practically important for innovation management. The paper also transparently acknowledges several limitations, including the student-only human baseline and the possibility that GPT-4's writing style drives the ratings. The main unresolved issue is whether the measured advantage reflects idea quality or presentation style, which directly bears on the paper's central claim.
major comments (4)
- [§8.1; Studies 1 and 3 (Fig. 4, Table 3)] The central quality claim depends on purchase-intent ratings elicited from text descriptions. The manuscript's own §8.1 concedes that "it is possible that the writing style ('pitch') convinces the customers rather than the idea itself." GPT-4 texts are uniform, fluent, and market-savvy, while student texts vary in clarity and polish. The robustness tests in §8.2 change weights and model families but never hold style fixed while varying idea content. Consequently, the average advantage (0.48 vs 0.40) and the 7:1 top-decile result may partly be a presentation effect. The authors should provide a style-controlled comparison—for example, rewriting student ideas in a uniform template or using an LLM to rephrase human ideas—or otherwise bound the effect. Without this, the headline claim about idea quality is not established.
- [§6.1, §6.2, Table 1] The overlap definition is stated inconsistently. §6.1 says an idea is overlapping if its pairwise similarity to any previously added idea exceeds θ=0.8, but Table 1's note defines the fraction as ideas whose average pairwise similarity to all other ideas exceeds 0.8. These are different statistics and can produce different overlap counts. The capture-recapture estimates in §6.2 (T=966 and 680) are fitted to observed overlap counts of 5 and 7 out of 100, so the exact counting rule matters. Please specify the algorithm used, report counts under both definitions, and, if the sequential definition is intended, state the order used.
- [§7, Table 2] Table 2 lists 42 ideas in the "top 10%" table, but the text reports "top 40 ideas" with 35 (87.5%) GPT-4-generated. If ties at purchase intent 0.61 produce 42 ideas, the table and the counts must be reconciled; as written, the reader cannot verify the headline 35/40 statistic from the displayed data. This is a direct issue for the paper's central quantitative claim.
- [§6.2, Eq. (1)-(2)] The capture-recapture estimate of the opportunity space T is based on only 5 and 7 overlapping observations out of 100. No standard error or bootstrap CI is reported, and the human pool, with zero observed overlaps, forces T=∞, so the "fewer discoverable ideas" comparison is not directly estimated for the human baseline. Given the importance of this quantity for Hypothesis 2a, please provide uncertainty bounds and discuss the sensitivity of T to the overlap threshold and to the zero-count problem.
minor comments (7)
- [§3] Typos: 'dimentions' and 'dimention' should be 'dimensions'.
- [§4] Typo: 'Speifically' should be 'Specifically'.
- [§6.1 and Figure 2] η² 95% CIs are reported as negative intervals (e.g., [-0.210, -0.204]); eta-squared cannot be negative. Please check the effect-size routine.
- [§6.1] The ANOVA reports F(2, 29598); based on the number of pairwise comparisons (29800 in total), the residual df should be 29,797 if all pairs are used. Please verify and correct.
- [Figure 3 note vs §6.3] The figure note credits mTurk, but the text says the novelty ratings were collected on Prolific. Align the descriptions.
- [Figure 3 note and References] The figure note cites Kwon, Kim, and Lee (2009), but this reference does not appear in the reference list. Add it or remove the citation.
- [§5 and References] The text cites 'Jameson and Bass (1989)' but the reference list has 'Jamieson and Bass (1989)'. Standardize the spelling.
Circularity Check
No circularity: the claims rest on independent empirical measurement; self-citations are methodological anchors, not derivation inputs.
full rationale
The paper's central claims are empirical comparisons of survey-measured purchase intent and novelty across three idea pools. The 7:1 top-decile result (35 of 40 top ideas from GPT-4) is a direct count of independently elicited purchase-intent ratings, not a quantity derived from a fitted parameter or defined into existence. The capture-recapture estimate of the opportunity space (T = 966 and 680) is fitted to observed overlap counts and explicitly presented as an estimate, not as a prediction; it is not used to generate the quality comparisons. Self-citations to Kornish and Ulrich (2011, 2014) and Girotra et al. (2010) are used to justify measurement choices and analytical conventions, but those prior results are external empirical benchmarks and do not encode the present paper's outcome. The acknowledged limitation in Section 8.1 that GPT-4's writing style ('pitch') might influence purchase intent is a validity threat to the construct being measured, but it is not a circularity: the comparison is still an empirical comparison of the rated texts, not a derivation that reduces to its inputs. No equation in the paper is equivalent by construction to a fitted parameter or to a self-citation chain. The study is self-contained as an empirical evaluation, so the circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- Capture-recapture decay a (equivalently T = 1/a) =
T ≈ 966 (zero-shot), T ≈ 680 (few-shot)
- Cosine similarity overlap threshold θ =
0.8
- GPT-4 temperature =
0.7
- Few-shot example count =
6
axioms (4)
- domain assumption Purchase-intent probability from five-box survey responses is a valid measure of idea quality and predicts future value.
- domain assumption Cosine similarity of Universal Sentence Encoder embeddings captures semantic overlap between product ideas.
- domain assumption Idea generation follows a capture-recapture process with equally likely draws and exponential decay of unique discoveries.
- domain assumption The 2021 student cohort is an adequate baseline for human idea-generation performance in this task.
read the original abstract
This research evaluates the efficacy of large language models (LLMs) in generating new product ideas. To do so, we compare three pools of ideas for new products targeted toward college students and priced at 50 dollars or less. The first pool of ideas was created by university students in a product design course before the availability of LLMs. The second and third pools of ideas were generated by GPT-4 from OpenAI using zero-shot and few-shot prompting, respectively. We evaluated idea quality using standard market research techniques to predict average purchase intent probability. We used text mining to assess idea similarity and human raters to evaluate idea novelty. We find that AI-generated ideas outperform human-generated ideas in terms of average purchase intent, with few-shot prompting yielding slightly higher intent than zero-shot prompting. However, AI-generated ideas are perceived as less novel and exhibit higher pairwise similarity, particularly with few-shot prompting, indicating a less diverse solution landscape. When focusing on the quality of the best ideas rather than the average ideas, we find that AI-generated ideas are seven times more likely to rank among the top 10 percent of ideas, demonstrating a significant advantage over human-generated ideas. We propose that this seven-to-one advantage is a conservative estimate because it does not account for the greater productivity of AI. Our findings suggest that despite some drawbacks, AI creativity presents a substantial benefit in generating high-quality ideas for new product development.
Reference graph
Works this paper leans on
-
[1]
Bellaiche L, Shahi R, Turpin MH, Ragnhildstveit A, Sprockett S, Barr N, Christensen A, Seli P (2023) Humans versus AI: whether and why we prefer human-created compared to AI-created artwork. Cogn. Research 8(1):42. Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Neelakantan A, et al. (2020) Language Models are Few-Shot Learners. (July
2023
-
[3]
Koenker R, Hallock KF (2001) Quantile Regression
http://arxiv.org/abs/2406.07016. Koenker R, Hallock KF (2001) Quantile Regression. Journal of Economic Perspectives 15(4):143–156. Koivisto M, Grassini S (2023) Best humans still outperform artificial intelligence in a creative divergent thinking task. Sci Rep 13(1):13601. Kornish LJ, Ulrich KT (2011) Opportunity Spaces in Innovation: Empirical Analysis o...
Pith/arXiv arXiv 2001
-
[4]
Osborn AF (1953) Applied imagination (Scribner’S, Oxford, England)
http://arxiv.org/abs/2303.08774. Osborn AF (1953) Applied imagination (Scribner’S, Oxford, England). Rashidi HH, Fennell BD, Albahra S, Hu B, Gorbett T (2023) The ChatGPT conundrum: Human- generated scientific manuscripts misidentified as AI creations by AI text detection tool. Journal of Pathology Informatics 14:100342. Shank DB, Stefanik C, Stuhlsatz C,...
Pith/arXiv arXiv 1953
-
[6]
Siangliulue P, Chan J, Dow SP, Gajos KZ (2016) IdeaHound: Improving Large-scale Collaborative Ideation with Crowd-Powered Real-time Semantic Modeling
https://papers.ssrn.com/abstract=4050940. Siangliulue P, Chan J, Dow SP, Gajos KZ (2016) IdeaHound: Improving Large-scale Collaborative Ideation with Crowd-Powered Real-time Semantic Modeling. Proceedings of the 29th Annual Symposium on User Interface Software and Technology. (ACM, Tokyo Japan), 609–624. Sommer SC, Loch CH (2004) Selectionism and Learning...
2016
-
[10]
Wang H, Zou J, Mozer M, Goyal A, Lamb A, Zhang L, Su WJ, et al
http://arxiv.org/abs/2310.06202. Wang H, Zou J, Mozer M, Goyal A, Lamb A, Zhang L, Su WJ, et al. (2024) Can AI Be as Creative as Humans? (January
Pith/arXiv arXiv 2024
-
[12]
Terwiesch C, Ulrich K (2023) The innovation tournament handbook: a step-by-step guide to finding exceptional solutions to any challenge (Wharton School Press, Philadelphia, PA)
https://www.ft.com/content/591ad272-6419-4f2c-9935-caff1d670f08. Terwiesch C, Ulrich K (2023) The innovation tournament handbook: a step-by-step guide to finding exceptional solutions to any challenge (Wharton School Press, Philadelphia, PA). Terwiesch C, Ulrich KT (2009) Innovation tournaments: creating and selecting exceptional opportunities (Harvard Bu...
2023
-
[13]
Weitzman ML (1979) Optimal Search for the Best Alternative
http://arxiv.org/abs/2201.11903. Weitzman ML (1979) Optimal Search for the Best Alternative. Econometrica 47(3):641. Zhou E, Lee D (2024) Generative artificial intelligence, human creativity, and art Harding M, ed. PNAS Nexus 3(3):pgae052. Zlatkov D, Ens J, Pasquier P (2023) Searching for Human Bias Against AI-Composed Music. Artificial Intelligence in Mu...
Pith/arXiv arXiv 1979
-
[15]
Doshi AR, Hauser OP (2024) Generative AI enhances individual creativity but reduces the collective diversity of novel content
https://papers.ssrn.com/abstract=4573321. Doshi AR, Hauser OP (2024) Generative AI enhances individual creativity but reduces the collective diversity of novel content. Sci. Adv. 10(28):eadn5290. Girotra K, Terwiesch C, Ulrich KT (2010) Idea Generation and the Quality of the Best Idea. Management Science 56(4):591–605. Goldenberg J, Mazursky D, Solomon S ...
2024
-
[22]
http://arxiv.org/abs/2005.14165. Chao RO, Kavadias S (2008) A Theoretical Framework for Managing the New Product Development Portfolio: When and How to Use Strategic Buckets. Management Science 54(5):907–921. Cochran WG (1978) Laplace’s Ratio Estimator. David HA, ed. Contributions to Survey Sampling and Applied Statistics. (Academic Press), 3–10. Connolly...
Pith/arXiv arXiv 2005
-
[23]
http://arxiv.org/abs/2301.11305. Meincke L, Carton A (2024) Beyond Multiple Choice: The Role of Large Language Models in Educational Simulations. (May
Pith/arXiv arXiv 2024
-
[25]
http://arxiv.org/abs/2401.01623. Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, Chi E, Le Q, Zhou D (2023) Chain-of- Thought Prompting Elicits Reasoning in Large Language Models. (January
Pith/arXiv arXiv 2023
-
[26]
OpenAI, Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman FL, et al
https://papers.ssrn.com/abstract=4873537. OpenAI, Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman FL, et al. (2024) GPT-4 Technical Report. (March
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.