Pith. sign in

REVIEW 4 major objections 7 minor 12 references

This paper claims that AI-generated product ideas outperform student ideas in the upper tail of quality: GPT-4 ideas are seven times more likely to rank among the top 10% of pooled ideas, despite being perceived as less novel and more simil

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 05:48 UTC pith:BS4YZI5E

load-bearing objection Solid empirical study on LLM ideation, but the 7:1 top-decile edge is partly confounded by GPT-4's writing fluency, so treat it as an upper bound. the 4 major comments →

arxiv 2607.27553 v1 pith:BS4YZI5E submitted 2026-07-30 cs.AI cs.CLecon.GNq-fin.EC

Using Large Language Models for Idea Generation in Innovation

classification cs.AI cs.CLecon.GNq-fin.EC
keywords LLM ideationidea generationGPT-4purchase intentcreativityinnovation tournamentsidea diversityproduct design
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether a large language model can generate product ideas that beat human-generated ones. It compares 200 student ideas for college-student products under $50 with 200 GPT-4 ideas (100 zero-shot, 100 few-shot). Based on purchase-intent surveys, it finds AI ideas score higher on average (0.48 vs 0.40), especially with few-shot prompting, and are seven times more likely to land in the top 10% of the pooled ideas—35 of the top 40 are AI-generated. The AI ideas are also perceived as less novel and are more similar to each other. The paper concludes that for innovation, where only the best ideas matter, the AI advantage is substantial despite lower diversity.

Core claim

On the paper's own terms, the discovery is that GPT-4-generated product ideas dominate student-generated ideas in the upper tail of quality: among the 400 ideas pooled, 35 of the top 40 (87.5%) came from GPT-4, a 7:1 ratio that a chi-square test rejects equal representation. Average purchase intent is also higher for AI ideas (0.48 vs 0.40, with few-shot 0.49). The same data show costs: human raters find AI ideas less novel (0.36 vs 0.41) and text-embedding similarity is higher among AI ideas (average cosine 0.42 vs 0.22), with a capture-recapture estimate of roughly 700-950 discoverable AI ideas in this space. The authors interpret the combination as evidence that AI ideation shifts the val

What carries the argument

The design's load-bearing piece is the idea-as-lottery-ticket model: innovation performance is governed by the upper tail of an unknown quality distribution, so the paper compares the 90th percentile rather than the mean. Quality is measured by average purchase-intent probability from a five-box survey, novelty by human ratings, and diversity by cosine similarity between text embeddings with a 0.8 overlap threshold. A capture-recapture model translates overlap counts into an estimate of the total discoverable idea space. These measures together let the paper separate the average effect (Study 1), the diversity cost (Study 2), and the top-decile advantage (Study 3).

Load-bearing premise

The entire comparison rests on treating average purchase-intent ratings by survey respondents as the value of an idea; the authors note that GPT-4's fluent phrasing, not the idea itself, may be driving those ratings.

What would settle it

Take the 200 human ideas, have GPT-4 rewrite each in a uniform LLM-style pitch, and re-run the purchase-intent survey. If the average gap and the 7:1 top-decile ratio shrink to near zero, the apparent AI advantage is a presentation effect. Alternatively, generate a fresh set of ideas with explicit novelty-seeking prompts; if they still dominate the top decile, the quality advantage is substantive.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the 7:1 top-decile advantage holds, an organization can get more exceptional candidate ideas per idea generated by using an LLM rather than a comparable number of student or novice ideators.
  • Because GPT-4 can produce the same number of ideas much faster, the true advantage in a fixed time budget is larger than 7:1; the paper calls its estimate conservative.
  • Few-shot prompting raises average purchase intent slightly but increases idea overlap, so the choice of prompt is a quality-versus-diversity tradeoff.
  • Lower perceived novelty does not strongly correlate with purchase intent, so the novelty deficit may not reduce the commercial value of the best ideas.
  • The shift from generation to evaluation emerges as a practical consequence: if high-quality drafts are cheap and abundant, resources should move to screening and refinement.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The purchase-intent measurement may reward GPT-4's fluent 'pitch' rather than the underlying concept; if so, the 7:1 ratio partly measures writing style, a possibility the paper flags.
  • The capture-recapture estimate implies a finite idea space of roughly 700-1,000 high-quality AI ideas for this product category; a testable extension is whether explicit diversity-forcing prompts enlarge that space.
  • If the results generalize beyond college products under $50, the comparative advantage may be largest in well-documented consumer domains and smaller in specialized fields where training data are thin.
  • A cheap falsifying check: ask GPT-4 to rewrite the human ideas in its own presentation style; if the purchase-intent gap disappears, the substance advantage vanishes.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The manuscript compares product-idea generation by university students (N=200) with GPT-4 ideas generated via zero-shot and few-shot prompting (N=100 each). Purchase-intent ratings from online respondents are used as the quality measure, text embeddings measure idea similarity, and human raters assess novelty. The main claims are that AI ideas have higher average purchase intent (0.48 vs 0.40), are somewhat less novel and more mutually similar, and account for 35 of the top 40 pooled ideas (a 7:1 advantage). The authors argue that this 7:1 ratio is conservative because it excludes AI's higher generation speed.

Significance. This is one of the first large-sample studies of LLM-based ideation in a product-development context. It uses a standard market-research quality metric, a documented experimental protocol, and a range of robustness checks (alternative weights, CLMMs, quantile regression, alternative similarity thresholds). If the purchase-intent measure is accepted as reflecting idea substance, the top-decile result is practically important for innovation management. The paper also transparently acknowledges several limitations, including the student-only human baseline and the possibility that GPT-4's writing style drives the ratings. The main unresolved issue is whether the measured advantage reflects idea quality or presentation style, which directly bears on the paper's central claim.

major comments (4)
  1. [§8.1; Studies 1 and 3 (Fig. 4, Table 3)] The central quality claim depends on purchase-intent ratings elicited from text descriptions. The manuscript's own §8.1 concedes that "it is possible that the writing style ('pitch') convinces the customers rather than the idea itself." GPT-4 texts are uniform, fluent, and market-savvy, while student texts vary in clarity and polish. The robustness tests in §8.2 change weights and model families but never hold style fixed while varying idea content. Consequently, the average advantage (0.48 vs 0.40) and the 7:1 top-decile result may partly be a presentation effect. The authors should provide a style-controlled comparison—for example, rewriting student ideas in a uniform template or using an LLM to rephrase human ideas—or otherwise bound the effect. Without this, the headline claim about idea quality is not established.
  2. [§6.1, §6.2, Table 1] The overlap definition is stated inconsistently. §6.1 says an idea is overlapping if its pairwise similarity to any previously added idea exceeds θ=0.8, but Table 1's note defines the fraction as ideas whose average pairwise similarity to all other ideas exceeds 0.8. These are different statistics and can produce different overlap counts. The capture-recapture estimates in §6.2 (T=966 and 680) are fitted to observed overlap counts of 5 and 7 out of 100, so the exact counting rule matters. Please specify the algorithm used, report counts under both definitions, and, if the sequential definition is intended, state the order used.
  3. [§7, Table 2] Table 2 lists 42 ideas in the "top 10%" table, but the text reports "top 40 ideas" with 35 (87.5%) GPT-4-generated. If ties at purchase intent 0.61 produce 42 ideas, the table and the counts must be reconciled; as written, the reader cannot verify the headline 35/40 statistic from the displayed data. This is a direct issue for the paper's central quantitative claim.
  4. [§6.2, Eq. (1)-(2)] The capture-recapture estimate of the opportunity space T is based on only 5 and 7 overlapping observations out of 100. No standard error or bootstrap CI is reported, and the human pool, with zero observed overlaps, forces T=∞, so the "fewer discoverable ideas" comparison is not directly estimated for the human baseline. Given the importance of this quantity for Hypothesis 2a, please provide uncertainty bounds and discuss the sensitivity of T to the overlap threshold and to the zero-count problem.
minor comments (7)
  1. [§3] Typos: 'dimentions' and 'dimention' should be 'dimensions'.
  2. [§4] Typo: 'Speifically' should be 'Specifically'.
  3. [§6.1 and Figure 2] η² 95% CIs are reported as negative intervals (e.g., [-0.210, -0.204]); eta-squared cannot be negative. Please check the effect-size routine.
  4. [§6.1] The ANOVA reports F(2, 29598); based on the number of pairwise comparisons (29800 in total), the residual df should be 29,797 if all pairs are used. Please verify and correct.
  5. [Figure 3 note vs §6.3] The figure note credits mTurk, but the text says the novelty ratings were collected on Prolific. Align the descriptions.
  6. [Figure 3 note and References] The figure note cites Kwon, Kim, and Lee (2009), but this reference does not appear in the reference list. Add it or remove the citation.
  7. [§5 and References] The text cites 'Jameson and Bass (1989)' but the reference list has 'Jamieson and Bass (1989)'. Standardize the spelling.

Circularity Check

0 steps flagged

No circularity: the claims rest on independent empirical measurement; self-citations are methodological anchors, not derivation inputs.

full rationale

The paper's central claims are empirical comparisons of survey-measured purchase intent and novelty across three idea pools. The 7:1 top-decile result (35 of 40 top ideas from GPT-4) is a direct count of independently elicited purchase-intent ratings, not a quantity derived from a fitted parameter or defined into existence. The capture-recapture estimate of the opportunity space (T = 966 and 680) is fitted to observed overlap counts and explicitly presented as an estimate, not as a prediction; it is not used to generate the quality comparisons. Self-citations to Kornish and Ulrich (2011, 2014) and Girotra et al. (2010) are used to justify measurement choices and analytical conventions, but those prior results are external empirical benchmarks and do not encode the present paper's outcome. The acknowledged limitation in Section 8.1 that GPT-4's writing style ('pitch') might influence purchase intent is a validity threat to the construct being measured, but it is not a circularity: the comparison is still an empirical comparison of the rated texts, not a derivation that reduces to its inputs. No equation in the paper is equivalent by construction to a fitted parameter or to a self-citation chain. The study is self-contained as an empirical evaluation, so the circularity score is 0.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The core result is empirical and depends on measurement and modeling assumptions rather than on free parameters in a derivation. The only fitted parameter is the capture-recapture decay a (inverse T), and the overlap threshold is a hand-set design choice. No invented entities are introduced.

free parameters (4)
  • Capture-recapture decay a (equivalently T = 1/a) = T ≈ 966 (zero-shot), T ≈ 680 (few-shot)
    Fit to a single observed unique fraction (u(100)=0.95 and 0.93) under the exponential model; no confidence interval is reported in Section 6.2.
  • Cosine similarity overlap threshold θ = 0.8
    Chosen by authors after pairwise experimentation; robustness checks in Section 8.2.2 show counts vary materially at θ = 0.7, 0.75, and 0.85.
  • GPT-4 temperature = 0.7
    Hand-set to balance randomness and coherence; affects the exact sample of ideas generated and is not optimized for this task.
  • Few-shot example count = 6
    Chosen due to context-window limits and prior few-shot learning results, not fitted to this experiment.
axioms (4)
  • domain assumption Purchase-intent probability from five-box survey responses is a valid measure of idea quality and predicts future value.
    Adopted from Kornish and Ulrich (2014); load-bearing because 'quality' is defined as purchase intent throughout the paper.
  • domain assumption Cosine similarity of Universal Sentence Encoder embeddings captures semantic overlap between product ideas.
    Used to operationalize flexibility/overlap in Study 2; depends on the choice of embedding model.
  • domain assumption Idea generation follows a capture-recapture process with equally likely draws and exponential decay of unique discoveries.
    Used in Section 6.2 to estimate T=966 and T=680; this model assumption is not independently tested in the paper.
  • domain assumption The 2021 student cohort is an adequate baseline for human idea-generation performance in this task.
    The authors acknowledge professionals might do better; the 7:1 claim depends on this baseline being representative enough.

pith-pipeline@v1.3.0-daily-deepseek · 21199 in / 10662 out tokens · 111849 ms · 2026-08-01T05:48:06.084623+00:00 · methodology

0 comments
read the original abstract

This research evaluates the efficacy of large language models (LLMs) in generating new product ideas. To do so, we compare three pools of ideas for new products targeted toward college students and priced at 50 dollars or less. The first pool of ideas was created by university students in a product design course before the availability of LLMs. The second and third pools of ideas were generated by GPT-4 from OpenAI using zero-shot and few-shot prompting, respectively. We evaluated idea quality using standard market research techniques to predict average purchase intent probability. We used text mining to assess idea similarity and human raters to evaluate idea novelty. We find that AI-generated ideas outperform human-generated ideas in terms of average purchase intent, with few-shot prompting yielding slightly higher intent than zero-shot prompting. However, AI-generated ideas are perceived as less novel and exhibit higher pairwise similarity, particularly with few-shot prompting, indicating a less diverse solution landscape. When focusing on the quality of the best ideas rather than the average ideas, we find that AI-generated ideas are seven times more likely to rank among the top 10 percent of ideas, demonstrating a significant advantage over human-generated ideas. We propose that this seven-to-one advantage is a conservative estimate because it does not account for the greater productivity of AI. Our findings suggest that despite some drawbacks, AI creativity presents a substantial benefit in generating high-quality ideas for new product development.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

12 extracted references · 7 linked inside Pith

  1. [1]

    Bellaiche L, Shahi R, Turpin MH, Ragnhildstveit A, Sprockett S, Barr N, Christensen A, Seli P (2023) Humans versus AI: whether and why we prefer human-created compared to AI-created artwork. Cogn. Research 8(1):42. Brown TB, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Neelakantan A, et al. (2020) Language Models are Few-Shot Learners. (July

  2. [3]

    Koenker R, Hallock KF (2001) Quantile Regression

    http://arxiv.org/abs/2406.07016. Koenker R, Hallock KF (2001) Quantile Regression. Journal of Economic Perspectives 15(4):143–156. Koivisto M, Grassini S (2023) Best humans still outperform artificial intelligence in a creative divergent thinking task. Sci Rep 13(1):13601. Kornish LJ, Ulrich KT (2011) Opportunity Spaces in Innovation: Empirical Analysis o...

  3. [4]

    Osborn AF (1953) Applied imagination (Scribner’S, Oxford, England)

    http://arxiv.org/abs/2303.08774. Osborn AF (1953) Applied imagination (Scribner’S, Oxford, England). Rashidi HH, Fennell BD, Albahra S, Hu B, Gorbett T (2023) The ChatGPT conundrum: Human- generated scientific manuscripts misidentified as AI creations by AI text detection tool. Journal of Pathology Informatics 14:100342. Shank DB, Stefanik C, Stuhlsatz C,...

  4. [6]

    Siangliulue P, Chan J, Dow SP, Gajos KZ (2016) IdeaHound: Improving Large-scale Collaborative Ideation with Crowd-Powered Real-time Semantic Modeling

    https://papers.ssrn.com/abstract=4050940. Siangliulue P, Chan J, Dow SP, Gajos KZ (2016) IdeaHound: Improving Large-scale Collaborative Ideation with Crowd-Powered Real-time Semantic Modeling. Proceedings of the 29th Annual Symposium on User Interface Software and Technology. (ACM, Tokyo Japan), 609–624. Sommer SC, Loch CH (2004) Selectionism and Learning...

  5. [10]

    Wang H, Zou J, Mozer M, Goyal A, Lamb A, Zhang L, Su WJ, et al

    http://arxiv.org/abs/2310.06202. Wang H, Zou J, Mozer M, Goyal A, Lamb A, Zhang L, Su WJ, et al. (2024) Can AI Be as Creative as Humans? (January

  6. [12]

    Terwiesch C, Ulrich K (2023) The innovation tournament handbook: a step-by-step guide to finding exceptional solutions to any challenge (Wharton School Press, Philadelphia, PA)

    https://www.ft.com/content/591ad272-6419-4f2c-9935-caff1d670f08. Terwiesch C, Ulrich K (2023) The innovation tournament handbook: a step-by-step guide to finding exceptional solutions to any challenge (Wharton School Press, Philadelphia, PA). Terwiesch C, Ulrich KT (2009) Innovation tournaments: creating and selecting exceptional opportunities (Harvard Bu...

  7. [13]

    Weitzman ML (1979) Optimal Search for the Best Alternative

    http://arxiv.org/abs/2201.11903. Weitzman ML (1979) Optimal Search for the Best Alternative. Econometrica 47(3):641. Zhou E, Lee D (2024) Generative artificial intelligence, human creativity, and art Harding M, ed. PNAS Nexus 3(3):pgae052. Zlatkov D, Ens J, Pasquier P (2023) Searching for Human Bias Against AI-Composed Music. Artificial Intelligence in Mu...

  8. [15]

    Doshi AR, Hauser OP (2024) Generative AI enhances individual creativity but reduces the collective diversity of novel content

    https://papers.ssrn.com/abstract=4573321. Doshi AR, Hauser OP (2024) Generative AI enhances individual creativity but reduces the collective diversity of novel content. Sci. Adv. 10(28):eadn5290. Girotra K, Terwiesch C, Ulrich KT (2010) Idea Generation and the Quality of the Best Idea. Management Science 56(4):591–605. Goldenberg J, Mazursky D, Solomon S ...

  9. [22]

    Chao RO, Kavadias S (2008) A Theoretical Framework for Managing the New Product Development Portfolio: When and How to Use Strategic Buckets

    http://arxiv.org/abs/2005.14165. Chao RO, Kavadias S (2008) A Theoretical Framework for Managing the New Product Development Portfolio: When and How to Use Strategic Buckets. Management Science 54(5):907–921. Cochran WG (1978) Laplace’s Ratio Estimator. David HA, ed. Contributions to Survey Sampling and Applied Statistics. (Academic Press), 3–10. Connolly...

  10. [23]

    Meincke L, Carton A (2024) Beyond Multiple Choice: The Role of Large Language Models in Educational Simulations

    http://arxiv.org/abs/2301.11305. Meincke L, Carton A (2024) Beyond Multiple Choice: The Role of Large Language Models in Educational Simulations. (May

  11. [25]

    Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, Chi E, Le Q, Zhou D (2023) Chain-of- Thought Prompting Elicits Reasoning in Large Language Models

    http://arxiv.org/abs/2401.01623. Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, Chi E, Le Q, Zhou D (2023) Chain-of- Thought Prompting Elicits Reasoning in Large Language Models. (January

  12. [26]

    OpenAI, Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman FL, et al

    https://papers.ssrn.com/abstract=4873537. OpenAI, Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman FL, et al. (2024) GPT-4 Technical Report. (March