{"id":"73f49755-ff2e-4d1f-a9a4-198101da00ba","arxiv_id":"2412.03963","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In flash fiction and song lyric experiments, generative AI improved creativity most for people high in general cognitive capital, and least for people with specialized writing expertise.","lead":"Two randomized experiments in China compared people writing flash fiction or song lyrics with and without generative AI help. The results suggest AI benefits people with stronger general cognitive and educational backgrounds more, while reducing the creative advantage of domain-specific expertise.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prompt-crafting instruction is bundled into the AI treatment, so the causal moderation claim is not identified.","rationale":"The reader's weakest assumption correctly identifies the treatment confound: the AI condition bundles the AI tool with prompt-crafting instruction, while the control condition receives no comparable instruction. This is the most load-bearing concern because it threatens the causal identification of the central moderation claim, including the replicated negative moderation by specific human capital. The paper's own theoretical mechanisms emphasize AI's lack of agency and expansive knowledge span, not instruction, so the empirical contrast does not cleanly test the theory. I agree with the reader that this confound should be addressed before the causal claim is accepted. Secondary concerns, such as the non-replication of the IQ moderation in Experiment 2 and the severe attrition in Study 2, further weaken the general-human-capital half of the claim, but the prompt-training confound is the more fundamental issue. Since the reader's verdict is already CONDITIONAL and my analysis identifies the same concern, I do not move the verdict.","tokens_in":27158,"tokens_out":7586,"duration_ms":79244,"concrete_test":"Run a three-arm experiment on the same flash-fiction or lyric task: (1) AI with prompt-crafting instruction (current treatment), (2) AI with no prompt-crafting instruction, and (3) no AI but with a matched-length general creative-strategy tutorial to hold instruction time and engagement constant. If the AI-by-general-human-capital positive interaction and the AI-by-specific-human-capital negative interaction are equivalent when comparing arm 1 vs 3 and arm 2 vs 3, the prompt-training confound is not the driver. If the moderation is substantially attenuated or absent in arm 2 vs 3, the reported effects are attributable to the instructional component rather than to generative AI itself. The design and interaction hypotheses should be pre-registered, including equivalence bounds for the key coefficients.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is that the AI-assist treatment is not just 'use GPT-4'; it also includes prompt-crafting instruction. Experiment 1 (Samples and Procedures) says the AI group 'was provided with information about effective prompt crafting,' while the control received only basic fiction-writing requirements; Experiment 2 does the same with 'additional guidance on using generative AI.' The central claim—that generative AI itself enhances the value of general human capital and diminishes domain-specific expertise—requires a contrast between AI and no AI. The actual contrast is between AI-plus-prompt-training and no-AI. The prompt-training component is a plausible active ingredient for the reported interactions: more educated/higher-IQ participants may be better at absorbing and applying prompt-crafting guidance, while domain experts may discount or resist it, generating exactly the positive general-human-capital and negative specific-human-capital interactions. Because the theoretical mechanisms in the paper (AI's lack of agency, expansive knowledge span) do not reference prompt-crafting instruction, the empirical test does not map onto the theory. The causal moderation claim is therefore not identified, and the abstract's language ('AI enhances... diminishes...') overstates what the design supports.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents two randomized experiments examining how generative AI (GPT-4) affects creativity in flash fiction (N=162) and song lyric writing (final N=299), with human capital as moderators. General human capital is measured by education and Raven IQ scores; specific human capital by self-reported writing skill (Study 1) and prior lyric publication (Study 2). The authors report that AI use improved all three creativity ratings in Study 1 but had no significant main effect in Study 2; education and IQ positively moderated AI effects in Study 1, education moderated effects on lyrics-only ratings in Study 2, and specific human capital negatively moderated AI effects in both studies. They interpret these results through an augmentation-automation framework and conclude that AI augments general cognitive skills while devaluing domain expertise.","tokens_in":27352,"tokens_out":7350,"duration_ms":68336,"significance":"Understanding who benefits from generative AI in creative work is an important and timely question, and the two-task design with public raters is a real strength. The consistent negative interaction between AI use and specific human capital across both studies and across lyric and song ratings in Study 2 is a meaningful empirical pattern that goes beyond simple average treatment effects. The authors also provide open data, materials, and analysis code, which facilitates verification. If the causal moderation claim were identified, the paper would be a valuable contribution to human-AI collaboration research. However, the headline claim that AI 'enhances general human capital ... but diminishes domain-specific expertise' is not supported by the full set of results, because the general-human-capital interactions are inconsistent across studies and outcome measures, and the treatment contrast conflates AI access with prompt-crafting instruction. The contribution would be strengthened by a more cautious framing that emphasizes the replicating specific-human-capital moderation and explicitly addresses the identification limitations.","major_comments":[{"comment":"The AI-assist condition is a bundle of (a) access to GPT-4 and (b) additional prompt-crafting instruction, while the control condition receives only basic task instructions. The abstract and Hypotheses 2 and 3 attribute the observed moderation to generative AI itself, but the contrast identifies the effect of an 'AI plus prompt training' package. The proposed theoretical mechanisms—lack of agency and expansive knowledge span—do not reference prompt-crafting instruction, so the empirical test does not map onto the theory. Please either add a control arm with prompt-crafting instruction without AI, or re-label the treatment as AI-assisted work with prompt guidance and temper all causal language about AI per se.","section":"Experiment 1, Samples and Procedures; Experiment 2, Sample and Procedures"},{"comment":"Study 2 experienced 43.04% attrition at the lyric-creation stage, and the final analyses use 329 observations from 299 participants, with no attrition analysis by condition or human capital and no apparent clustering of standard errors by participant. Non-random attrition can break the initial randomization for the moderators, and treating repeated submissions as independent can understate standard errors and inflate significance. Please report attrition balance tests and re-estimate Tables 4-8 with participant-level clustering (or a mixed model).","section":"Experiment 2, Sample and Procedures; Tables 4-8"},{"comment":"The lyrics-only ratings in Experiment 2 have low inter-rater reliability (ICC2 ranges .26-.59), yet the education moderation supporting Hypothesis 2a is significant only on these lyrics-only ratings and not on the more reliable song ratings. This pattern does not constitute consistent support for a general-human-capital moderation effect. The manuscript should report both sets of ratings as equally important, discuss the reliability difference explicitly, and avoid concluding that education moderates AI's effect on creativity when the effect fails to appear in the song ratings.","section":"Experiment 2, Measures; Tables 6-8"},{"comment":"The headline claim overstates the evidence. In Experiment 1, AI use has significant main effects on all outcomes, but in Experiment 2 the main effects are not significant (p = .061 and .075 for song novelty and usefulness). The IQ moderation is significant in Experiment 1 but not in Experiment 2. The education moderation is significant for novelty only in Experiment 1, and in Experiment 2 only for lyrics ratings, not song ratings. The specific-human-capital moderation is the only pattern that replicates consistently. Please reframe the abstract and general discussion around this replicating pattern and describe the general-human-capital results as task-dependent and partially supported.","section":"Abstract; Tables 2-8"},{"comment":"All regressions control for the AI identification ratio, a post-treatment variable measured after the creative output exists. Conditioning on a post-treatment outcome can induce collider or overcontrol bias in the estimated treatment effect and in the interactions, especially because raters' AI perception is likely correlated with output quality and style, which are themselves affected by the treatment. Please re-estimate all main and moderation models without this control, and justify the variable's role (e.g., as a robustness check or a placebo outcome) rather than treating it as a standard covariate.","section":"Experiment 1, Results; Experiment 2, Results (Tables 2, 4-8)"}],"minor_comments":[{"comment":"The phrase 'conventional wisedom' should be 'conventional wisdom'.","section":"Introduction"},{"comment":"The term 'random controlled experiments' should be 'randomized controlled experiments'.","section":"Abstract"},{"comment":"The column header 'AI Identification Raio_L' contains a typo; it should be 'AI Identification Ratio_L'.","section":"Table 3"},{"comment":"No randomization balance table is provided for either experiment; please add a table of covariate means by condition with tests to support the claim of successful random assignment.","section":"Experiment 1 and Experiment 2"},{"comment":"The prompt-length and interaction-round comparisons report t(195) with Nlow=131 and Nhigh=66, but the final sample is N=299; clarify which subsample (apparently only the AI-assist condition) these analyses use and why the degrees of freedom are 195.","section":"Experiment 2, Supplementary Analysis"},{"comment":"Please define 'AI Use' explicitly as the binary assignment indicator in the text and table notes, and state how the variable is coded (e.g., 1 = AI-assist, 0 = control).","section":"Measures and Tables"}],"recommendation":"major_revision","confidential_remarks":"The treatment-bundle issue and the post-treatment control are the two changes I would require before considering the paper publishable; the inconsistent general-human-capital results also need to be reflected in the framing. The authors' own Limitations section is candid about task differences and measurement issues, but it does not mention the prompt-crafting confound or the bad control. In addition, the manuscript cites Li et al. (2024), a working paper whose author list includes the corresponding author; given the topical overlap, the editor should verify that the present submission's claims are distinct from that working paper and that prior dissemination is disclosed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth reading for one empirical pattern: across two randomized experiments, domain-specific expertise consistently moderates the AI-creativity relationship negatively, while education/IQ moderate positively in the first study but not reliably in the second. That dual moderation, if it holds up in cleaner designs, is a genuinely interesting result for the human-AI collaboration literature. The authors also did some things right: they ran two real tasks (flash fiction, song lyrics), used public raters, posted data and code on OSF, reported supplementary analyses, and acknowledged several inconsistencies between studies. The specific-human-capital replication across both studies is the strongest part of the empirical contribution.\n\nThe soft spots are serious, though. The biggest one is the treatment confound: in both experiments, the AI-assist arm received prompt-crafting instruction or additional guidance on using generative AI, while the control arm got only basic task instructions. That makes the contrast AI-plus-training versus no-AI, not AI versus no-AI. The theoretical mechanisms (AI's lack of agency, expansive knowledge span) don't mention prompt training, so the empirical test doesn't cleanly map onto the theory. The prompt training itself could plausibly interact with education, IQ, and domain expertise, generating exactly the observed moderation patterns. The paper never flags this as a limitation, and it should.\n\nThe abstract also overstates the findings. It says AI 'enhances general human capital' and 'diminishes the value of domain-specific expertise,' but in Experiment 2 the IQ moderation was not significant on any outcome, and the education moderation was significant only for lyric-only ratings, not for the complete-song ratings. The specific-human-capital moderation is more consistent, but the general-human-capital side is partial at best. I'd also note the high attrition in Experiment 2 (43% dropped out before submission), low rater reliability on the lyric-only ratings (ICCs in the .26-.59 range), and the use of mean splits for some supplementary analyses, which can inflate or distort interaction estimates. None of these are fatal alone, but together they mean the headline claim needs substantial qualification.\n\nWho is this for? Organizational behavior and management researchers studying AI at work, and empirical creativity researchers. It deserves a serious referee—the question is important and the raw design is ambitious—but the referee should push for a cleaner separation of AI access from prompt-training, or at minimum a much more cautious framing of what the design can identify. I'd like to see the authors reanalyze with the confound acknowledged, report all outcomes without selective emphasis, and revise the abstract to match the results. Send it to peer review, but expect major revision.","headline":"Potentially important empirical pattern about AI and human capital, but the headline causal claim is confounded by prompt-crafting instruction bundled into the AI treatment, and the abstract overstates what the data actually show.","tokens_in":27848,"tokens_out":1485,"would_cite":false,"duration_ms":17042,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Generative AI raises creative output most for people with strong general cognitive skills and least for domain experts, two randomized experiments in fiction and songwriting find.","keywords":["generative AI","human capital theory","creativity","augmentation-automation paradox","randomized controlled experiment","cognitive ability","domain-specific expertise","creative performance"],"falsifier":"A three-arm experiment, with AI plus prompt training, AI without prompt training, and a no-AI control with equally long writing instruction, would settle the central claim: if the education and expertise interactions persist when prompt training is identical, the mechanism is AI, and if they vanish, the mechanism is instruction.","tokens_in":26960,"feed_emoji":"✍️","tokens_out":9396,"duration_ms":80656,"temperature":0.7,"pith_summary":"The paper tries to establish that generative AI does not help everyone equally in creative work: it amplifies the creative output of people with high general human capital, such as cognitive ability and education, while shrinking the performance advantage of domain-specific expertise. Across two randomized experiments, one in flash fiction and one in song lyric writing, the gain in public ratings of novelty, usefulness, and appeal from using AI was largest for participants with stronger general cognitive resources and smallest, often absent, for those with the most domain expertise. The authors argue that this happens because AI lacks agency, so human judgment and idea integration become central, while its expansive knowledge base makes specialized know-how less scarce. If correct, the finding overturns the simple picture of AI as a universal equalizer and predicts that AI adoption will widen gaps tied to adaptability and education while compressing returns to specialized training.","feed_headline":"AI helps flexible thinkers more than domain experts","feed_subtitle":"Experiments in fiction and songwriting show AI widens the creativity gap by education and shrinks the expert edge.","key_machinery":"The machinery is the two-type human capital distinction crossed with the augmentation-automation paradox. General human capital, meaning cognitive ability and education, is treated as the resource that lets a user guide, evaluate, and integrate AI output, so AI augments it; specific human capital, meaning domain expertise such as writing skill or lyric publication history, is treated as knowledge AI can reproduce, so AI automates it away. The empirical engine is a pair of randomized experiments in which creativity is scored by external raters using the consensual assessment technique, with interactions between AI use and measured human capital estimated in OLS regressions.","core_discovery":"The paper's central claim is that generative AI acts as an augmenter for general human capital and an automator for specific human capital. In Experiment 1 (flash fiction, N=162), AI use raised novelty, usefulness, and overall impression on average, and the gains grew with education and IQ; higher self-rated writing skill weakened the benefit, significantly for usefulness and overall impression. In Experiment 2 (lyric writing, N=299 with 329 rated works), AI use alone did not significantly lift lyric ratings, but education still positively moderated the lyric-only ratings, while prior lyric publication consistently and negatively moderated AI's effect on both lyric-only and full-song ratings. The authors interpret the pattern as a shift in the locus of creative advantage: what matters most in AI-assisted creative work is broad cognitive adaptability and the ability to integrate and evaluate ideas, not accumulated domain expertise.","pith_inferences":["The authors do not test this directly, but the mechanism implies that the negative moderation by expertise should be strongest in formulaic creative genres, where AI's knowledge span most closely substitutes for accumulated experience; a genre-by-genre replication could test that.","Because both experiments gave prompt-crafting guidance only to the AI arm, one reading is that the interactions with education and expertise could partly reflect training rather than AI itself; a clean test would vary prompt training independently of AI access.","The reported drop in psychological ownership suggests a possible dynamic cost: over repeated tasks, lower ownership could reduce intrinsic motivation to iterate, which would mute the augmentation effect observed in a single session.","Read as a labor-market prediction, the framework implies that credentials and portfolios signaling domain craft should lose relative value against measures of fluid intelligence and learning agility; longitudinal wage or hiring data after AI adoption could check this."],"forward_implications":["For organizations, the value of hiring and training for general cognitive ability should rise as generative AI spreads, while roles defined mainly by narrow domain mastery face downward pressure on their creative premium.","For individuals, access to the same AI tool will not equalize creative performance; gains depend on the user's cognitive adaptability, so simple tool access is not a sufficient equity intervention.","In specialized creative fields, novices stand to gain the most from AI assistance, whereas experienced experts may see little or no improvement unless they change how they interact with the tool.","AI's benefit is task-contingent: in the songwriting study the average effect of AI was not significant, so in tasks where emotional expression and idea generation dominate over writing fluency, AI may not raise average creativity at all."],"supporting_citations":[{"why":"Supplies the human capital theory and the general-versus-specific distinction the hypotheses are built on.","marker":"(Becker, 1962)"},{"why":"Provides the automation-augmentation paradox that predicts AI will complement general capital and substitute for specific capital.","marker":"(Raisch & Krakowski, 2021)"},{"why":"Supplies the consensual assessment technique used to obtain external creativity ratings.","marker":"(Amabile et al., 1996)"},{"why":"Supplies the approach of using audience-style raters to evaluate creative products.","marker":"(Berg, 2019)"},{"why":"Baseline experimental evidence that generative AI raises productivity and compresses performance variance, which the paper challenges and extends.","marker":"(Noy & Zhang, 2023)"},{"why":"Recent evidence that generative AI boosts individual creativity while homogenizing collective output, a reference point for the main effect and similarity analyses.","marker":"(Doshi & Hauser, 2024)"},{"why":"Field evidence on AI and knowledge-worker productivity that motivates the emphasis on general human capital and adaptability.","marker":"(Dell'Acqua et al., 2023)"},{"why":"Supplies the specialization-at-the-knowledge-frontier logic that the automation side of the argument inverts.","marker":"(Teodoridis et al., 2019)"}],"fun_headline_variants":["AI rewards flexible thinkers, punishes domain experts","Creative AI: education boosts gains, expertise cuts them","Generative AI widens skill gaps in fiction and song","AI in creative work: adaptability trumps specialization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that the AI-assisted group's extra prompt-crafting instructions are not doing the work, because only that group received the tutorial and the effects attributed to AI could therefore come from the training instead.","fun_headline_variants_meta":{"raw":{"variants":["AI rewards flexible thinkers, punishes domain experts","Creative AI: education boosts gains, expertise cuts them","Generative AI widens skill gaps in fiction and song","AI in creative work: adaptability trumps specialization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000511,"raw_usage":{"total_tokens":2466,"prompt_tokens":909,"completion_tokens":1557,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":1495}},"tokens_in":525,"tokens_out":1557,"duration_ms":10260,"temperature":1.0,"reasoning_tokens":1495,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:54:15.408485+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A three-arm experiment, with AI plus prompt training, AI without prompt training, and a no-AI control with equally long writing instruction, would settle the central claim: if the education and expertise interactions persist when prompt training is identical, the mechanism is AI, and if they vanish, the mechanism is instruction.","supporting_citations":[],"review_version":1}