{"id":"96f3eabf-e32a-4a99-b284-0877a7e0ce42","arxiv_id":"2506.20803","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A randomized execution study with 43 experts shows that LLM-generated research ideas lose more of their appeal than human ideas when actually implemented, reversing part of their ideation-stage advantage.","lead":"Researchers had 43 experts spend over 100 hours each turning research ideas into real papers. AI-written ideas scored higher than human ideas before execution, then dropped much more than human ideas after the work was done, showing that ideas need to be tested, not just judged on paper.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 5's gap may partly reflect regression to the mean: AI ideas start higher, ideation-execution correlations are near zero, and the difference-score analysis never conditions on baseline ideation score.","rationale":"The reader's identified weakest assumption is cross-study scale comparability. I agree that is a threat, but the sharper, more decisive threat is that the Table 5 estimator is biased by baseline imbalance: AI ideas score higher at ideation, execution scores are only weakly correlated with ideation scores, and a difference score with baseline imbalance creates regression to the mean. A crossover calibration of the two scales would not fix this artifact, so I treat the statistical issue as the load-bearing one. The paper otherwise has real strengths: randomized assignment, pre-registration, blind review, expert executors, and a sensitivity analysis in Appendix F. Those strengths do not address the baseline confound, because none of the robustness checks conditions on ideation score. Keeping the conditional verdict is appropriate until the released data are re-analyzed with the proposed baseline-adjusted test; the concern is concrete and falsifiable rather than a rejection of the study's value.","tokens_in":49016,"tokens_out":6355,"duration_ms":79823,"concrete_test":"Re-analyze the released per-idea data. For each of the four metrics, fit execution score ~ condition + ideation score (plus topic and reviewer random effects if identifiable) and report the condition coefficient. Additionally, simulate the null: keep observed ideation scores and permutation-shuffle execution scores across conditions, then compute the distribution of the Table 5 gap difference. If the observed Delta falls within the permutation null or the ANCOVA condition coefficient is not significant after baseline adjustment, the ideation-execution gap is not established; if the coefficient remains significant and close to Delta, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is statistical rather than semantic: Table 5's change-score comparison is confounded by regression to the mean. Table 4 shows AI ideas were significantly higher at ideation on all four metrics (e.g., effectiveness 6.003 vs 4.833; overall 5.382 vs 4.596). Appendix D (Table 9) then reports near-zero ideation-execution correlations (AI novelty r=-0.019, excitement r=-0.321; human |r|<=0.21). When the outcome is a difference D=E-I and I contains measurement error, D is mechanically negatively correlated with I; a group that scored higher at ideation will show larger negative mean drops even if its true execution outcomes are identical. The claim in Section 4.2 that the gap 'controls for heterogeneity in idea quality' is the opposite of what a difference score does under baseline imbalance. Using Table 4 means, a null model with execution scores independent of ideation would already predict gap differences of roughly 0.79-1.25 points (the ideation gaps), i.e., a large fraction of the observed Delta=1.35-1.84 in Table 5, before any condition effect. The Table 5 p-values therefore do not establish that AI ideas lose more value after execution unless the analysis conditions on ideation score (e.g., ANCOVA) or shows the effect within ideation-score strata.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a pre-registered execution study in which 43 expert researchers were randomly assigned to execute 19 human-written and 24 LLM-generated NLP research ideas taken from the authors' prior ideation study (Si et al., 2025). Each participant spent roughly 100 hours implementing the assigned idea and writing a 4-page paper, and 58 blind reviewers scored the resulting papers and codebases (181 reviews). The authors compare execution scores with the ideation scores from the prior study and find that AI ideas drop more between ideation and execution than human ideas on novelty, excitement, effectiveness, and overall (e.g., -1.049 vs -0.010 for novelty, Table 5). They also analyze executor changes, reviewer rationales, and robustness to removing six AI ideas whose human-evaluation components were dropped.","tokens_in":49159,"tokens_out":5132,"duration_ms":61115,"significance":"If the result holds, it is an important and timely caveat to claims that LLM-generated research ideas are superior to human ideas: idea quality judged without execution would overstate LLM idea value. The study's strengths include pre-registration, randomization and blinding, release of data, substantial per-project execution effort, and a useful robustness check. Its main quantitative claim, however, rests on a change-score comparison that is vulnerable to regression to the mean and to uncalibrated evaluation scales; the current analyses do not rule out these artifacts. The study is nonetheless a valuable and original contribution to the methodology of evaluating AI-generated research ideas, and the authors' qualitative analyses (Section 5) are informative.","major_comments":[{"comment":"The change-score analysis is confounded by regression to the mean. Because the outcome D = E - I is mechanically negatively correlated with I whenever I contains measurement error, and Table 9 reports near-zero ideation-execution correlations (AI novelty r = -0.019, excitement r = -0.321; human |r| <= 0.21), a group with higher mean ideation scores will tend to show larger negative mean drops even if its true execution outcomes are identical. Table 4 shows AI ideas start significantly higher at ideation (e.g., effectiveness 6.003 vs 4.833); a null model with execution scores independent of ideation would already predict gap differences of roughly 0.79-1.25 points, which is a large fraction of the observed Delta = 1.35-1.84 in Table 5. The statement in Section 4.2 that the gap 'controls for heterogeneity in idea quality' is therefore inaccurate for this design: difference scores remove stable idea-level effects only under strong assumptions that are not met here. Please re-analyze the data with ANCOVA on execution score adjusting for ideation score, or with within-ideation-strata matching, and report what the null model predicts for Table 5 before attributing the gap to condition.","section":"Section 4.2, Tables 4 and 5, Appendix D"},{"comment":"The ideation and execution evaluations are not on a common, calibrated scale. Ideation reviewers scored short idea proposals, whereas execution reviewers scored 4-page papers and codebases using a different review form that adds soundness, codebase quality, and faithfulness, and the overall-score rubric is anchored to a 'short paper track at *ACL' (Appendix A). The gap D = execution score minus ideation score therefore conflates genuine idea-quality changes with any shift in reviewer standards, rubric anchors, or task demands between the two studies. Because the AI and human conditions have different ideation score distributions, even a common additive or multiplicative shift in scoring standards would produce differential observed gaps. The authors state that the review guidelines 'closely match' the ideation study, but this is not a calibration. Please provide evidence of scale comparability, for example through a crossover review of a common set of ideas or a sensitivity analysis allowing for study-specific shifts in score distributions, or temper the causal interpretation of the gap accordingly.","section":"Section 4.2 and Appendix A"},{"comment":"The randomization of ideas to executors was within preferred topics, but the realized samples show large imbalances in potentially important covariates: execution participants in the Human condition spent more hours on average (112.6 vs 93.7) and reported lower topic familiarity (2.9 vs 3.4) than those in the AI condition. Since execution quality is partly a function of executor effort and familiarity, these imbalances could contribute to the observed execution-score differences and hence to the gap. Please report balance tests for these covariates and show that the Table 5 conclusions survive adjustment for time spent and familiarity, or at least discuss the sensitivity of the results to these imbalances.","section":"Section 2.1 and Table 2"}],"minor_comments":[{"comment":"There are several typos and formatting errors: 'heterogenity' in Section 4.2, 'reponsibility' in Section 7.1, 'Learning positive' in the Appendix A excitement scale, and spacing issues in Table 4 (e.g., '4 .404').","section":"Throughout"},{"comment":"The axis label '(Study2 Study1)' in Figure 2 is malformed; it should read something like '(Study 2 - Study 1)'.","section":"Figure 2"},{"comment":"The manual categorization of reviewer rationales into the ten categories in Figure 3 would benefit from reporting inter-annotator agreement, since the categories are subjective and the figure underlies the qualitative claim that execution reviews consider more factors.","section":"Section 5.2, Figure 3"},{"comment":"The review-level t-tests in Table 3 treat each review as independent, which ignores clustering by idea and reviewer; the authors do report idea-level results, but the text should be explicit that the review-level p-values are exploratory.","section":"Table 3"},{"comment":"Wen et al. (2025) is cited in the Future Work section on proxy reward models but is not discussed in Related Work; a brief mention there would help contextualize the execution-outcome-prediction thread.","section":"Related Work"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The primary issue is statistical. The change-score analysis in Table 5 must be re-run conditioning on ideation score (e.g., ANCOVA) before the paper's central claim can be accepted; if the effect disappears, the title and abstract claims should be weakened accordingly. The paper transparently reuses the authors' own prior ideation study, but the degree of dependence on that study's measurements and on uncalibrated cross-study scores is greater than the text acknowledges. The qualitative analyses and the study design are genuinely valuable, so I see this as a fixable major revision rather than a rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Main punchline: this is the first large-scale execution comparison of LLM vs human research ideas, and the study design is genuinely impressive. But I don't trust Table 5's central claim. The AI ideas start significantly higher on ideation, and the ideation-execution correlations are near zero (Table 9). When the post-score is essentially uncorrelated with the pre-score, a difference score mechanically favors the lower-baseline group: both groups regress toward the same execution mean, so the group that started higher shows a larger drop. Using Table 4's means, a null model with execution independent of ideation predicts gap differences of 0.79–1.25 points, versus the observed 1.35–1.84. That's most of the effect. The paper says the gap 'controls for heterogeneity in idea quality,' but it does the opposite; it needs an ANCOVA conditioning on ideation score, or within-stratum analysis, before the gap can be attributed to idea origin.\n\nWhat the paper does well: the execution RCT is a real contribution. 43 researchers, 100+ hours each, blinded review, pre-registered, data released. The qualitative analysis in Section 5 is concrete and useful—AI ideas more often proposed human evaluations that had to be swapped for LLM-as-a-judge, and execution reviewers catch missing baselines and resource issues that ideation reviewers missed. The direct execution comparison shows a consistent (though underpowered) trend favoring human ideas on effectiveness and excitement.\n\nSoft spots beyond the main one: the two evaluations used different reviewer pools, different artifacts (idea proposals vs papers plus code), and arguably different constructs (expected vs realized effectiveness), with no anchor-item calibration. The abstract's 'flip' language overstates a non-significant result. Attrition by condition is not reported.\n\nWho this is for: anyone building AI-scientist systems or idea benchmarks. The paper deserves a serious referee because the study is a first and the released data enable re-analysis. My recommendation: engage with it, but require a re-analysis conditioning on baseline ideation score before the headline claim is accepted.","headline":"A first-of-its-kind execution study whose central gap result may be a regression-to-the-mean artifact rather than a real ideation-execution gap.","tokens_in":49817,"tokens_out":3908,"would_cite":false,"duration_ms":43859,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM-generated research ideas suffer far larger review-score drops than human expert ideas once they are actually executed, erasing the ideation-stage advantage AI ideas had shown.","keywords":["LLM-generated research ideas","ideation-execution gap","research idea evaluation","execution study","human evaluation","randomized controlled trial","AI scientists","research idea quality"],"falsifier":"Score a subsample of the same executed projects under the ideation protocol — reviewers see only the original idea text, without the paper or codebase — and check whether AI ideas' drop relative to human ideas persists when the evaluation protocol is held fixed; if the differential drop disappears, the gap is an artifact of the two evaluation regimes rather than of idea quality. The converse check is to have a second team of executors independently re-execute the same ideas: if the larger AI drop reproduces, the gap is robust.","tokens_in":48692,"feed_emoji":"📉","tokens_out":5277,"duration_ms":56137,"temperature":0.7,"pith_summary":"This paper asks whether AI-generated research ideas that look good on paper still look good after real researchers spend months executing them. The authors took research ideas written by expert NLP researchers and by an LLM, randomly assigned 43 qualified researchers to execute them blind, and had 58 expert reviewers score the resulting projects. The central finding is an ideation–execution gap: on a 10-point scale, the LLM ideas' scores fell by roughly 1.0 to 2.0 points after execution on novelty, excitement, effectiveness, and overall quality, while human ideas barely moved, shrinking and on some metrics reversing AI's ideation-stage advantage. The paper argues that judging ideas without executing them systematically overstates the value of LLM-generated ideas, because only implementation exposes weaknesses such as infeasible plans, optimistic human-evaluation proposals, and unverified assumptions.","feed_headline":"AI research ideas lose up to 2 points once executed","feed_subtitle":"In 43 blind executions, LLM ideas fall 1.0–2.0 points, erasing AI's edge over human ideas.","key_machinery":"The load-bearing object is the ideation–execution gap: for each idea, the difference between its post-execution review score and its pre-execution ideation score, computed on the four metrics shared by both evaluations (novelty, excitement, effectiveness, overall). Measuring the within-idea change from before to after execution removes much of the heterogeneity in idea quality that makes direct human-versus-AI comparisons statistically noisy, and this paired difference is what yields significant results at a sample size of 43. The comparison rests on a blinded, pre-registered randomized design: 43 expert executors each spent roughly 100 hours implementing a randomly assigned idea and wrote a 4-page paper, and 58 blinded expert reviewers produced 181 reviews, with the human and AI conditions indistinguishable on the control metrics of faithfulness and codebase quality.","core_discovery":"The paper's central claim is that LLM-generated research ideas incur a significantly larger drop in expert review scores once they are executed than human expert ideas do. Using 43 executed projects (19 human, 24 AI) drawn from a prior blinded ideation study, the authors compute, for each idea, the difference between its execution review score and its ideation score on the four metrics both evaluations share. AI ideas drop by 1.049 to 1.976 points across novelty, excitement, effectiveness, and overall score; human ideas change by only −0.010 to −0.628, and the difference between the two conditions is significant on every metric (FDR-corrected p < 0.05). Because both conditions were executed under identical instructions, reviewed blind, and matched on faithfulness to the original idea and codebase quality, the authors attribute the differential drop to the origin of the idea. They further show that the drop is not explained by the executors' minor experiment-detail changes, and that execution reviewers weigh empirical performance, experimental rigor, and feasibility, factors that are almost invisible at the ideation stage.","pith_inferences":["The result suggests a proxy-reward failure that likely generalizes beyond this study: LLMs may be optimizing ideas for surface signals that impress reviewers, such as novelty and excitement, rather than for executability, so any automated pipeline that scores ideas without executing them will drift toward ideas that look good and fail under implementation.","A testable extension: if the gap is driven by feasibility mismatches, stratifying AI ideas by the resources they propose (for example, planned human evaluations or heavy computation) should predict gap size; AI ideas that proposed human studies were indeed more likely to be altered by executors, and removing those six ideas did not change the result.","A crossover replication, in which the same ideas are executed by a second independent set of researchers, would separate idea quality from executor variance and reveal how much of the gap is attributable to the idea itself rather than its implementer."],"forward_implications":["Idea-stage reviews, by themselves, overstate the value of LLM-generated research ideas relative to human expert ideas; execution outcomes are needed to see their true quality.","The gap closes, and on excitement and effectiveness even reverses, the ranking between AI and human ideas observed at ideation, although the reversed ranking is not statistically significant at this sample size.","Ideation scores are weak predictors of execution outcomes, with correlations weak in most cases and moderately negative for AI ideas on excitement, so ideation evaluation should not be used as a proxy for research impact.","Evaluations of AI-scientist systems that stop at proposal or paper review without independent execution will likely overestimate these systems' research capability.","Reviewer rationales show that execution evaluation surfaces empirical performance, experiment rigor, and resource feasibility, factors that ideation evaluation cannot see because it implicitly assumes the proposed method will work."],"supporting_citations":[{"why":"Supplies the human and AI idea pool and the pre-execution ideation review scores; the gap measure subtracts these scores from the new execution scores.","marker":"(Si et al., 2025)"},{"why":"Shows that expert evaluation of ideas before execution is unreliable, motivating the execution study as the ground truth for idea quality.","marker":"(Simsek et al., 2024)"},{"why":"An AI-scientist system whose wet-lab validation is limited to a handful of ideas; the contrast that this study scales up.","marker":"(Ghareeb et al., 2025)"},{"why":"A similar small-scale execution validation of AI hypotheses, illustrating why prior work could not draw statistical conclusions.","marker":"(Gottweis et al., 2025)"},{"why":"An end-to-end AI Scientist whose papers are evaluated without independent execution; the target of the paper's critique that idea review overestimates quality.","marker":"(Lu et al., 2024)"}],"fun_headline_variants":["LLM ideas lose up to 2 points once executed","Executing AI ideas flips rankings toward human experts","Human ideas outrank AI after execution, reversing novelty gap","AI research ideas drop 1-2 points when put to test","Ideation-execution gap: AI ideas lose ground to human ones"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The most fragile premise is that the ideation-stage scores (from the earlier study) and the execution-stage scores (from this study) rate the same quality on the same numeric scale, so that a larger pre-post drop for AI ideas reflects worse ideas rather than a shift in reviewer standards, rubric anchors, or evaluation context between the two studies.","fun_headline_variants_meta":{"raw":{"variants":["LLM ideas lose up to 2 points once executed","Executing AI ideas flips rankings toward human experts","Human ideas outrank AI after execution, reversing novelty gap","AI research ideas drop 1-2 points when put to test","Ideation-execution gap: AI ideas lose ground to human ones"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000373,"raw_usage":{"total_tokens":2043,"prompt_tokens":1042,"completion_tokens":1001,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":927}},"tokens_in":658,"tokens_out":1001,"duration_ms":11276,"temperature":1.0,"reasoning_tokens":927,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:41:54.084600+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Score a subsample of the same executed projects under the ideation protocol — reviewers see only the original idea text, without the paper or codebase — and check whether AI ideas' drop relative to human ideas persists when the evaluation protocol is held fixed; if the differential drop disappears, the gap is an artifact of the two evaluation regimes rather than of idea quality. The converse check is to have a second team of executors independently re-execute the same ideas: if the larger AI drop reproduces, the gap is robust.","supporting_citations":[{"cited_title":"Do grant proposal texts matter for funding decisions? a field experiment","cited_arxiv_id":null,"evidence_quote":"Shows that expert evaluation of ideas before execution is unreliable, motivating the execution study as the ground truth for idea quality."}],"review_version":1}