{"id":"7d47b283-33d1-4ebe-bf0b-a7421e4db436","arxiv_id":"2607.25881","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"AI-generated one-page research proposals are scored about the same as human-written ones by human reviewers, but AI reviewers favor AI-written proposals by roughly one point and detect authorship perfectly.","lead":"Scientists asked humans and three AI chatbots to each write one-page research proposals for eight physics projects, then had human and AI reviewers score them blindly. Human reviewers found the AI plans on par with human ones, while AI reviewers preferred AI-written plans and always guessed the writer correctly.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human reviewer independence is the load-bearing weak point: all four human reviewers are co-authors, so parity and origin-detection results may reflect recognition or conflict of interest rather than an unbiased panel.","rationale":"I agree with the reader's weakest assumption. The co-author reviewer panel is the most load-bearing issue because it directly undermines the abstract's primary result of human/AI parity and the human origin-detection rates. The paper is transparent but does not address this conflict of interest. No other concern (AI-polishing confound, small sample, classification cues) is as central to the main inference. The proposed concrete test—an independent replication of the human review—would settle whether the parity is robust. Therefore the reader's CONDITIONAL verdict should remain unchanged pending that test.","tokens_in":16480,"tokens_out":9720,"duration_ms":92228,"concrete_test":"Recruit 4-6 independent expert reviewers with no affiliation to the paper or its authors. Blind them to the study design and authorship; have them score all 32 proposals using the same rubric and origin-judgment task. Compute mean human-vs-AI score differences and detection accuracies, plus inter-rater reliability (e.g., ICC or Fleiss' kappa) for both original and independent panels. If independent reviewers find similar parity (within 0.2 points) and detection rates within ±10%, the concern is resolved; if not, the central claim must be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of human/AI parity rests on the four human reviewers, all of whom are co-authors (Sec. II B). They may recognize the human planners' writing styles or project identities, biasing quality ratings and origin judgments. No inter-rater reliability is reported, and the authors note reviewers 'often disagreed with one another' (Sec. III B). Origin-judgment accuracy across reviewers spans 59%-88%, so the panel is heterogeneous. With only 8 projects per condition, one or two biased reviewers could determine the parity result. This threatens the abstract's first clause: 'Human reviewers rated human- and AI-written proposals similarly overall.' If the human panel is biased, the conclusion that LLMs produce comparable project plans is unsupported. The disclosure that the reviewers are co-authors does not mitigate the threat; it identifies a limitation that requires external verification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a controlled, blinded study of AI-assisted scientific project planning and proposal evaluation. Eight expert-written project seeds in physics, astrophysics, and cosmology were each expanded into one human-written and three LLM-written one-page proposals, yielding 32 proposals. Four human reviewers (all co-authors) and two frontier LLM reviewers evaluated all proposals on a four-aspect rubric and also made a binary human-versus-AI authorship judgment. The reported results are: human reviewers rated human- and AI-written proposals similarly overall; both AI reviewers scored AI-written proposals roughly one point higher than human-written ones; AI reviewers classified all 32 proposals correctly, while human reviewers averaged 72% and 79% correct on human- and AI-written proposals, respectively; and the AI-human score gap is strongly anti-correlated with the quality of the human-written proposal. The authors conclude that LLMs can currently produce project plans comparable to human ones in the eyes of human reviewers, but that AI reviewers exhibit a systematic pro-AI bias that warrants caution in deploying LLMs in proposal preparation and review. The manuscript is transparent about the small sample, the rapid model turnover, and the fact that the human reviewers are authors, and it includes appendices with project seeds, prompts, and sample proposals.","tokens_in":16677,"tokens_out":6589,"duration_ms":65674,"significance":"If the results hold, this is a useful and timely controlled contribution to the growing literature on LLMs in scientific workflows. The study design has real strengths: a fixed template and common starting point for human and AI proposals, uniform anonymized formatting, explicit disclosure of evaluator identity, and rich qualitative material on the cues used for authorship judgments. The finding of a systematic preference by AI reviewers for AI-written proposals, if robust, has direct policy relevance for grant review, and the authors appropriately cite agency guidelines that restrict AI use in peer review. However, the paper's central quantitative claims rest on a small sample with no inferential statistics, and the human-reviewer panel consists entirely of co-authors, which is a serious independence risk. The human-written condition is also contaminated by a ChatGPT grammar-correction pass, and one of the paper's key project-level correlations is partly mechanical. These issues are substantive but addressable through re-analysis and more careful framing, so the paper is best treated as a strong pilot study requiring major revision rather than a definitive measurement.","major_comments":[{"comment":"The human panel that grounds the parity claim consists of four co-authors. The manuscript discloses this but does not control for or test the obvious recognition risk: reviewers may recognize the human planners' prose or project identities, which could inflate the human-written proposals' quality ratings and contaminate the 72%/79% origin-detection rates. No inter-rater reliability statistic is reported; the paper even notes that the four reviewers 'often disagreed' (Sec. III B). With eight projects per condition, one or two non-independent judges can determine the aggregate parity result. This is load-bearing because the abstract's first claim rests entirely on this panel. Please report per-reviewer scores and agreement (e.g., ICC or Fleiss' kappa), and either add independent external reviewers or re-frame the human-panel component as a pilot with the independence threat as a central li","section":"Sec. II B / III B"},{"comment":"The 'human-written' proposals were post-processed by ChatGPT 4o. Thus the human condition is not purely human-authored; AI stylistic normalization may introduce the very template-like, polished features that the reviewers (human and AI) report using to flag AI text (Sec. III A). This can inflate both the AI reviewers' 100% classification and the human reviewers' perception of parity. The manuscript should quantify the edit distance or amount of rewriting, justify that 'minimal changes' indeed preserved content and style, and preferably include a no-AI-touch control arm, or at least treat the condition as 'human-written then AI-normalized' in all claims.","section":"Sec. II C"},{"comment":"The central quantitative claims—that human reviewers rate human and AI proposals similarly and that AI reviewers give AI proposals about one point more—are supported only by group means and standard deviations across n=8 projects. No confidence intervals, p-values, or effect sizes are given, and the small sample means the mean difference can be dominated by one project (e.g., GW, which is called out in Sec. III C). Please provide per-project data, bootstrapped or permutation confidence intervals for the gaps, and ideally a mixed-effects model with reviewer and project as random effects. Without this, the 'similarly overall' and 'one point higher' phrasing overstates the precision of the measurements.","section":"Sec. III B, Tables IV and V"},{"comment":"The reported r=-0.95 between the AI-human score gap and the human-written proposal score is largely mechanical, because the gap subtracts the human score from itself. Even if the AI-written score were statistically independent of the human-written score, a strong negative correlation would appear; the reported value is therefore not evidence that the AI advantage is concentrated in projects with weak human plans. The correct diagnostic is a scatter plot of mean AI score versus human score with a fitted slope, or a formal model. Please re-analyze before drawing the conclusion that the parity and bias are 'driven by the same handful of projects.'","section":"Sec. III C, Fig. 4"}],"minor_comments":[{"comment":"The abstract says 'two AI reviewers' but Table III also lists Codex 5.5 Pro and Claude Sonnet 4.6; clarify which are primary and which are supplementary, and make the table caption consistent with the main text.","section":"Abstract / Sec. III A"},{"comment":"The AI evaluation prompt is not included; Appendix B gives proposal-generation prompts only. Since the AI reviewers' behavior depends on the exact instruction, please include the reviewer prompt in the appendix or a supplement.","section":"Sec. II D / Appendix B"},{"comment":"Minor typo in the rubric header: 'W eak' should be 'Weak'.","section":"Table II"},{"comment":"No data/code availability statement appears. Since the study is empirical and based on LLM outputs, depositing de-identified proposals, reviewer scores, and prompts would substantially strengthen reproducibility.","section":"Reproducibility"},{"comment":"Reference [20] is listed as arXiv:2607.xxxx with a placeholder; this should be updated to the actual companion paper ID. Also, the paper's concluding caveats do not mention the ChatGPT grammar-correction pass on human proposals or the co-author reviewer bias; these belong in the limitations paragraph.","section":"Sec. IV / References"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent and methodologically inventive, but the human-reviewer independence threat is serious and is compounded by the ChatGPT post-processing of human proposals. The descriptive patterns are suggestive, especially the consistent pro-AI bias across four AI evaluators, but the paper currently overstates the precision of the human/AI parity result. I recommend major revision: require the per-reviewer and per-project analyses, inferential statistics, a mechanical-correlation check for Fig. 4, and a substantially hedged abstract/conclusion. This is not a reject because the core question is timely and the reported effect is consistent with prior work; with re-analysis and caveats the paper could make a valid contribution as a pilot study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful part of this paper is the controlled comparison: eight expert-conceived projects, one human and three AI one-page proposals per seed, blinded scoring on a four-aspect rubric, plus origin judgments. That's a clean design for a hard task, and the appendices actually give you the prompts, rubric, and sample proposals, so you can reproduce the protocol. The project-level analysis is the most interesting bit: the AI-human score gap tracks the quality of the human plan (r = -0.95), meaning the AI advantage is concentrated where the human proposal was weak. That nuance is more informative than the average parity headline.\n\nThe headline result—human reviewers see no difference, AI reviewers favor AI by about a point—is plausible, but it rests on four human reviewers who are also co-authors. The paper says they are senior postdocs/faculty and mirror a typical panel, but they could recognize the planners or their own projects. The stress-test is right to call this load-bearing: with eight projects per cell, one or two biased reviewers could flip the parity result. The authors don't report inter-rater reliability and even mention reviewers often disagreed. I'd want an independent panel, or at minimum a leave-one-reviewer-out sensitivity check.\n\nOther soft spots are proportionally smaller: no inferential statistics (means and SDs only, no CIs or tests), and the human proposals were lightly polished by ChatGPT to normalize style, which slightly muddies the origin-detection comparison. The sample is small, but the authors are candid about that.\n\nThis deserves a serious referee. The design is transparent, the limitations are addressable, and the AI-reviewer pro-AI bias is consistent with earlier work but this domain-specific evidence is worth having. Send it to review with a request for independent reviewer data or analysis, confidence intervals, and ideally release of the corpus.","headline":"A transparent, well-designed controlled study whose central parity claim leans on a human reviewer panel made up entirely of co-authors; fix that and the paper has real value.","tokens_in":17233,"tokens_out":2883,"would_cite":true,"duration_ms":31135,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI-written research proposals pass human review, but AI reviewers favor AI-authored text by about one point.","keywords":["large language models","scientific proposals","project planning","proposal evaluation","peer review bias","authorship detection","physics research","AI-assisted research"],"falsifier":"Run the identical 32-proposal, blinded review with eight reviewers from outside the author list who have never seen the proposals, and check whether human-rated AI-vs-human means remain within 0.2 points and origin-detection accuracy lands near 72/79%. A second decisive test: give both AI reviewers the same proposals with instructions to ignore style and score only scientific substance; if the roughly one-point pro-AI gap vanishes, the bias is stylistic rather than substantive.","tokens_in":16402,"feed_emoji":"⚖️","tokens_out":4179,"duration_ms":39376,"temperature":0.7,"pith_summary":"The paper asks whether large language models can plan real research projects and whether they can be trusted to review such plans. Eight expert-conceived projects in physics, astrophysics, and cosmology were each turned into four one-page proposals: one by a human expert, three by mid-2025 LLMs. Four human reviewers, blind to origin, rated the human and AI proposals statistically indistinguishable overall, while two frontier AI reviewers gave AI-written proposals roughly one point higher on a five-point scale. Human reviewers identified authorship about three-quarters of the time; the AI reviewers did so perfectly. The paper concludes that LLMs can already produce competitive short research plans, but that deploying AI as reviewer carries a systematic pro-AI bias that would disadvantage human applicants.","feed_headline":"AI proposals pass human review, but AI judges favor AI","feed_subtitle":"Human reviewers rated AI and human plans equally; AI judges scored AI plans ~1 point higher and caught every AI author.","key_machinery":"The controlled corpus of 32 one-page proposals: eight expert-conceived project seeds (title, background, goal) each expanded by one human expert and three LLMs under an identical four-section template and fixed prompt, plus the same four-aspect scoring rubric (clarity and structure, appropriateness of methods, resource and tool planning, feasibility/timeline/risk) and binary origin-judgment task given to all reviewers. The design isolates content from stylistic tells — uniform formatting, grammar normalization of human text — so that observed differences in scores and origin judgments can be attributed to author type rather than formatting.","core_discovery":"The central discovery is a three-part asymmetry. First, capability: when blinded expert reviewers score one-page research plans on a four-aspect rubric, AI-generated plans are rated no worse than plans written by human experts, with mean totals around 3.5 out of 5 for both. Second, detectability: human experts can tell AI from human authorship about 72–79% of the time, while the two most capable AI reviewers classified all 32 proposals correctly, 100%. Third, bias: every AI reviewer tested scored AI-written proposals about one point higher than human-written ones, a pro-AI preference absent from the human panel; human-written proposals received nearly identical scores from every evaluator, s","pith_inferences":["The near-constant vertical offset between human and AI panels suggests a simple calibration: shifting AI-reviewer scores down by about one point for AI-flagged proposals would align their project-level rankings with the human panel's; this is testable on the paper's own data.","Because the pro-AI bias appeared across all four AI models and both vendors, it likely stems from stylistic regularities of LLM output — template-like five-phase structure, exhaustive tool lists, round-number timelines — rather than any single model's idiosyncrasy, so the bias may persist even as models improve unless review instructions explicitly penalize those markers.","The perfect AI detection rate may be an artifact of the constrained one-page template, since LLM proposals in this corpus share a detectable 'too clean' structure; running the same origin-judgment task on full-length, unconstrained proposals would test whether accuracy collapses.","A decisive external check the authors did not run: repeat the review with non-author reviewers who have no stake in the outcome, to rule out recognition effects on the 72/79% human detection rates and the human parity ratings."],"forward_implications":["If short structured proposals are representative, researchers can use LLMs to draft competitive project plans without losing quality in human review.","AI-reviewer pro-AI bias means incorporating LLM evaluators into grant review without calibration would systematically disadvantage human-written proposals.","Perfect AI authorship detection on this corpus suggests current frontier models can identify AI-generated planning text at high accuracy, though the paper cautions against generalizing from 32 proposals.","The tight anti-correlation between AI advantage and human-plan quality implies AI assistance may be most valuable where the human baseline plan is weak, and least valuable where it is already strong.","Per-aspect scores show AI plans are weakest on feasibility, timeline, and risk awareness under both human and AI reviewers, so AI assistance is least reliable for realistic scheduling and contingency planning."],"fun_headline_variants":["AI plans match human experts, but AI judges play favorites","AI proposals score equal with humans, yet AI reviewers boost AI","Human judges fair to AI plans; AI judges favor AI","AI-written plans pass blind human review, but AI reviewers show bias","LLMs write plans that pass human review; AI critics catch every AI"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The valid comparison assumes the four human reviewers — co-authors of this paper — judge the proposals exactly as an independent grant panel would, with no recognition of their colleagues' writing and no stake in the outcome; if they recognized or favored particular proposals, both the human/AI parity and the reported detection rates could be inflated.","fun_headline_variants_meta":{"raw":{"variants":["AI plans match human experts, but AI judges play favorites","AI proposals score equal with humans, yet AI reviewers boost AI","Human judges fair to AI plans; AI judges favor AI","AI-written plans pass blind human review, but AI reviewers show bias","LLMs write plans that pass human review; AI critics catch every AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000709,"raw_usage":{"total_tokens":3042,"prompt_tokens":772,"completion_tokens":2270,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":516,"completion_tokens_details":{"reasoning_tokens":2183}},"tokens_in":516,"tokens_out":2270,"duration_ms":15317,"temperature":1.0,"reasoning_tokens":2183,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T01:13:40.347272+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical 32-proposal, blinded review with eight reviewers from outside the author list who have never seen the proposals, and check whether human-rated AI-vs-human means remain within 0.2 points and origin-detection accuracy lands near 72/79%. A second decisive test: give both AI reviewers the same proposals with instructions to ignore style and score only scientific substance; if the roughly one-point pro-AI gap vanishes, the bias is stylistic rather than substantive.","supporting_citations":[],"review_version":1}