REVIEW 4 major objections 5 minor 41 references
AI-written research proposals pass human review, but AI reviewers favor AI-authored text by about one point.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 01:13 UTC pith:XMOCQVYB
load-bearing objection A transparent, well-designed controlled study whose central parity claim leans on a human reviewer panel made up entirely of co-authors; fix that and the paper has real value. the 4 major comments →
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is a three-part asymmetry. First, capability: when blinded expert reviewers score one-page research plans on a four-aspect rubric, AI-generated plans are rated no worse than plans written by human experts, with mean totals around 3.5 out of 5 for both. Second, detectability: human experts can tell AI from human authorship about 72–79% of the time, while the two most capable AI reviewers classified all 32 proposals correctly, 100%. Third, bias: every AI reviewer tested scored AI-written proposals about one point higher than human-written ones, a pro-AI preference absent from the human panel; human-written proposals received nearly identical scores from every evaluator, s
What carries the argument
The controlled corpus of 32 one-page proposals: eight expert-conceived project seeds (title, background, goal) each expanded by one human expert and three LLMs under an identical four-section template and fixed prompt, plus the same four-aspect scoring rubric (clarity and structure, appropriateness of methods, resource and tool planning, feasibility/timeline/risk) and binary origin-judgment task given to all reviewers. The design isolates content from stylistic tells — uniform formatting, grammar normalization of human text — so that observed differences in scores and origin judgments can be attributed to author type rather than formatting.
Load-bearing premise
The valid comparison assumes the four human reviewers — co-authors of this paper — judge the proposals exactly as an independent grant panel would, with no recognition of their colleagues' writing and no stake in the outcome; if they recognized or favored particular proposals, both the human/AI parity and the reported detection rates could be inflated.
What would settle it
Run the identical 32-proposal, blinded review with eight reviewers from outside the author list who have never seen the proposals, and check whether human-rated AI-vs-human means remain within 0.2 points and origin-detection accuracy lands near 72/79%. A second decisive test: give both AI reviewers the same proposals with instructions to ignore style and score only scientific substance; if the roughly one-point pro-AI gap vanishes, the bias is stylistic rather than substantive.
If this is right
- If short structured proposals are representative, researchers can use LLMs to draft competitive project plans without losing quality in human review.
- AI-reviewer pro-AI bias means incorporating LLM evaluators into grant review without calibration would systematically disadvantage human-written proposals.
- Perfect AI authorship detection on this corpus suggests current frontier models can identify AI-generated planning text at high accuracy, though the paper cautions against generalizing from 32 proposals.
- The tight anti-correlation between AI advantage and human-plan quality implies AI assistance may be most valuable where the human baseline plan is weak, and least valuable where it is already strong.
- Per-aspect scores show AI plans are weakest on feasibility, timeline, and risk awareness under both human and AI reviewers, so AI assistance is least reliable for realistic scheduling and contingency planning.
Where Pith is reading between the lines
- The near-constant vertical offset between human and AI panels suggests a simple calibration: shifting AI-reviewer scores down by about one point for AI-flagged proposals would align their project-level rankings with the human panel's; this is testable on the paper's own data.
- Because the pro-AI bias appeared across all four AI models and both vendors, it likely stems from stylistic regularities of LLM output — template-like five-phase structure, exhaustive tool lists, round-number timelines — rather than any single model's idiosyncrasy, so the bias may persist even as models improve unless review instructions explicitly penalize those markers.
- The perfect AI detection rate may be an artifact of the constrained one-page template, since LLM proposals in this corpus share a detectable 'too clean' structure; running the same origin-judgment task on full-length, unconstrained proposals would test whether accuracy collapses.
- A decisive external check the authors did not run: repeat the review with non-author reviewers who have no stake in the outcome, to rule out recognition effects on the 72/79% human detection rates and the human parity ratings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a controlled, blinded study of AI-assisted scientific project planning and proposal evaluation. Eight expert-written project seeds in physics, astrophysics, and cosmology were each expanded into one human-written and three LLM-written one-page proposals, yielding 32 proposals. Four human reviewers (all co-authors) and two frontier LLM reviewers evaluated all proposals on a four-aspect rubric and also made a binary human-versus-AI authorship judgment. The reported results are: human reviewers rated human- and AI-written proposals similarly overall; both AI reviewers scored AI-written proposals roughly one point higher than human-written ones; AI reviewers classified all 32 proposals correctly, while human reviewers averaged 72% and 79% correct on human- and AI-written proposals, respectively; and the AI-human score gap is strongly anti-correlated with the quality of the human-written proposal. The authors conclude that LLMs can currently produce project plans comparable to human ones in the eyes of human reviewers, but that AI reviewers exhibit a systematic pro-AI bias that warrants caution in deploying LLMs in proposal preparation and review. The manuscript is transparent about the small sample, the rapid model turnover, and the fact that the human reviewers are authors, and it includes appendices with project seeds, prompts, and sample proposals.
Significance. If the results hold, this is a useful and timely controlled contribution to the growing literature on LLMs in scientific workflows. The study design has real strengths: a fixed template and common starting point for human and AI proposals, uniform anonymized formatting, explicit disclosure of evaluator identity, and rich qualitative material on the cues used for authorship judgments. The finding of a systematic preference by AI reviewers for AI-written proposals, if robust, has direct policy relevance for grant review, and the authors appropriately cite agency guidelines that restrict AI use in peer review. However, the paper's central quantitative claims rest on a small sample with no inferential statistics, and the human-reviewer panel consists entirely of co-authors, which is a serious independence risk. The human-written condition is also contaminated by a ChatGPT grammar-correction pass, and one of the paper's key project-level correlations is partly mechanical. These issues are substantive but addressable through re-analysis and more careful framing, so the paper is best treated as a strong pilot study requiring major revision rather than a definitive measurement.
major comments (4)
- [Sec. II B / III B] The human panel that grounds the parity claim consists of four co-authors. The manuscript discloses this but does not control for or test the obvious recognition risk: reviewers may recognize the human planners' prose or project identities, which could inflate the human-written proposals' quality ratings and contaminate the 72%/79% origin-detection rates. No inter-rater reliability statistic is reported; the paper even notes that the four reviewers 'often disagreed' (Sec. III B). With eight projects per condition, one or two non-independent judges can determine the aggregate parity result. This is load-bearing because the abstract's first claim rests entirely on this panel. Please report per-reviewer scores and agreement (e.g., ICC or Fleiss' kappa), and either add independent external reviewers or re-frame the human-panel component as a pilot with the independence threat as a central li
- [Sec. II C] The 'human-written' proposals were post-processed by ChatGPT 4o. Thus the human condition is not purely human-authored; AI stylistic normalization may introduce the very template-like, polished features that the reviewers (human and AI) report using to flag AI text (Sec. III A). This can inflate both the AI reviewers' 100% classification and the human reviewers' perception of parity. The manuscript should quantify the edit distance or amount of rewriting, justify that 'minimal changes' indeed preserved content and style, and preferably include a no-AI-touch control arm, or at least treat the condition as 'human-written then AI-normalized' in all claims.
- [Sec. III B, Tables IV and V] The central quantitative claims—that human reviewers rate human and AI proposals similarly and that AI reviewers give AI proposals about one point more—are supported only by group means and standard deviations across n=8 projects. No confidence intervals, p-values, or effect sizes are given, and the small sample means the mean difference can be dominated by one project (e.g., GW, which is called out in Sec. III C). Please provide per-project data, bootstrapped or permutation confidence intervals for the gaps, and ideally a mixed-effects model with reviewer and project as random effects. Without this, the 'similarly overall' and 'one point higher' phrasing overstates the precision of the measurements.
- [Sec. III C, Fig. 4] The reported r=-0.95 between the AI-human score gap and the human-written proposal score is largely mechanical, because the gap subtracts the human score from itself. Even if the AI-written score were statistically independent of the human-written score, a strong negative correlation would appear; the reported value is therefore not evidence that the AI advantage is concentrated in projects with weak human plans. The correct diagnostic is a scatter plot of mean AI score versus human score with a fitted slope, or a formal model. Please re-analyze before drawing the conclusion that the parity and bias are 'driven by the same handful of projects.'
minor comments (5)
- [Abstract / Sec. III A] The abstract says 'two AI reviewers' but Table III also lists Codex 5.5 Pro and Claude Sonnet 4.6; clarify which are primary and which are supplementary, and make the table caption consistent with the main text.
- [Sec. II D / Appendix B] The AI evaluation prompt is not included; Appendix B gives proposal-generation prompts only. Since the AI reviewers' behavior depends on the exact instruction, please include the reviewer prompt in the appendix or a supplement.
- [Table II] Minor typo in the rubric header: 'W eak' should be 'Weak'.
- [Reproducibility] No data/code availability statement appears. Since the study is empirical and based on LLM outputs, depositing de-identified proposals, reviewer scores, and prompts would substantially strengthen reproducibility.
- [Sec. IV / References] Reference [20] is listed as arXiv:2607.xxxx with a placeholder; this should be updated to the actual companion paper ID. Also, the paper's concluding caveats do not mention the ChatGPT grammar-correction pass on human proposals or the co-author reviewer bias; these belong in the limitations paragraph.
Circularity Check
No significant circularity: the study is an empirical measurement, not a derivation, and no reported finding reduces by construction to its inputs.
full rationale
The paper reports a controlled empirical comparison: fixed expert-provided seeds (title, background, goal) are given to human and AI writers, and the resulting proposals are scored by human and AI reviewers. The central claims—human/AI parity, AI reviewers' pro-AI bias, and near-perfect AI authorship detection—are measured outcomes, not quantities constructed from the inputs. No parameter is fitted and then renamed as a prediction; no equation sets a reported result equal to an input; the rubric and template are shared inputs rather than derived outputs. The cited companion Paper I [20] merely notes that the same project seeds are used there and does not carry any load-bearing justification; it is a normal companion reference, not a circularity. The paper's disclosed limitation that four human reviewers are co-authors (Sec. II B) is a genuine external-validity concern, since those reviewers might recognize human-written plans or have conflicts of interest, but this is not circularity: it does not make the human-parity rating equal by construction to the study design. The paper's own caveats—small sample, reviewer disagreement, rapidly changing AI models, and evaluation limited to planning rather than novelty—further show that the findings are contingent empirical results rather than tautologies.
Axiom & Free-Parameter Ledger
free parameters (1)
- Equal rubric aspect weights =
w = 1/4 for each of four aspects
axioms (6)
- domain assumption Human expert rubric scores are the validity standard for proposal quality
- domain assumption The four human reviewers, though co-authors, evaluate proposals impartially and without recognizing the human planners
- domain assumption Blind formatting plus a light ChatGPT grammar pass removes superficial writing cues without changing content
- domain assumption One-page templated proposals are a representative proxy for research project planning
- domain assumption The four rubric criteria are the relevant dimensions of proposal quality and are equally important
- domain assumption AI reviewers were not exposed to the true origin labels or authors' identities during evaluation
read the original abstract
We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.
Figures
Reference graph
Works this paper leans on
-
[1]
This demonstrates AI’s potential to assist with project planning in scientific workflows
On capability, at the level of a short, structured research proposal, current LLMs produce research plans comparable to those of human researchers, as rated by expert human reviewers (Figures 2, 3). This demonstrates AI’s potential to assist with project planning in scientific workflows
-
[2]
Hu- man reviewers were well above chance (≳70%) but relied on noisy and sometimes conflicting cues
Humans and AI are not equally good at telling if a proposal is written by human or AI (Table III). Hu- man reviewers were well above chance (≳70%) but relied on noisy and sometimes conflicting cues. The two most capable AI reviewers applied essentially the same cues as human reviewers but did so con- sistently and correctly, classifying all 32 proposals w...
-
[3]
AI style
Pro-AI bias was observed in all AI evaluations, where all four AI reviewers scored AI-written pro- posals roughly a point higher than human-written ones (Fig. 2). Such bias is absent in human eval- uation. This pro-AI bias is a concrete risk for any review system that incorporates LLMs: an AI reviewer may favor AI-written proposals, poten- tially disadvan...
-
[4]
III C, Fig
These aggregate patterns hide strong project-to- project variation (Sec. III C, Fig. 4): the human panel preferred the human-written proposal in five of the eight projects, and the AI–human score gap is tightly anti-correlated with the quality of the hu- man plan (r=−0.95). Both the human–AI parity and the pro-AI bias are thus concentrated in a few projec...
2020
-
[5]
E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn, inProceedings of the 40th International Con- ference on Machine Learning (ICML), PMLR, Vol. 202 (2023) pp. 24950–24962, arXiv:2301.11305
Pith/arXiv arXiv 2023
-
[6]
Y. D. Hezaveh, L. Perreault Levasseur, and P. J. Mar- shall, Nature548, 555 (2017)
2017
-
[7]
Villaescusa-Navarro, D
F. Villaescusa-Navarro, D. Angl´ es-Alc´ azar, S. Genel, D. N. Spergel, R. S. Somerville, R. Dave, A. Pillepich, L. Hernquist, D. Nelson, P. Torrey,et al., The Astro- physical Journal915, 71 (2021)
2021
-
[8]
T. D. Nguyen, Y.-S. Ting, I. Ciuc˘ a, C. O’Neill, Z.-C. Sun, M. Jab lo´ nska, S. Kruk, E. Perkowski, J. Miller, J. Li,et al., inProceedings of the Second Work- shop on Information Extraction from Scientific Publica- tions (WIESP), IJCNLP-AACL 2023(2023) pp. 49–55, arXiv:2309.06126
Pith/arXiv arXiv 2023
-
[9]
H. Zhou, H. Huang, Y. Long, B. Xu, C. Zhu, H. Cao, M. Yang, and T. Zhao, inProceedings of the 23rd Chi- nese National Conference on Computational Linguistics (CCL)(2024) pp. 1310–1319, arXiv:2409.16788
Pith/arXiv arXiv 2024
-
[10]
C. Si, D. Yang, and T. Hashimoto, inInternational Conference on Learning Representations (ICLR)(2025) arXiv:2409.04109
Pith/arXiv arXiv 2025
-
[11]
C. A. Gao, F. M. Howard, N. S. Markov, E. C. Dyer, S. Ramesh, Y. Luo, and A. T. Pearson, npj Digital Medicine6, 75 (2023)
2023
-
[12]
found that LLMs can provide feedback on papers that researchers often find useful; [8] reported low inter- reviewer agreement together with quality-rating inflation by AI (also see [9, 11]). What remains comparatively unexplored is a controlled, blinded comparison of human and AI performance on an open-ended research-planning task, evaluated by both human...
Pith/arXiv arXiv 2026
-
[13]
C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha, arXiv e-prints (2024), arXiv:2408.06292
Pith/arXiv arXiv 2024
-
[14]
Shcherbiak, H
A. Shcherbiak, H. Habibnia, R. B¨ ohm, and S. Fiedler, Judgment and Decision Making19, e21 (2024)
2024
-
[15]
A. Panickssery, S. R. Bowman, and S. Feng, inAdvances in Neural Information Processing Systems (NeurIPS), Vol. 37 (2024) arXiv:2404.13076
Pith/arXiv arXiv 2024
-
[16]
K. Wataoka, T. Takahashi, and R. Ri, arXiv e-prints (2024), arXiv:2410.21819. Presented at the NeurIPS 2024 Safe Generative AI Workshop
Pith/arXiv arXiv 2024
-
[17]
W. Liang, Y. Zhang, H. Cao, B. Wang, D. Y. Ding, X. Yang, K. Vodrahalli, S. He, D. S. Smith, Y. Yin, D. A. McFarland, and J. Zou, NEJM AI1, 10.1056/AIoa2400196 (2024), arXiv:2310.01783
Pith/arXiv arXiv 2024
-
[18]
V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, and S. Feizi, arXiv e-prints (2023), arXiv:2303.11156
Pith/arXiv arXiv 2023
- [19]
-
[20]
Sikimi´ c, Synthese206, 282 (2025)
V. Sikimi´ c, Synthese206, 282 (2025). 10
2025
-
[21]
U. Sandstr¨ om and M. Thelwall, arXiv e-prints (2026), arXiv:2603.14565
arXiv 2026
- [22]
-
[23]
F. Villaescusa-Navarro, B. Bolliet, P. Villanueva- Domingo, A. E. Bayer, A. Acquah, C. Amancharla, A. Barzilay-Siegal, P. Bermejo, C. Bilodeau, P. C. Ram ´ ırez, M. Cranmer, U. L. Fran¸ ca, C. Hahn, Y.- F. Jiang, R. Jimenez, J.-Y. Lee, A. Lerario, O. Ma- mun, T. Meier, A. A. Ojha, P. Protopapas, S. Roy, D. N. Spergel, P. Taranc´ on-´Alvarez, U. Tiwari, M....
arXiv 2025
-
[24]
A. Hell and L. Thiele, LLMs with in-context learn- ing for Algorithmic Theoretical Physics (2026), arXiv:2605.08212 [cs.LG]
Pith/arXiv arXiv 2026
-
[25]
A. Hell, K. Vovk, V. Krishnaraj, J. Liu, K. Aizawa, A. E. Bayer, L. Blot, J. Cowell, S. Garg, J. Gr´ ee, B. Horowitz, M. Ichikawa, K. Iemoto, K. Kondo, Z. Lorsin, K. Mc- Carthy, J. Robinson, M. Ruiz-Granda, L. Thiele, I. Vovk, and M. Zhou, AI’s Capability in Assisting Scientific Re- search in Physics, Astrophysics, and Cosmology I: Liter- ature Review, ar...
-
[26]
National Institutes of Health, The use of generative artificial intelligence technologies is prohibited for the NIH peer review process, NIH Guide Notice NOT- OD-23-149 (2023),https://grants.nih.gov/grants/ guide/notice-files/NOT-OD-23-149.html
2023
-
[27]
National Science Foundation, Notice to the re- search community: Use of generative artificial intelligence technology in the NSF merit re- view process (2023),https://www.nsf.gov/news/ notice-to-the-research-community-on-ai
2023
-
[28]
See alsohttps://erc.europa.eu/news-events/news/ erc-clarifies-limits-ai-use-grant-evaluation
European Research Council, The use of AI in grant proposal evaluation (2026),https: //erc.europa.eu/system/files/2026-03/ Use-AI-grant-proposal-evaluation.pdf. See alsohttps://erc.europa.eu/news-events/news/ erc-clarifies-limits-ai-use-grant-evaluation. Appendix A: Background and goals of the eight research projects The following project titles, backgroun...
2026
-
[29]
AGN – MaNGA: AGN Duty Cycle Background:The time a galaxy spends in the AGN phase, from both general arguments and ensemble studies such as quasar clustering and black-hole mass-function studies and Heiiproximity-zone analysis, is suggested to last∼10 6–109 yr. Ionization studies of AGN host galax- ies and their surroundings indicate that active nuclei can...
-
[30]
drop-out
LBG – The galaxy–dark matter halo connection of Lyman-break galaxies Background:A Lyman-break galaxy (LBG) is a galaxy whose broadband photometry shows a “drop-out” in the bluest bands as features in its spectrum move from blue to red through the filter set due to cosmic expansion; the reduction in flux blueward of the Lyman-αand Lyman- limit frequencies ...
-
[31]
Many studies have therefore investigated the characteristics of 11 IA in order to eliminate it from the data
IA – Intrinsic alignments in varying environments Background:Weak-lensing surveys are one of the most powerful probes in cosmology; however, the intrinsic alignment (IA) of galaxies contaminates the signal. Many studies have therefore investigated the characteristics of 11 IA in order to eliminate it from the data. So far, re- searchers believe IA is rela...
-
[32]
AR – Prediction of debris emergence on laser-ablated sub-wavelength shapes Background:We have been developing methods to fabricate sub-wavelength structures (SWS) for anti- reflective coating in the millimeter-wave region on hard materials such as ceramics, using ultra-short-pulse laser ablation, which is crucial for machining materials with relatively wi...
-
[33]
RG – Radio Galaxies with HalfDome Background:At low CMB frequencies (around 100 GHz), high-energy radio galaxies act as bright point- source contaminants to CMB maps. The locations of these galaxies are likely correlated with features in the underlying large-scale structure as well as with galaxy properties (e.g., the CIB, radio continuum, X-ray). Goal:Ad...
-
[34]
However, the origins of these binary black holes and the environments they reside in remain unknown
GW – Environment of gravitational-wave black hole binaries with weak-lensing maps Background:The first detection of a gravitational wave (GW) by LIGO opened a new era of multi- messenger astronomy, and around 300 GW events from binary black hole (BBH) mergers have now been ob- served. However, the origins of these binary black holes and the environments t...
-
[35]
PT A – F orecasting pulsar timing array sensitivity to deviations from general relativity Background:The Pulsar Timing Array (PTA) is a measurement method relying on the observation of pul- sars, fast-rotating neutron stars with well-known timing models. By measuring slight perturbations in the times of arrival (ToAs) of each pulse, computing residuals, a...
2023
-
[36]
For instance, introducing a Chern–Simons cou- pling between a pseudo-scalar field and a non-Abelian gauge field can lead to slow-roll inflation, as in Chromo- 12 Natural Inflation
SU2 – Massive Y ang–Mills theory Background:When exploring mechanisms that drive inflation, non-Abelian gauge fields – such as SU(2) Yang– Mills fields – have been proposed as alternatives to scalar fields. For instance, introducing a Chern–Simons cou- pling between a pseudo-scalar field and a non-Abelian gauge field can lead to slow-roll inflation, as in...
-
[37]
Project Title, copied from the title provided
-
[38]
Background, a one-sentence summary
-
[39]
Goal, a one-sentence summary; and
-
[40]
Methodology, broken into no more than five ma- jor steps or phases (e.g., data preparation, model- ing, analysis), each with an approximate comple- tion time, in under 300 words. All parties were told that the proposal would be evalu- ated on four criteria: clarity and structure of the research plan; appropriateness of methods to the scientific goal; reso...
-
[41]
Correct typo and grammar mis- takes, with minimal change to the content: [text]
Methodology: Break down the work into major steps or phases (e.g., data preparation, modeling, analysis). Include theoretical, computational, or experimental techniques if relevant. Include the approximate time for each step. No more than 5 steps and keep it under 300 words. Use formal scientific tone appropriate for a grant or academic setting. Your prop...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.