REVIEW 3 major objections 5 minor 43 references
A controlled study of eight physics research projects finds that mid-2025 AI assistants rarely match expert literature selections and that most AI-supplied references are real papers with corrupted metadata rather than invented ones.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 01:48 UTC pith:CBLLWQW7
load-bearing objection Useful controlled study; the 64% metadata-mismatch headline needs a normalization sensitivity check before trusting it. the 3 major comments →
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper 's own terms, the central discovery is that the dominant failure mode of LLM-generated bibliographies in physics is not the invention of papers (3%) but the corruption of real references: 64% are papers that exist with at least one wrong title, author, year, journal, DOI, or link field. Among mid-2025 models, ChatGPT Deep Research was the most reliable, producing no fabrications and only 22% mismatches, while Gemini produced the most errors. The paper also reports that a single-project test of ChatGPT Pro 5.5 produced only perfect references, though the authors caution this may reflect stronger tool use rather than a change in the underlying language model alone.
What carries the argument
The method is a controlled, parallel literature search with a standardized prompt. For each of eight expert-conceived projects, a human expert and three LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini) independently built bibliographies capped at 50 papers, and experts judged the AI suggestions for relevance. Hallucination was then measured by resolving each AI-only reference's DOI against a bibliographic metadata database, falling back to the provided link or a title search, and classifying each reference as perfect, metadata mismatch, or fabrication. This DOI-first verification chain is the mechanism that separates real-but-corrupted references from invented ones.
Load-bearing premise
The central 64% statistic relies on counting any difference from the metadata database's copy as an error, even though the prompt asked for last-name-only authors; relax that rule and the headline number could change.
What would settle it
Take the 408 references labeled metadata mismatches, normalize author fields to last-name-only, and recompute the mismatch rate while ignoring year and journal formatting differences; if the rate falls below 50%, the '64% require verification' claim would need a strong caveat.
If this is right
- If the 64% mismatch rate holds, LLM-generated reference lists in physics cannot be trusted at face value; each entry must be checked field by field before citation.
- The less-than-6% overlap with expert selection implies human and AI searches are largely complementary, so a combined human-plus-AI search should find more relevant work than either alone.
- The strong difference between plain chat models and a tool-augmented research model shows that verification-oriented architectures dramatically reduce hallucination.
- The single-project zero-error result from ChatGPT Pro 5.5 points to rapid improvement, but the paper itself flags that this is not yet a systematic benchmark.
Where Pith is reading between the lines
- If the recent-model trend generalizes, the near-term role of LLMs in literature review will shift from producing cite-able entries to generating discovery candidates, with humans or retrieval tools verifying metadata.
- The paper's counting rule treats any field disagreement, however trivial (such as a last-name-only author), as a mismatch; the true rate of practically harmful errors may be lower than 64%.
- The observation that AIs are keyword-driven while humans search more broadly suggests that prompting for adjacent fields and foundational works could close part of the relevance gap.
- Comparing against a metadata database rather than the original published version may itself introduce mismatches; a check against the publisher's own records would test that.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a controlled study of LLM-assisted literature review in physics, astrophysics, and cosmology. For eight expert-conceived projects, a human expert and three mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, Gemini) independently produced bibliographies using a standardized prompt. The authors compare human/AI overlap, then classify the 641 AI references not found by humans as perfect, metadata mismatch, or fabrication using a DOI/link/title verification pipeline against Crossref. Main quantitative findings are <6% human-AI overlap, 3% fabrications, and 64% metadata mismatches among analyzed AI references; ChatGPT Deep Research is more reliable, and a single-project ChatGPT Pro 5.5 spot-test yields zero errors. The paper concludes that mid-2025 models require systematic verification of AI-generated references.
Significance. If the quantitative claims hold, the paper provides useful, field-specific evidence on LLM bibliographic reliability in physics: the distinction between fabricated references and real-but-corrupted references is valuable, and the explicit verification pipeline (Fig. 1), per-model tables, and standardized prompt in Appendix B are strengths. The study is non-circular: AI output is compared against an external registry (Crossref), and the definitions are not outcome-dependent. However, the headline 64% mismatch rate and the <6% overlap claim depend on comparison conventions that are not fully specified or tested, so the central quantitative result needs robustness checks before being accepted at face value.
major comments (3)
- [§II.E, Appendix B, Table IV] The 64% metadata-mismatch rate is load-bearing and may be inflated by prompt-conformant formatting. Appendix B instructs models to write "only the last name of the first author," while the evaluation compares the generated first author against Crossref's full formatted name. Section II.E counts any "partial or complete" disagreement as a mismatch. If the comparison is raw string equality, every "Smith" vs "Smith, John" is scored as a first-author mismatch, and Table IV reports 306 first-author mismatches among 399 resolved mismatches. The paper does not state whether author names were normalized to surnames, whether journal names were canonicalized, or whether title punctuation/capitalization was normalized. Please specify the exact matching procedure and provide a sensitivity analysis, e.g., recounting mismatches with surname-only author comparison, journal-name canonicalization, and ti
- [§II.D, §III.A, Table II] The human-AI overlap definition requires a match in both title and category: a paper placed by the human as "recent" and by the AI as "highly cited" is counted as two different references, and the same paper placed by one model in two categories counts twice. This convention can artificially lower the measured overlap and affects the abstract's "<6%" claim. The authors should report a title-only overlap sensitivity analysis (ignoring category) and state how often human and AI agree on the title but disagree on the category. Without this, the overlap statistic conflates bibliographic coverage with subjective categorization.
- [§III.B, Table III, Abstract] The 3% fabrication and 64% mismatch rates are computed on 641 AI references "not found by a human," not on all 701 AI-generated references. The abstract's phrase "of the AI-generated references" is therefore imprecise; the 60 excluded references are a systematically different subset (human-confirmed relevant papers). Please either report the rates on the full 701-reference set or consistently qualify the denominator in the abstract and Section III.B. The magnitude of the effect is likely modest, but the current wording overstates the scope of the measurement.
minor comments (5)
- [Table III / text] The table header calls the row "AI-generated references" while the text specifies "those not found by a human." Make the qualifier explicit in the table itself to avoid misreading.
- [Section III.B / Fig. 3] The text refers to "Fig. III B" in two places; this should be "Fig. 3." The figure caption also does not mention that ChatGPT Pro 5.5 is a single-project spot-test; add that to the caption.
- [Section II.D] The sentence "The overall results are presented in subsection III A" appears twice in slightly different forms; remove the duplicate.
- [Section III.C] The handling of near-duplicate references for ChatGPT-4o (5 papers differing only in link or DOI) is described only here. State in Section II how near-duplicates are treated in the main analysis, since the same issue could affect the 701 total.
- [Table IV] The rows "1/4 mismatch" through "4/4 mismatch" should explicitly define the denominator as the four compared fields (title, first author, year, journal).
Circularity Check
No significant circularity: the central measurements are externally grounded (human expert lists and Crossref) and the only self-citations are contextual.
full rationale
The paper's derivation chain is a controlled measurement rather than a derivation from an assumed conclusion. Eight expert-written project backgrounds are turned into a standardized prompt (Appendix B); human experts and three LLMs independently produce reference lists; the human-vs-AI overlap is scored by title-plus-category matching; AI hallucinations are judged by resolving DOIs/links and comparing title, first author, year, and journal against Crossref entries (Sec. II.E). The 33% perfect / 3% fabrication / 64% metadata-mismatch breakdown is an empirical tally under an explicit classification rule, with no fitted parameter later renamed as a prediction and no equation reducing the result to the input. The only self-citations — [26] (an earlier LLM-physics study by two of the authors, cited as one supporting example) and [27] (the companion paper) — are contextual and do not carry the central claim. The acknowledged limitations (eight projects; single-project Pro 5.5 test) are scope caveats, not circular dependencies. A separate methodological concern is that the mismatch rule may count prompt-conformant forms ('write only the last name of the first author', Appendix B) as partial disagreements against Crossref's full author and journal strings, potentially inflating the 64% figure; however, that is a measurement-validity/sensitivity issue, not circularity, since the mismatch statistic is not constructed from the headline conclusion. Therefore no circular steps are flagged.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Human-expert lists of up to 50 'most important' references are a valid reference standard for a competent literature search.
- domain assumption Crossref metadata is the ground truth for judging AI field accuracy, and any partial disagreement in title, first author, year, or journal counts as a mismatch.
- ad hoc to paper The same paper placed in two different categories counts as two references in the overlap statistics.
- domain assumption Each LLM's output in mid-2025 is deterministic enough that running it once represents that model's performance.
- domain assumption The eight expert-conceived projects span representative frontier physics/astro/cosmology work.
read the original abstract
We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astrophysics, and cosmology. Each project has a defined background and goal, and human experts and AI prompters are asked to perform identical literature review tasks in parallel. We compare the relevant literature selected by humans with that selected by mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini). We find the overlap between human- and AI-selected references to be small ($<$6\%), indicating that AI models do not yet reproduce a competent expert search on their own, though they have the potential to complement literature searches by humans. We then assess the reliability and completeness of AI-generated candidate references, distinguishing two types of hallucination: fabrications (references to nonexistent papers) and metadata mismatches (real papers with one or more incorrect fields). We find that while fabricated references make up 3\% of the AI-generated references, 64\% are real papers with at least one incorrect field (title, author, year, journal, DOI, or link), indicating that the mid-2025 models require systematic verification. However, the performance is significantly improved for the 2026 model ChatGPT Pro 5.5, with a single-project test showing zero fabrication or metadata mismatches.
Figures
Reference graph
Works this paper leans on
-
[1]
Both AI and human experts generated a large set of relevant references that were essentially not over- lapping, indicating that a collaboration between humans and AI may yield a more complete refer- ence overview
-
[2]
For the mid-2025 models, 3% of the references are fabrications (nonexistent papers), while 64% are real papers with at least one incorrect field (a meta- data mismatch); both are forms of hallucination. This is one of the main weaknesses of AI, which has persisted in recent years in scientific literature searches, and suggests that one should carefully ch...
2025
-
[3]
Our findings also come with two important caveats
AIs appear to prefer more specialized search guided by the project keywords, focused narrowly on the project topic, while human experts appear to search more broadly, taking into account also re- lated fields or more fundamental papers. Our findings also come with two important caveats. First, we have limited our investigation to only eight projects, whil...
2020
-
[4]
Woesle, L
C. Woesle, L. Fischer-Brandies, and R. Buettner, IEEE Access13, 148231 (2025)
2025
-
[5]
W. H. Walters and E. I. Wilder, Sci Rep13, 14045 (2023)
2023
-
[6]
Scherbakov, N
D. Scherbakov, N. Hubig, V. Jansari, A. Bakumenko, and L. A. Lenert, Journal of the American Medical Informat- ics Association32, 1071–1086 (2025)
2025
-
[7]
W. Liang, Y. Zhang, H. Cao, B. Wang, D. Y. Ding, X. Yang, K. Vodrahalli, S. He, D. S. Smith, Y. Yin, D. A. McFarland, and J. Zou, NEJM AI1, 10.1056/AIoa2400196 (2024), arXiv:2310.01783
Pith/arXiv arXiv 2024
-
[8]
M. Wang, R. Lin, K. Hu, J. Jiao, N. Chowd- hury, E. Chang, and T. Patwardhan, arXiv e-prints , arXiv:2601.21165 (2026), arXiv:2601.21165 [cs.AI]
arXiv 2026
-
[9]
Villaescusa-Navarro, D
F. Villaescusa-Navarro, D. Angl´ es-Alc´ azar, S. Genel, D. N. Spergel, R. S. Somerville, R. Dave, A. Pillepich, L. Hernquist, D. Nelson, P. Torrey,et al., The Astro- physical Journal915, 71 (2021)
2021
-
[10]
T. D. Nguyen, Y.-S. Ting, I. Ciuc˘ a, C. O’Neill, Z.-C. Sun, M. Jab lo´ nska, S. Kruk, E. Perkowski, J. Miller, J. Li,et al., inProceedings of the Second Work- shop on Information Extraction from Scientific Publica- tions (WIESP), IJCNLP-AACL 2023(2023) pp. 49–55, arXiv:2309.06126
Pith/arXiv arXiv 2023
-
[11]
A. L. Franzoni Vel´ azquez, E. Huerta, and S. Jensen, Dis- cov Educ 3226, 226 (2024)
2024
-
[12]
M. e. a. Pastucha, Medical science monitor : interna- tional medical journal of experimental and clinical re- search32, e950916 (2026)
2026
-
[13]
Y. D. Hezaveh, L. Perreault Levasseur, and P. J. Mar- shall, Nature548, 555 (2017)
2017
-
[14]
C. Lu, C. Lu, R. T. Lange, J. Foerster, J. Clune, and D. Ha, arXiv e-prints (2024), arXiv:2408.06292
Pith/arXiv arXiv 2024
-
[15]
Shcherbiak, H
A. Shcherbiak, H. Habibnia, R. B¨ ohm, and S. Fiedler, Judgment and Decision Making19, e21 (2024)
2024
-
[16]
H. Zhou, H. Huang, Y. Long, B. Xu, C. Zhu, H. Cao, M. Yang, and T. Zhao, inProceedings of the 23rd Chi- nese National Conference on Computational Linguistics (CCL)(2024) pp. 1310–1319, arXiv:2409.16788
Pith/arXiv arXiv 2024
-
[17]
E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn, inProceedings of the 40th International Con- ference on Machine Learning (ICML), PMLR, Vol. 202 (2023) pp. 24950–24962, arXiv:2301.11305
Pith/arXiv arXiv 2023
-
[18]
C. A. Gao, F. M. Howard, N. S. Markov, E. C. Dyer, S. Ramesh, Y. Luo, and A. T. Pearson, npj Digital Medicine6, 75 (2023)
2023
-
[19]
V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, and S. Feizi, arXiv e-prints (2023), arXiv:2303.11156
Pith/arXiv arXiv 2023
- [20]
-
[21]
A. Panickssery, S. R. Bowman, and S. Feng, inAdvances in Neural Information Processing Systems (NeurIPS), Vol. 37 (2024) arXiv:2404.13076
Pith/arXiv arXiv 2024
-
[22]
C. Si, D. Yang, and T. Hashimoto, inInternational Conference on Learning Representations (ICLR)(2025) 10 arXiv:2409.04109
Pith/arXiv arXiv 2025
-
[23]
K. Wataoka, T. Takahashi, and R. Ri, arXiv e-prints (2024), arXiv:2410.21819. Presented at the NeurIPS 2024 Safe Generative AI Workshop
Pith/arXiv arXiv 2024
-
[24]
F. Villaescusa-Navarro, B. Bolliet, P. Villanueva- Domingo, A. E. Bayer, A. Acquah, C. Amancharla, A. Barzilay-Siegal, P. Bermejo, C. Bilodeau, P. C. Ram ´ ırez, M. Cranmer, U. L. Fran¸ ca, C. Hahn, Y.- F. Jiang, R. Jimenez, J.-Y. Lee, A. Lerario, O. Ma- mun, T. Meier, A. A. Ojha, P. Protopapas, S. Roy, D. N. Spergel, P. Taranc´ on-´Alvarez, U. Tiwari, M....
arXiv 2025
-
[25]
T. Miaoet al., PRL-Bench: A Comprehensive Bench- mark Evaluating LLMs’ Capabilities in Frontier Physics Research (2026), arXiv:2604.15411 [cs.LG]
Pith/arXiv arXiv 2026
-
[26]
Sikimi´ c, Synthese206, 282 (2025)
V. Sikimi´ c, Synthese206, 282 (2025)
2025
-
[27]
U. Sandstr¨ om and M. Thelwall, arXiv e-prints (2026), arXiv:2603.14565
arXiv 2026
- [28]
-
[29]
Hell, JHEP03, 167, arXiv:2111.00017 [hep-th]
A. Hell, JHEP03, 167, arXiv:2111.00017 [hep-th]. Appendix A: Background and goals of the eight research projects The following project titles, backgrounds, and goals were written by the human experts without any AI as- sistance, and were the common input provided to the human planners and the AI prompters for each project (Sec. II A)
-
[31]
A. Hell and L. Thiele, LLMs with in-context learn- ing for Algorithmic Theoretical Physics (2026), arXiv:2605.08212 [cs.LG]
Pith/arXiv arXiv 2026
-
[32]
J. Liu, V. Krishnaraj, K. Vovk, , K. Aizawa, A. E. Bayer, L. Blot, J. Cowell, S. Garg, J. Gr´ ee, A. Hell, B. Horowitz, M. Ichikawa, K. Iemoto, K. Kondo, Z. Lorsin, K. McCarthy, J. Robinson, M. Ruiz-Granda, L. Thiele, I. Vovk, and M. Zhou, AI’s Capability in As- sisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and P...
-
[33]
A. E. Bayer, Y. Zhong, Z. Li, J. DeRose, Y. Feng, and J. Liu, JCAP05, 016, arXiv:2407.17462 [astro-ph.CO]
-
[35]
AGN – MaNGA: AGN Duty Cycle Background:The time a galaxy spends in the AGN phase, from both general arguments and ensemble studies such as quasar clustering and black-hole mass-function studies and Heiiproximity-zone analysis, is suggested to last∼10 6–109 yr. Ionization studies of AGN host galax- ies and their surroundings indicate that active nuclei can...
-
[36]
drop-out
LBG – The galaxy–dark matter halo connection of Lyman-break galaxies Background:A Lyman-break galaxy (LBG) is a galaxy whose broadband photometry shows a “drop-out” in the bluest bands as features in its spectrum move from blue to red through the filter set due to cosmic expansion; the reduction in flux blueward of the Lyman-αand Lyman- limit frequencies ...
-
[37]
Many studies have therefore investigated the characteristics of 11 IA in order to eliminate it from the data
IA – Intrinsic alignments in varying environments Background:Weak-lensing surveys are one of the most powerful probes in cosmology; however, the intrinsic alignment (IA) of galaxies contaminates the signal. Many studies have therefore investigated the characteristics of 11 IA in order to eliminate it from the data. So far, re- searchers believe IA is rela...
-
[38]
AR – Prediction of debris emergence on laser-ablated sub-wavelength shapes Background:We have been developing methods to fabricate sub-wavelength structures (SWS) for anti- reflective coating in the millimeter-wave region on hard materials such as ceramics, using ultra-short-pulse laser ablation, which is crucial for machining materials with relatively wi...
-
[39]
RG – Radio Galaxies with HalfDome Background:At low CMB frequencies (around 100 GHz), high-energy radio galaxies act as bright point- source contaminants to CMB maps. The locations of these galaxies are likely correlated with features in the underlying large-scale structure as well as with galaxy properties (e.g., the CIB, radio continuum, X-ray). Goal:Ad...
-
[40]
However, the origins of these binary black holes and the environments they reside in remain unknown
GW – Environment of gravitational-wave black hole binaries with weak-lensing maps Background:The first detection of a gravitational wave (GW) by LIGO opened a new era of multi- messenger astronomy, and around 300 GW events from binary black hole (BBH) mergers have now been ob- served. However, the origins of these binary black holes and the environments t...
-
[41]
PT A – F orecasting pulsar timing array sensitivity to deviations from general relativity Background:The Pulsar Timing Array (PTA) is a measurement method relying on the observation of pul- sars, fast-rotating neutron stars with well-known timing models. By measuring slight perturbations in the times of arrival (ToAs) of each pulse, computing residuals, a...
2023
-
[42]
For instance, introducing a Chern–Simons cou- pling between a pseudo-scalar field and a non-Abelian gauge field can lead to slow-roll inflation, as in Chromo- 12 Natural Inflation
SU2 – Massive Y ang–Mills theory Background:When exploring mechanisms that drive inflation, non-Abelian gauge fields – such as SU(2) Yang– Mills fields – have been proposed as alternatives to scalar fields. For instance, introducing a Chern–Simons cou- pling between a pseudo-scalar field and a non-Abelian gauge field can lead to slow-roll inflation, as in...
-
[43]
Author” write only the last name of the first author of the paper. “Year
Recent papers/results on this topic (past 10 years) 4. Other papers that may be relevant The total number of papers you find (all 4 categories combined) should be no more than 50 papers. Limit to refereed papers only. As for the number of papers in each individual category it is up to you to determine what is appropriate. I want you to for each of the 4 c...
-
[641]
Human & AI
One should note that this number is smaller than in the previous subsection because here we do not dis- tinguish between the different categories into which the same paper may have been placed. Following the verifica- tion procedure of subsection II E (summarized in Fig. 1), each reference is classified as aperfectreference (all fields match the true pape...
2025
-
[2025]
drop-out
– supplemented with project-specific context. The relevance of the AI-generated candidate references was then evaluated by an expert. The overall material was also compared to human output and further analyzed to assess reliability and possible hallucinations. Finally, the literature search was repeated for one of the projects us- ing ChatGPT Pro 5.5 to s...
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.