{"id":"7206004a-9501-449a-b057-72e53aa22fee","arxiv_id":"2607.25672","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a controlled test, three mid-2025 LLMs shared under 6% of literature references with physics experts, and 64% of their real references had at least one metadata error.","lead":"Eight physics, astrophysics, and cosmology research projects were given to both human experts and three AI chatbots to find relevant papers: the two groups shared less than 6% of their references. Most real papers that the AI found had at least one wrong detail, so AI reference lists still need systematic checking before use.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported 64% metadata-mismatch rate may be inflated by counting prompt-conformant author/journal formatting as errors; needs normalization sensitivity check.","rationale":"The paper is a carefully designed, self-aware empirical study with internally consistent numbers and appropriate caveats about the single-project Pro 5.5 test. The reader's CONDITIONAL verdict is appropriate. The load-bearing weakness is indeed the metadata-mismatch definition: the comparison rule counts 'partial' disagreement as a mismatch, while the prompt explicitly asked models to provide only the first author's last name. Since first-author mismatches dominate Table IV, a modest normalization could substantially reduce the reported 64%. This is not an accusation of fraud or carelessness; it is a missing sensitivity analysis that is directly addressable. The central claim would still hold in spirit if the mismatch rate dropped to, say, 50%, but the abstract's precise '64%' is not robust without a canonicalization check. I agree with the reader's weakest_assumption, and the recommended verdict remains CONDITIONAL rather than ACCEPT, REJECT, or UNVERDICTED, because the issue is testable and the rest of the analysis is sound.","tokens_in":14798,"tokens_out":6494,"duration_ms":71585,"concrete_test":"Run a reanalysis of Table III/IV using the raw AI-vs-Crossref comparison data: canonicalize each author cell by extracting the surname token from Crossref's 'Family, Given' and comparing with the AI cell after removing initials; canonicalize journal names to the standard abbreviation (e.g., map 'Physical Review D' to 'Phys. Rev. D'); ignore case/punctuation in titles. Recompute the number of perfect references and the aggregate mismatch rate under (i) exact-string rule as currently implied, (ii) surname-only + journal-abbreviation-tolerant rule, and (iii) a semantic rule requiring genuinely wrong authors/years/journals. If rule (ii) or (iii) drops the 64% below ~50% (or especially if author mismatches drop from 306 to a small number), the headline is a comparison artifact; if the rate stays near 64%, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim (Abstract; §III.B; Table III) that 64% of mid-2025 AI-generated references are real papers with incorrect metadata rests on the classification rule in §II.E: a reference is a 'metadata mismatch' if any of title, first author, year, or journal 'disagree[s], either partially or completely' with Crossref. However, Appendix B instructs the models to write 'only the last name of the first author', while Crossref's author field is a full formatted name (e.g., 'Smith, John'). If the evaluation compares raw strings, every 'Smith' vs 'Smith, John' is scored as a first-author mismatch. First-author mismatch is reported for 306 of the 399 resolvable mismatches (Table IV), so this one formatting issue alone can move a large fraction of the 408 mismatches into the 'perfect' category. Journal abbreviations ('Phys. Rev. D' vs 'Physical Review D') and title punctuation/capitalization are analogous. The paper does not report any normalization of author names to surnames, journal-name canonicalization, or sensitivity analysis around the partial-disagreement rule. Because the 64% headline is the paper's main quantitative finding and the basis for the 'systematic verification' conclusion, the missing normalization assumption is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a controlled study of LLM-assisted literature review in physics, astrophysics, and cosmology. For eight expert-conceived projects, a human expert and three mid-2025 LLMs (ChatGPT-4o, ChatGPT Deep Research, Gemini) independently produced bibliographies using a standardized prompt. The authors compare human/AI overlap, then classify the 641 AI references not found by humans as perfect, metadata mismatch, or fabrication using a DOI/link/title verification pipeline against Crossref. Main quantitative findings are <6% human-AI overlap, 3% fabrications, and 64% metadata mismatches among analyzed AI references; ChatGPT Deep Research is more reliable, and a single-project ChatGPT Pro 5.5 spot-test yields zero errors. The paper concludes that mid-2025 models require systematic verification of AI-generated references.","tokens_in":15057,"tokens_out":4891,"duration_ms":53408,"significance":"If the quantitative claims hold, the paper provides useful, field-specific evidence on LLM bibliographic reliability in physics: the distinction between fabricated references and real-but-corrupted references is valuable, and the explicit verification pipeline (Fig. 1), per-model tables, and standardized prompt in Appendix B are strengths. The study is non-circular: AI output is compared against an external registry (Crossref), and the definitions are not outcome-dependent. However, the headline 64% mismatch rate and the <6% overlap claim depend on comparison conventions that are not fully specified or tested, so the central quantitative result needs robustness checks before being accepted at face value.","major_comments":[{"comment":"The 64% metadata-mismatch rate is load-bearing and may be inflated by prompt-conformant formatting. Appendix B instructs models to write \"only the last name of the first author,\" while the evaluation compares the generated first author against Crossref's full formatted name. Section II.E counts any \"partial or complete\" disagreement as a mismatch. If the comparison is raw string equality, every \"Smith\" vs \"Smith, John\" is scored as a first-author mismatch, and Table IV reports 306 first-author mismatches among 399 resolved mismatches. The paper does not state whether author names were normalized to surnames, whether journal names were canonicalized, or whether title punctuation/capitalization was normalized. Please specify the exact matching procedure and provide a sensitivity analysis, e.g., recounting mismatches with surname-only author comparison, journal-name canonicalization, and ti","section":"§II.E, Appendix B, Table IV"},{"comment":"The human-AI overlap definition requires a match in both title and category: a paper placed by the human as \"recent\" and by the AI as \"highly cited\" is counted as two different references, and the same paper placed by one model in two categories counts twice. This convention can artificially lower the measured overlap and affects the abstract's \"<6%\" claim. The authors should report a title-only overlap sensitivity analysis (ignoring category) and state how often human and AI agree on the title but disagree on the category. Without this, the overlap statistic conflates bibliographic coverage with subjective categorization.","section":"§II.D, §III.A, Table II"},{"comment":"The 3% fabrication and 64% mismatch rates are computed on 641 AI references \"not found by a human,\" not on all 701 AI-generated references. The abstract's phrase \"of the AI-generated references\" is therefore imprecise; the 60 excluded references are a systematically different subset (human-confirmed relevant papers). Please either report the rates on the full 701-reference set or consistently qualify the denominator in the abstract and Section III.B. The magnitude of the effect is likely modest, but the current wording overstates the scope of the measurement.","section":"§III.B, Table III, Abstract"}],"minor_comments":[{"comment":"The table header calls the row \"AI-generated references\" while the text specifies \"those not found by a human.\" Make the qualifier explicit in the table itself to avoid misreading.","section":"Table III / text"},{"comment":"The text refers to \"Fig. III B\" in two places; this should be \"Fig. 3.\" The figure caption also does not mention that ChatGPT Pro 5.5 is a single-project spot-test; add that to the caption.","section":"Section III.B / Fig. 3"},{"comment":"The sentence \"The overall results are presented in subsection III A\" appears twice in slightly different forms; remove the duplicate.","section":"Section II.D"},{"comment":"The handling of near-duplicate references for ChatGPT-4o (5 papers differing only in link or DOI) is described only here. State in Section II how near-duplicates are treated in the main analysis, since the same issue could affect the 701 total.","section":"Section III.C"},{"comment":"The rows \"1/4 mismatch\" through \"4/4 mismatch\" should explicitly define the denominator as the four compared fields (title, first author, year, journal).","section":"Table IV"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and likely to be of interest to astro-ph.IM readers. The central concern is not circularity but matching normalization: the 64% claim and the <6% overlap claim both rest on comparison conventions that need to be made explicit and stress-tested. This is fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful controlled study of LLM literature search in physics/astro/cosmology. The <6% human-AI overlap and the 3% fabrication rate are believable and worth knowing. The 64% metadata-mismatch number is the one to be careful with.\n\nWhat it does well: eight expert-conceived projects, standardized prompt, independent human and AI searches, expert relevance judgment, and a clear DOI→link→title verification pipeline. The two-type hallucination split (fabrication vs metadata mismatch) is a good taxonomy, and the per-model breakdown—Deep Research clearly better than plain chat models—is informative. The study is self-aware, with explicit caveats in the conclusion.\n\nThe soft spot that matters: the mismatch classification. Appendix B tells models to write only the last name of the first author. Crossref returns full names. If the comparison in §II.E compares raw strings, every 'Smith' vs 'Smith, John' counts as a first-author mismatch, and Table IV says 306 of 399 resolvable mismatches involve the first author. That could move a large chunk of the 408 mismatches into 'perfect.' The paper doesn't say whether surnames were normalized, nor does it report any sensitivity analysis around the partial-disagreement rule. Journal abbreviations and title punctuation are analogous. This is load-bearing for the abstract's 64% claim. I don't think the paper is wrong in direction—the mismatch rate is high no matter what—but the specific number could shift meaningfully.\n\nOther gaps are minor and addressable: the Pro 5.5 'significant improvement' is one project; there are no error bars anywhere; and the full reference lists are not released, so independent re-checking isn't possible. The counting of the same title in two categories as two references is explained and defensible.\n\nBottom line: this is a solid empirical study that deserves a serious referee. The main action items for revision are to document how fields were normalized before comparison, run a sensitivity analysis around the partial-disagreement rule, and release the reference lists. I'd cite it once those are in.","headline":"Useful controlled study; the 64% metadata-mismatch headline needs a normalization sensitivity check before trusting it.","tokens_in":15626,"tokens_out":2123,"would_cite":true,"duration_ms":22513,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A controlled study of eight physics research projects finds that mid-2025 AI assistants rarely match expert literature selections and that most AI-supplied references are real papers with corrupted metadata rather than invented ones.","keywords":["literature review","large language models","hallucination","metadata mismatch","reference verification","physics research","astrophysics","cosmology"],"falsifier":"Take the 408 references labeled metadata mismatches, normalize author fields to last-name-only, and recompute the mismatch rate while ignoring year and journal formatting differences; if the rate falls below 50%, the '64% require verification' claim would need a strong caveat.","tokens_in":14681,"feed_emoji":"🤖","tokens_out":5112,"duration_ms":52217,"temperature":0.7,"pith_summary":"The paper asks whether large language models can help physicists discover relevant literature at the research frontier. It sets up a controlled test: eight expert-designed projects in physics, astrophysics, and cosmology, each searched independently by a human expert and by three mid-2025 AI assistants. The AI produced many references, but almost none overlapped with the experts' choices, and a detailed check of the AI-only references found most were real papers carrying at least one incorrect field, not outright fabrications. A single-project test of a later 2026 model produced only perfect references, suggesting rapid improvement. The takeaway: AI can flag extra papers a human might miss, but its bibliographies need systematic verification.","feed_headline":"64% of AI-generated citations have metadata errors","feed_subtitle":"Only 3% are invented, but under 6% overlap with expert picks; a 2026 test run was perfect.","key_machinery":"The method is a controlled, parallel literature search with a standardized prompt. For each of eight expert-conceived projects, a human expert and three LLMs (ChatGPT-4o, ChatGPT Deep Research, and Gemini) independently built bibliographies capped at 50 papers, and experts judged the AI suggestions for relevance. Hallucination was then measured by resolving each AI-only reference's DOI against a bibliographic metadata database, falling back to the provided link or a title search, and classifying each reference as perfect, metadata mismatch, or fabrication. This DOI-first verification chain is the mechanism that separates real-but-corrupted references from invented ones.","core_discovery":"On the paper 's own terms, the central discovery is that the dominant failure mode of LLM-generated bibliographies in physics is not the invention of papers (3%) but the corruption of real references: 64% are papers that exist with at least one wrong title, author, year, journal, DOI, or link field. Among mid-2025 models, ChatGPT Deep Research was the most reliable, producing no fabrications and only 22% mismatches, while Gemini produced the most errors. The paper also reports that a single-project test of ChatGPT Pro 5.5 produced only perfect references, though the authors caution this may reflect stronger tool use rather than a change in the underlying language model alone.","pith_inferences":["If the recent-model trend generalizes, the near-term role of LLMs in literature review will shift from producing cite-able entries to generating discovery candidates, with humans or retrieval tools verifying metadata.","The paper's counting rule treats any field disagreement, however trivial (such as a last-name-only author), as a mismatch; the true rate of practically harmful errors may be lower than 64%.","The observation that AIs are keyword-driven while humans search more broadly suggests that prompting for adjacent fields and foundational works could close part of the relevance gap.","Comparing against a metadata database rather than the original published version may itself introduce mismatches; a check against the publisher's own records would test that."],"forward_implications":["If the 64% mismatch rate holds, LLM-generated reference lists in physics cannot be trusted at face value; each entry must be checked field by field before citation.","The less-than-6% overlap with expert selection implies human and AI searches are largely complementary, so a combined human-plus-AI search should find more relevant work than either alone.","The strong difference between plain chat models and a tool-augmented research model shows that verification-oriented architectures dramatically reduce hallucination.","The single-project zero-error result from ChatGPT Pro 5.5 points to rapid improvement, but the paper itself flags that this is not yet a systematic benchmark."],"fun_headline_variants":["AI citations: 64% real papers with wrong metadata","AI literature picks barely overlap experts, 64% have errors","Deep Research safest: 0 fakes, 22% metadata misses","Pro 5.5 perfect in single citation test","LLM bibliographies: 3% fake, 64% wrong fields"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central 64% statistic relies on counting any difference from the metadata database's copy as an error, even though the prompt asked for last-name-only authors; relax that rule and the headline number could change.","fun_headline_variants_meta":{"raw":{"variants":["AI citations: 64% real papers with wrong metadata","AI literature picks barely overlap experts, 64% have errors","Deep Research safest: 0 fakes, 22% metadata misses","Pro 5.5 perfect in single citation test","LLM bibliographies: 3% fake, 64% wrong fields"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000948,"raw_usage":{"total_tokens":3907,"prompt_tokens":791,"completion_tokens":3116,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":3028}},"tokens_in":535,"tokens_out":3116,"duration_ms":19734,"temperature":1.0,"reasoning_tokens":3028,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T01:48:36.015823+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 408 references labeled metadata mismatches, normalize author fields to last-name-only, and recompute the mismatch rate while ignoring year and journal formatting differences; if the rate falls below 50%, the '64% require verification' claim would need a strong caveat.","supporting_citations":[],"review_version":1}