REVIEW 4 major objections 5 minor 4 cited by
Every 2025 high-performance computing conference proceedings examined contains 'mysterious citations' to papers with no evidence of existence, affecting up to 6% of published papers.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Fabricated “mysterious” citations appear in every 2025 HPC conference proceeding examined, affecting 2-6% of papers, while the 2021 baseline showed none.
T0 review reviewed 2026-08-03 challenge →
load-bearing objection A useful, honest measurement of fabricated citations in published HPC proceedings, with a rate estimate that is plausible but not fully auditable. the 4 major comments →
The Case of the Mysterious Citations
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that fabricated or 'mysterious' citations — references whose titles do not match any existing publication and whose cited locations hold no related work — have entered peer-reviewed proceedings since generative AI became widespread. Comparing the same four HPC conferences in 2021 and 2025, the authors report zero mysterious citations in the 2021 baseline but at least one in every 2025 proceeding, affecting 2–6% of papers. In one paper, over half of the citations did not exist at the cited location. The authors also report rephrased titles (real papers cited under altered titles), frequent missing or extra authors, incorrect DOIs, and a citation containing a URL w
What carries the argument
The central mechanism is the paper's citation-verification pipeline combined with the definition of a 'mysterious citation.' The pipeline extracts a PDF's bibliography, parses each entry into title, authors, venue, DOI, and arXiv identifier, then searches six external bibliographic aggregators and follows up with manual web searches, checking cited page locations directly when aggregators fail. A citation is classified as mysterious when no paper with a similar title exists, and the cited location either does not exist or contains an unrelated paper. This classification — the binary of found vs. not-found across aggregators — is what bears the paper's empirical weight.
Load-bearing premise
The classification of a citation as 'mysterious' rests on the completeness of the six bibliographic aggregators and the authors' manual searches; if a genuine but poorly indexed or nonstandard publication exists that those sources miss, it would be counted as fabricated, inflating the reported rates.
What would settle it
A reader could take a random sample of citations flagged as mysterious and check them in physical library catalogs and by directly contacting the cited authors, or calibrate the pipeline on a set of known obscure but real publications; if a substantial share turn out to exist, the 'no evidence of existence' conclusion weakens.
If this is right
- If the central claim holds, current peer review is not catching fabricated references, meaning published literature contains unverifiable citations that undermine the traceability of scientific claims.
- Conference policies requiring AI disclosure are widely ignored: no paper in the dataset disclosed AI-generated citations despite all four venues requiring it.
- A concrete fingerprint of AI-generated content (e.g., a URL containing utm_source=chatgpt.com) can be used to flag suspicious bibliographies for manual review.
- Citation-verification tools like the one built for this study could be adopted as a standard pre-publication check, since manual verification does not scale.
- The reported rise in title and author errors, independent of full fabrication, suggests LLM-assisted bibliographies degrade citation quality even when references are real.
Where Pith is reading between the lines
- The 2–6% figure is likely a lower bound on the true rate of problematic citations: the pipeline flags only citations that fail verification against aggregators and manual search, and some real but obscure works may be falsely classified as mysterious, while other fabricated citations that happen to point to real-sounding but wrong venues might be missed.
- If the pattern generalizes beyond HPC proceedings, the same methodology could be applied to medical, social science, or humanities venues, where citation practices differ and fabricated references could more severely distort evidence synthesis.
- The ResearchGate case — a fake paper uploaded the same day as the conference deadline and then cited — suggests a new failure mode: authors citing non-peer-reviewed, plausibly fabricated sources that appear online, which requires reviewers to scrutinize the provenance of cited items.
- Replicated over a larger sample with independent verification, this measurement approach could serve as a routine quality metric for conference programs, such as a 'citation integrity score' per venue or per year.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an automated pipeline (PyPDF extraction, heuristic bibliography splitting, regex parsing of 12 reference formats, lookup in arXiv/Crossref/dblp/Google Books/OpenAlex/OSTI, then manual validation) applied to proceedings of four HPC conferences (anonymized as Conferences 0–3), comparing 2021 and 2025. It reports no mysterious citations in 2021, at least one in every 2025 conference, affecting 2–6% of papers; increased author-list errors; and no author acknowledgments of AI-generated citations despite the four conference policies. It explicitly disclaims proof that AI caused the errors.
Significance. If the central existence claim holds, it provides rare systematic evidence that fabricated/unverifiable citations are present in peer-reviewed proceedings, with direct policy implications. Strengths include the explicit manual validation step, the detailed ResearchGate case, the use of multiple external aggregators, and the honest Limitations section. However, the rates are sensitive to the negative-existential classification rule and the incomplete 2021 baseline, so the numerical claims should be treated as provisional pending data release and independent audit.
major comments (4)
- [Methodology, steps 5–6 and Fig. 1] The definition of a 'mysterious citation' is the absence of evidence across six aggregators plus manual web search. This is a negative existential claim: real but under-indexed works (workshop papers, institutional technical reports, non-English venues, removed preprints, undigitized books) could be missed. The paper gives no formal standard for when 'no evidence of existence' is conclusive, no inter-rater reliability check, and no full list of flagged citations or search queries. Please publish the flagged entries and a search log, and have at least two independent raters classify ambiguous cases, reporting agreement. The ResearchGate case is compelling, but one documented case does not by itself establish the 2–6% range.
- [Figure 2 and Data collection] The text states 'The 2021 proceedings were analyzed as a baseline while the 2025 proceedings of the same conferences were examined,' but Figure 2 reports 2021 data only for Conferences 0 and 1. Please clarify whether the 2021 proceedings of Conferences 2 and 3 were analyzed. If not, the abstract/conclusion claim that 'none of the 2021 papers contained mysterious citations' is not supported for those venues. Report the 2021 data for all four conferences, or restrict the claim accordingly.
- [AI-Generated Citations / Abstract] The abstract and conclusion claim that 'all four conference policies required it' — i.e., required authors to acknowledge AI-generated citations — is stronger than the body's statement that the conferences 'required authors to acknowledge any generative AI usage.' A general AI-disclosure policy does not necessarily require specific acknowledgment of AI-generated citations. Quote the exact policy language for each of the four conferences. Also, the single 'utm_source=chatgpt.com' URL is anecdotal; it cannot by itself substantiate the claim that the observed errors stem from undisclosed AI use, though the Limitations section appropriately acknowledges this.
- [Results / Figures 2–4] The reported rates are based on small counts (2–6% of proceedings with <50–100+ papers, so per-conference counts are likely single digits), and no error bars or statistical tests are given. The classification also depends on the 12 empirically identified reference formats, which may handle 2021 and 2025 templates differentially. Please provide per-conference raw counts, the number of papers analyzed, the threshold separating 'rephrased title' from 'mysterious citation,' and an error or sensitivity analysis.
minor comments (5)
- [Figure 2] The y-axis label says 'Number of Citation Errors Across All Papers' while the caption describes the number of papers with at least one citation issue; align the axis label with the caption.
- [Methodology, page 2] 'it’s relationship' should be 'its relationship.'
- [References] Reference 16 uses a DOI URL ('https://www.doi.org/10.1126/science.zl99qni'); provide the standard DOI link and an access date consistent with the other references.
- [Figure 1] The aggregator list inconsistently spells 'Arxiv' instead of 'arXiv'; use one spelling consistently.
- [Discussion] The statement that 'fully hallucinated citations are likely to disappear as generative AI continues to improve' is speculative and not supported by the data; consider removing it or clearly labeling it as opinion.
Circularity Check
No significant circularity; the central claim is a directly operationalized measurement, with causal attribution explicitly disclaimed.
full rationale
The paper contains no derived equations, fitted parameters, or self-citation chain that would make the result equivalent to its inputs. The central claim is an empirical measurement: PDFs are parsed, bibliographies are isolated, entries are searched against six external aggregators (ArXiv, Crossref, dblp, Google Books, OpenAlex, OSTI), and ambiguous cases are manually checked. The 'mysterious citation' category is defined operationally in the Methodology/Mysterious Citations sections ('No paper a similar enough title exists'), and the reported finding restates that operationalized classification rather than deriving it from an independent premise, which is normal for a measurement study and not circular. The paper explicitly disclaims causal attribution to LLMs in the LIMITATIONS section: 'It is not possible to prove that a citation error was generated through an LLM hallucination.' It also hedges its strongest claim as 'citations to papers for which we were able to find no evidence of existence,' acknowledging the absence-of-evidence nature of the check. The only soft point is that the abstract and conclusion call the citations 'AI-hallucinated' despite the disclaimer, but that is an interpretive overreach, not a circular reduction. The measurement is externally grounded, so there is no load-bearing circular step.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption External bibliographic aggregators (ArXiv, Crossref, dblp, Google Books, OpenAlex, OSTI) plus the authors' manual web searches are complete enough to establish non-existence of a cited work.
- domain assumption The PDF-to-text and bibliography-isolation heuristics (last occurrence of 'references'/'bibliography', numbered-entry splitting, 12 empirically-identified formats) do not systematically miss or corrupt citations.
- ad hoc to paper All four anonymized conference policies required authors to acknowledge AI-generated citations, not merely general AI use.
- domain assumption No unmodeled human-factor confound explains the 2021 to 2025 rise in rephrased titles and mysterious citations.
Cite this review
Pith. "Pith review of The Case of the Mysterious Citations." pith.science (2026). https://pith.science/paper/XSR6W2GF
@misc{pith2026260205867,
author = {Pith},
title = {Pith review of: The Case of the Mysterious Citations},
year = {2026},
howpublished = {\url{https://pith.science/paper/XSR6W2GF}},
note = {Machine review of arXiv:2602.05867}
}
read the original abstract
Mysterious citations are routinely appearing in peer-reviewed publications throughout the scientific community. In this paper, we developed an automated pipeline and examine the proceedings of four major high-performance computing conferences, comparing the accuracy of citations between the 2021 and 2025 proceedings. While none of the 2021 papers contained mysterious citations, every 2025 proceeding did, impacting 2-6% of published papers. In addition, we observe a sharp rise in paper title and authorship errors, motivating the need for stronger citation-verification practice. No author within our dataset acknowledged using AI to generate citations even though all four conference policies required it, indicating current policies are insufficient.
Forward citations
Cited by 4 Pith papers
-
HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers
HALLMARK shows that citation verifiers' false-positive rate, not recall, is what determines whether their flags are mostly real catches or mostly noise at realistic hallucination rates.
-
sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing
An open-source local linter verifies reference integrity and claim support in scientific manuscripts using public databases and consumer hardware, with an experimental contribution scoring extension.
-
HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
HalluCiteChecker is a lightweight, offline, CPU-only toolkit that detects hallucinated citations in AI-assisted scientific papers.
-
sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing
A local lint pipeline verifies paper citations and internal consistency and proposes a SciLint Score combining citation-chain integrity with philosophy-of-science criteria.
Reference graph
Works this paper leans on
-
[1]
ChatGPT
OpenAI, “ChatGPT.” https://openai.com/blog/chatgpt,
-
[2]
Accessed: 2026-01-13
Anthropic, “Claude.” https://www.anthropic.com/ claude, 2023. Accessed: 2026-01-13
2023
-
[3]
Accessed: 2026-01-13
Google, “Gemini.” https://deepmind.google/ technologies/gemini/, 2023. Accessed: 2026-01-13
2023
-
[4]
Quantifying large language model usage in scientific papers,
W. Liang, Y . Zhang, Z. Wu, H. Lepp, W. Ji, X. Zhao, H. Cao, S. Liu, S. He, Z. Huang,et al., “Quantifying large language model usage in scientific papers,” Nature Human Behaviour, pp. 1–11, 2025
2025
-
[5]
ACM pol- icy on authorship
Association for Computing Machinery, “ACM pol- icy on authorship.” https://www.acm.org/publications/ policies/new-acm-policy-on-authorship, 2023. In- 6 preprint. SAND2026-16932O Feb 2026 THEME cludes guidance on the use of generative AI in manuscript preparation. Accessed: 2026-01-13
2023
-
[6]
Submission and peer review policies: Guidelines on AI-generated text
Institute of Electrical and Electronics Engineers, “Submission and peer review policies: Guidelines on AI-generated text.” https://journals.ieeeauthorcenter.ieee.org/become- an-ieee-journal-author/publishing-ethics/guidelines- and-policies/submission-and-peer-review-policies/,
-
[7]
Acm policy on plagiarism, misrepresentation, and content fal- sification
Association for Computing Machinery, “Acm policy on plagiarism, misrepresentation, and content fal- sification.” https://www.acm.org/publications/policies/ plagiarism-overview, 2025. ACM defines content fal- sification as intentional misrepresentation of results, supporting materials, or references, including manu- factured or irrelevant citations that do...
2025
-
[8]
Ieee ethical requirements for authors: Avoiding fabrication, falsification, and citation misconduct
Institute of Electrical and Electronics Engineers, “Ieee ethical requirements for authors: Avoiding fabrication, falsification, and citation misconduct.” https://journals.ieeeauthorcenter.ieee.org/become- an-ieee-journal-author/publishing-ethics/ethical- requirements/, 2025. IEEE author ethics policies prohibit fabrication (inventing data/results), falsif...
2025
-
[9]
The ChatGPT lawyer explains himself
B. Weiser and N. Schweber, “The ChatGPT lawyer explains himself.” https://www.nytimes.com/2023/06/ 08/nyregion/lawyer-chatgpt-sanctions.html, 6 2023. Accessed: 2026-01-14
2023
-
[10]
GPTZero finds over 50 new hallucinations in ICLE2026 submis- sions
P . Esau, N. Shmatko, and A. Adam, “GPTZero finds over 50 new hallucinations in ICLE2026 submis- sions.” https://gptzero.me/news/iclr-2026/, 2025. Ac- cessed: 2026-01-14
2026
-
[11]
Gptzero: Detection of ai- generated text,
E. Tian and A. Cui, “Gptzero: Detection of ai- generated text,” 2026. Accessed: 2026-01-15
2026
-
[12]
Startup investigation reveals 50 peer-reviewed papers contained ai-hallucinated citations
S. Martin, “Startup investigation reveals 50 peer-reviewed papers contained ai-hallucinated citations.” https://betakit.com/start-up-investigation- reveals-50-peer-reviewed-papers-contained- hallucinated-citations/?utm_source=chatgpt.com,
-
[13]
Accessed: 2026-01-15
PyPDF Contributors, “Pypdf.” https://github.com/py- pdf/pypdf, 2026. Accessed: 2026-01-15
2026
-
[14]
ERA conference rankings
“ERA conference rankings.” http://www. conferenceranks.com. Accessed: 2026-01-21
2026
-
[15]
From content creation to citation inflation: A genai case study
H. S. Al-Sinani and C. J. Mitchell, “From content creation to citation inflation: A genai case study.” https://arxiv.org/abs/2503.23414, 2025. Accessed: 2026-01-16
Pith/arXiv arXiv 2025
-
[16]
How easy is it to fudge your scientific rank? meet larry, the world’s most cited cat
C. Wilcox, “How easy is it to fudge your scientific rank? meet larry, the world’s most cited cat.” https: //www.doi.org/10.1126/science.zl99qni, 2024. Ac- cessed: 2026-01-16. Amanda Bienzis an Assistant Professor in the Com- puter Science Department at the University of New Mexico. Her research interests include improving the performance and scalability o...
-
[2022]
Accessed: 2026-01-13
2026
-
[2023]
Accessed: 2026-01-13
Defines disclosure requirements for AI- generated content in IEEE publications. Accessed: 2026-01-13
2026
-
[2025]
Published December 16, 2025; accessed 2026-01-15
2025
This paper was first reviewed by deepseek-v4-flash on August 3, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.