Pith. sign in

REVIEW 4 major objections 5 minor 4 cited by

Every 2025 high-performance computing conference proceedings examined contains 'mysterious citations' to papers with no evidence of existence, affecting up to 6% of published papers.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 04:04 UTC pith:XSR6W2GF

load-bearing objection A useful, honest measurement of fabricated citations in published HPC proceedings, with a rate estimate that is plausible but not fully auditable. the 4 major comments →

arxiv 2602.05867 v2 pith:XSR6W2GF submitted 2026-02-05 cs.DL

The Case of the Mysterious Citations

classification cs.DL
keywords citation fabricationhallucinated citationsgenerative AIbibliographic verificationpeer reviewscholarly integrityhigh-performance computingproceedings analysis
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether generative AI has introduced fabricated references into peer-reviewed conference proceedings. It builds an automated citation-verification pipeline and checks every paper in the 2021 and 2025 proceedings of four high-performance computing conferences. The 2021 baseline contains no unverifiable citations, while every 2025 proceeding contains at least one 'mysterious citation' — a reference for which no matching publication could be found. The paper also documents a sharp increase in title and authorship errors, and finds that no author acknowledged using AI to generate citations even though all four venues required disclosure. The authors argue these errors are passing through peer review, and current policies are insufficient to catch them.

Core claim

The paper's central claim is that fabricated or 'mysterious' citations — references whose titles do not match any existing publication and whose cited locations hold no related work — have entered peer-reviewed proceedings since generative AI became widespread. Comparing the same four HPC conferences in 2021 and 2025, the authors report zero mysterious citations in the 2021 baseline but at least one in every 2025 proceeding, affecting 2–6% of papers. In one paper, over half of the citations did not exist at the cited location. The authors also report rephrased titles (real papers cited under altered titles), frequent missing or extra authors, incorrect DOIs, and a citation containing a URL w

What carries the argument

The central mechanism is the paper's citation-verification pipeline combined with the definition of a 'mysterious citation.' The pipeline extracts a PDF's bibliography, parses each entry into title, authors, venue, DOI, and arXiv identifier, then searches six external bibliographic aggregators and follows up with manual web searches, checking cited page locations directly when aggregators fail. A citation is classified as mysterious when no paper with a similar title exists, and the cited location either does not exist or contains an unrelated paper. This classification — the binary of found vs. not-found across aggregators — is what bears the paper's empirical weight.

Load-bearing premise

The classification of a citation as 'mysterious' rests on the completeness of the six bibliographic aggregators and the authors' manual searches; if a genuine but poorly indexed or nonstandard publication exists that those sources miss, it would be counted as fabricated, inflating the reported rates.

What would settle it

A reader could take a random sample of citations flagged as mysterious and check them in physical library catalogs and by directly contacting the cited authors, or calibrate the pipeline on a set of known obscure but real publications; if a substantial share turn out to exist, the 'no evidence of existence' conclusion weakens.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • If the central claim holds, current peer review is not catching fabricated references, meaning published literature contains unverifiable citations that undermine the traceability of scientific claims.
  • Conference policies requiring AI disclosure are widely ignored: no paper in the dataset disclosed AI-generated citations despite all four venues requiring it.
  • A concrete fingerprint of AI-generated content (e.g., a URL containing utm_source=chatgpt.com) can be used to flag suspicious bibliographies for manual review.
  • Citation-verification tools like the one built for this study could be adopted as a standard pre-publication check, since manual verification does not scale.
  • The reported rise in title and author errors, independent of full fabrication, suggests LLM-assisted bibliographies degrade citation quality even when references are real.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The 2–6% figure is likely a lower bound on the true rate of problematic citations: the pipeline flags only citations that fail verification against aggregators and manual search, and some real but obscure works may be falsely classified as mysterious, while other fabricated citations that happen to point to real-sounding but wrong venues might be missed.
  • If the pattern generalizes beyond HPC proceedings, the same methodology could be applied to medical, social science, or humanities venues, where citation practices differ and fabricated references could more severely distort evidence synthesis.
  • The ResearchGate case — a fake paper uploaded the same day as the conference deadline and then cited — suggests a new failure mode: authors citing non-peer-reviewed, plausibly fabricated sources that appear online, which requires reviewers to scrutinize the provenance of cited items.
  • Replicated over a larger sample with independent verification, this measurement approach could serve as a routine quality metric for conference programs, such as a 'citation integrity score' per venue or per year.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents an automated pipeline (PyPDF extraction, heuristic bibliography splitting, regex parsing of 12 reference formats, lookup in arXiv/Crossref/dblp/Google Books/OpenAlex/OSTI, then manual validation) applied to proceedings of four HPC conferences (anonymized as Conferences 0–3), comparing 2021 and 2025. It reports no mysterious citations in 2021, at least one in every 2025 conference, affecting 2–6% of papers; increased author-list errors; and no author acknowledgments of AI-generated citations despite the four conference policies. It explicitly disclaims proof that AI caused the errors.

Significance. If the central existence claim holds, it provides rare systematic evidence that fabricated/unverifiable citations are present in peer-reviewed proceedings, with direct policy implications. Strengths include the explicit manual validation step, the detailed ResearchGate case, the use of multiple external aggregators, and the honest Limitations section. However, the rates are sensitive to the negative-existential classification rule and the incomplete 2021 baseline, so the numerical claims should be treated as provisional pending data release and independent audit.

major comments (4)
  1. [Methodology, steps 5–6 and Fig. 1] The definition of a 'mysterious citation' is the absence of evidence across six aggregators plus manual web search. This is a negative existential claim: real but under-indexed works (workshop papers, institutional technical reports, non-English venues, removed preprints, undigitized books) could be missed. The paper gives no formal standard for when 'no evidence of existence' is conclusive, no inter-rater reliability check, and no full list of flagged citations or search queries. Please publish the flagged entries and a search log, and have at least two independent raters classify ambiguous cases, reporting agreement. The ResearchGate case is compelling, but one documented case does not by itself establish the 2–6% range.
  2. [Figure 2 and Data collection] The text states 'The 2021 proceedings were analyzed as a baseline while the 2025 proceedings of the same conferences were examined,' but Figure 2 reports 2021 data only for Conferences 0 and 1. Please clarify whether the 2021 proceedings of Conferences 2 and 3 were analyzed. If not, the abstract/conclusion claim that 'none of the 2021 papers contained mysterious citations' is not supported for those venues. Report the 2021 data for all four conferences, or restrict the claim accordingly.
  3. [AI-Generated Citations / Abstract] The abstract and conclusion claim that 'all four conference policies required it' — i.e., required authors to acknowledge AI-generated citations — is stronger than the body's statement that the conferences 'required authors to acknowledge any generative AI usage.' A general AI-disclosure policy does not necessarily require specific acknowledgment of AI-generated citations. Quote the exact policy language for each of the four conferences. Also, the single 'utm_source=chatgpt.com' URL is anecdotal; it cannot by itself substantiate the claim that the observed errors stem from undisclosed AI use, though the Limitations section appropriately acknowledges this.
  4. [Results / Figures 2–4] The reported rates are based on small counts (2–6% of proceedings with <50–100+ papers, so per-conference counts are likely single digits), and no error bars or statistical tests are given. The classification also depends on the 12 empirically identified reference formats, which may handle 2021 and 2025 templates differentially. Please provide per-conference raw counts, the number of papers analyzed, the threshold separating 'rephrased title' from 'mysterious citation,' and an error or sensitivity analysis.
minor comments (5)
  1. [Figure 2] The y-axis label says 'Number of Citation Errors Across All Papers' while the caption describes the number of papers with at least one citation issue; align the axis label with the caption.
  2. [Methodology, page 2] 'it’s relationship' should be 'its relationship.'
  3. [References] Reference 16 uses a DOI URL ('https://www.doi.org/10.1126/science.zl99qni'); provide the standard DOI link and an access date consistent with the other references.
  4. [Figure 1] The aggregator list inconsistently spells 'Arxiv' instead of 'arXiv'; use one spelling consistently.
  5. [Discussion] The statement that 'fully hallucinated citations are likely to disappear as generative AI continues to improve' is speculative and not supported by the data; consider removing it or clearly labeling it as opinion.

Circularity Check

0 steps flagged

No significant circularity; the central claim is a directly operationalized measurement, with causal attribution explicitly disclaimed.

full rationale

The paper contains no derived equations, fitted parameters, or self-citation chain that would make the result equivalent to its inputs. The central claim is an empirical measurement: PDFs are parsed, bibliographies are isolated, entries are searched against six external aggregators (ArXiv, Crossref, dblp, Google Books, OpenAlex, OSTI), and ambiguous cases are manually checked. The 'mysterious citation' category is defined operationally in the Methodology/Mysterious Citations sections ('No paper a similar enough title exists'), and the reported finding restates that operationalized classification rather than deriving it from an independent premise, which is normal for a measurement study and not circular. The paper explicitly disclaims causal attribution to LLMs in the LIMITATIONS section: 'It is not possible to prove that a citation error was generated through an LLM hallucination.' It also hedges its strongest claim as 'citations to papers for which we were able to find no evidence of existence,' acknowledging the absence-of-evidence nature of the check. The only soft point is that the abstract and conclusion call the citations 'AI-hallucinated' despite the disclaimer, but that is an interpretive overreach, not a circular reduction. The measurement is externally grounded, so there is no load-bearing circular step.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The paper introduces no fitted numbers or new physical/formal entities. Its central measurement depends on domain assumptions about search completeness, parsing correctness, the content of anonymized conference policies, and the absence of unmodeled confounds; these are documented in the Methodology and Limitations.

axioms (4)
  • domain assumption External bibliographic aggregators (ArXiv, Crossref, dblp, Google Books, OpenAlex, OSTI) plus the authors' manual web searches are complete enough to establish non-existence of a cited work.
    Methodology steps 5-6. If a cited paper is real but unindexed or not findable, it is misclassified as a mysterious citation, inflating the reported 2-6% rates.
  • domain assumption The PDF-to-text and bibliography-isolation heuristics (last occurrence of 'references'/'bibliography', numbered-entry splitting, 12 empirically-identified formats) do not systematically miss or corrupt citations.
    Methodology steps 1-4. Parsing failures would change the numerator and denominator of the reported fractions.
  • ad hoc to paper All four anonymized conference policies required authors to acknowledge AI-generated citations, not merely general AI use.
    Abstract and 'AI-Generated Citations' section. No conference policy documents are cited because venues are anonymized, so the requirement is asserted rather than verified; the policy-insufficiency conclusion depends on it.
  • domain assumption No unmodeled human-factor confound explains the 2021 to 2025 rise in rephrased titles and mysterious citations.
    LIMITATIONS paragraph concedes the authors 'do not explicitly account for' confounding factors; the rise is attributed to LLM-era publication practices rather than, e.g., changes in editorial thoroughness or indexing coverage.

pith-pipeline@v1.3.0-alltime-deepseek · 6968 in / 12893 out tokens · 131679 ms · 2026-08-03T04:04:33.676083+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of The Case of the Mysterious Citations." pith.science (2026). https://pith.science/paper/XSR6W2GF

@misc{pith2026260205867,
  author       = {Pith},
  title        = {Pith review of: The Case of the Mysterious Citations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XSR6W2GF}},
  note         = {Machine review of arXiv:2602.05867}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Mysterious citations are routinely appearing in peer-reviewed publications throughout the scientific community. In this paper, we developed an automated pipeline and examine the proceedings of four major high-performance computing conferences, comparing the accuracy of citations between the 2021 and 2025 proceedings. While none of the 2021 papers contained mysterious citations, every 2025 proceeding did, impacting 2-6% of published papers. In addition, we observe a sharp rise in paper title and authorship errors, motivating the need for stronger citation-verification practice. No author within our dataset acknowledged using AI to generate citations even though all four conference policies required it, indicating current policies are insufficient.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers

    cs.CR 2026-07 conditional novelty 6.0

    HALLMARK shows that citation verifiers' false-positive rate, not recall, is what determines whether their flags are mostly real catches or mostly noise at realistic hallucination rates.

  2. sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing

    cs.DL 2026-04 unverdicted novelty 6.0

    An open-source local linter verifies reference integrity and claim support in scientific manuscripts using public databases and consumer hardware, with an experimental contribution scoring extension.

  3. HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists

    cs.CL 2026-04 unverdicted novelty 5.0

    HalluCiteChecker is a lightweight, offline, CPU-only toolkit that detects hallucinated citations in AI-assisted scientific papers.

  4. sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing

    cs.DL 2026-04 unverdicted novelty 5.0

    A local lint pipeline verifies paper citations and internal consistency and proposes a SciLint Score combining citation-chain integrity with philosophy-of-science criteria.

Reference graph

Works this paper leans on

19 extracted references · 1 canonical work pages · cited by 3 Pith papers

  1. [1]

    ChatGPT

    OpenAI, “ChatGPT.” https://openai.com/blog/chatgpt,

  2. [2]

    Accessed: 2026-01-13

    Anthropic, “Claude.” https://www.anthropic.com/ claude, 2023. Accessed: 2026-01-13

  3. [3]

    Accessed: 2026-01-13

    Google, “Gemini.” https://deepmind.google/ technologies/gemini/, 2023. Accessed: 2026-01-13

  4. [4]

    Quantifying large language model usage in scientific papers,

    W. Liang, Y . Zhang, Z. Wu, H. Lepp, W. Ji, X. Zhao, H. Cao, S. Liu, S. He, Z. Huang,et al., “Quantifying large language model usage in scientific papers,” Nature Human Behaviour, pp. 1–11, 2025

  5. [5]

    ACM pol- icy on authorship

    Association for Computing Machinery, “ACM pol- icy on authorship.” https://www.acm.org/publications/ policies/new-acm-policy-on-authorship, 2023. In- 6 preprint. SAND2026-16932O Feb 2026 THEME cludes guidance on the use of generative AI in manuscript preparation. Accessed: 2026-01-13

  6. [6]

    Submission and peer review policies: Guidelines on AI-generated text

    Institute of Electrical and Electronics Engineers, “Submission and peer review policies: Guidelines on AI-generated text.” https://journals.ieeeauthorcenter.ieee.org/become- an-ieee-journal-author/publishing-ethics/guidelines- and-policies/submission-and-peer-review-policies/,

  7. [7]

    Acm policy on plagiarism, misrepresentation, and content fal- sification

    Association for Computing Machinery, “Acm policy on plagiarism, misrepresentation, and content fal- sification.” https://www.acm.org/publications/policies/ plagiarism-overview, 2025. ACM defines content fal- sification as intentional misrepresentation of results, supporting materials, or references, including manu- factured or irrelevant citations that do...

  8. [8]

    Ieee ethical requirements for authors: Avoiding fabrication, falsification, and citation misconduct

    Institute of Electrical and Electronics Engineers, “Ieee ethical requirements for authors: Avoiding fabrication, falsification, and citation misconduct.” https://journals.ieeeauthorcenter.ieee.org/become- an-ieee-journal-author/publishing-ethics/ethical- requirements/, 2025. IEEE author ethics policies prohibit fabrication (inventing data/results), falsif...

  9. [9]

    The ChatGPT lawyer explains himself

    B. Weiser and N. Schweber, “The ChatGPT lawyer explains himself.” https://www.nytimes.com/2023/06/ 08/nyregion/lawyer-chatgpt-sanctions.html, 6 2023. Accessed: 2026-01-14

  10. [10]

    GPTZero finds over 50 new hallucinations in ICLE2026 submis- sions

    P . Esau, N. Shmatko, and A. Adam, “GPTZero finds over 50 new hallucinations in ICLE2026 submis- sions.” https://gptzero.me/news/iclr-2026/, 2025. Ac- cessed: 2026-01-14

  11. [11]

    Gptzero: Detection of ai- generated text,

    E. Tian and A. Cui, “Gptzero: Detection of ai- generated text,” 2026. Accessed: 2026-01-15

  12. [12]

    Startup investigation reveals 50 peer-reviewed papers contained ai-hallucinated citations

    S. Martin, “Startup investigation reveals 50 peer-reviewed papers contained ai-hallucinated citations.” https://betakit.com/start-up-investigation- reveals-50-peer-reviewed-papers-contained- hallucinated-citations/?utm_source=chatgpt.com,

  13. [13]

    Accessed: 2026-01-15

    PyPDF Contributors, “Pypdf.” https://github.com/py- pdf/pypdf, 2026. Accessed: 2026-01-15

  14. [14]

    ERA conference rankings

    “ERA conference rankings.” http://www. conferenceranks.com. Accessed: 2026-01-21

  15. [15]

    From content creation to citation inflation: A genai case study

    H. S. Al-Sinani and C. J. Mitchell, “From content creation to citation inflation: A genai case study.” https://arxiv.org/abs/2503.23414, 2025. Accessed: 2026-01-16

  16. [16]

    How easy is it to fudge your scientific rank? meet larry, the world’s most cited cat

    C. Wilcox, “How easy is it to fudge your scientific rank? meet larry, the world’s most cited cat.” https: //www.doi.org/10.1126/science.zl99qni, 2024. Ac- cessed: 2026-01-16. Amanda Bienzis an Assistant Professor in the Com- puter Science Department at the University of New Mexico. Her research interests include improving the performance and scalability o...

  17. [2022]

    Accessed: 2026-01-13

  18. [2023]

    Accessed: 2026-01-13

    Defines disclosure requirements for AI- generated content in IEEE publications. Accessed: 2026-01-13

  19. [2025]

    Published December 16, 2025; accessed 2026-01-15