Frontier LLMs generate BibTeX entries at 83.6% field accuracy but only 50.9% fully correct; two-stage clibib revision raises accuracy to 91.5% and fully correct entries to 78.3% with 0.8% regression.
Hallucitation matters: Revealing the impact of hallu- cinated references with 300 hallucinated papers in acl conferences
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 9roles
background 2representative citing papers
MBR decoding is reformulated via a noisy-channel decomposition into four weighted probabilistic terms, revealing that channel importance is metric-specific and task-agnostic, and that reweighting can improve performance.
Under a strict identity-level definition, roughly one in twenty 2025 NeurIPS and USENIX Security papers carries at least two likely hallucinated academic citations that survived peer review.
A new multi-domain benchmark shows scientifically fine-tuned LLMs have degraded factual reliability and are less confident yet more assertive than their base models.
CiteTracer detects hallucinated citations with a 12-code field-level taxonomy and a cascading multi-agent retrieval pipeline, reporting 97.1% accuracy on a synthetic benchmark.
CiteAudit supplies a human-validated benchmark and multi-agent verification system that outperforms existing LLMs and commercial tools at detecting hallucinated scientific references.
LLMs hallucinate citations at rates from 14.23% to 94.93%, with 1.07% of papers containing invalid citations and an 80.9% increase in 2025.
AI lowers the cost of generating plausible scientific artifacts without lowering verification costs, so the paper proposes blueprints as typed graph components that decompose claims, evidence, and assumptions to enable cheaper downstream verification.
HalluCiteChecker is a lightweight, offline, CPU-only toolkit that detects hallucinated citations in AI-assisted scientific papers.
citing papers explorer
-
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
Frontier LLMs generate BibTeX entries at 83.6% field accuracy but only 50.9% fully correct; two-stage clibib revision raises accuracy to 91.5% and fully correct entries to 78.3% with 0.8% regression.
-
Noisy-Channel Minimum Bayes Risk Decoding
MBR decoding is reformulated via a noisy-channel decomposition into four weighted probabilistic terms, revealing that channel importance is metric-specific and task-agnostic, and that reweighting can improve performance.
-
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
Under a strict identity-level definition, roughly one in twenty 2025 NeurIPS and USENIX Security papers carries at least two likely hallucinated academic citations that survived peer review.
-
Finetuning with Scientific Data Increases Hallucinations: A Multi-domain Factuality Evaluation of LLMs
A new multi-domain benchmark shows scientifically fine-tuned LLMs have degraded factual reliability and are less confident yet more assertive than their base models.
-
Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection
CiteTracer detects hallucinated citations with a 12-code field-level taxonomy and a cascading multi-agent retrieval pipeline, reporting 97.1% accuracy on a synthetic benchmark.
-
CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
CiteAudit supplies a human-validated benchmark and multi-agent verification system that outperforms existing LLMs and commercial tools at detecting hallucinated scientific references.
-
GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models
LLMs hallucinate citations at rates from 14.23% to 94.93%, with 1.07% of papers containing invalid citations and an 80.9% increase in 2025.
-
Toward an Engineering of Science: Rebalancing Generation and Verification in the Age of AI
AI lowers the cost of generating plausible scientific artifacts without lowering verification costs, so the paper proposes blueprints as typed graph components that decompose claims, evidence, and assumptions to enable cheaper downstream verification.
-
HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
HalluCiteChecker is a lightweight, offline, CPU-only toolkit that detects hallucinated citations in AI-assisted scientific papers.