REVIEW 4 major objections 5 minor 40 references
ArxEval: Evaluating Retrieval and Generation in Language Models for Scientific Literature
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper reports that across fifteen open language models, every valid DOI generated in its Mixed Titles task was paired with a wrong title, and that only 17.75% of generated DOIs existed.
desk verdict Plausible qualitative signal, but the benchmark's central quantitative claim is unverified and the main results table has arithmetic errors. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the ArxEval pipeline, a two-task benchmark over 528 real titles from 176 preprint categories. The Jumbled Titles task shuffles the words of a title, asks the model to describe the paper, and scores the response against the true abstract with cosine similarity, BERTScore, and semantic textual similarity. The Mixed Titles task shuffles two titles into one string, asks for two paper titles plus DOIs, checks each DOI against bibliographic databases, and then compares the model's title for each valid DOI against the database title. The 100% title-mismatch finding is computed against exact title equality after DOI validation.
What would settle it
Take a random sample of the Mixed Titles responses, rescore the titles against the validated DOIs using fuzzy string similarity, and have human raters judge whether one-DOI or zero-DOI answers to a shuffled word salad are reasonable; a large jump in the title-match rate or many accepted non-compliant answers would show the reported 100% and 17.75% figures depend on exact matching and on the fixed two-paper expectation.
Extended reading notes
Core claim
The central claim is that current open language models systematically hallucinate when retrieving and generating scientific literature under input distortion. In the Jumbled Titles task, the best model averaged 0.585 similarity, and models clustered in a moderate band rather than recovering the papers. In the Mixed Titles task, none of the fifteen models produced the requested two correct title–DOI pairs; valid DOIs appeared only 17.75% of the time on average, and every valid DOI was associated with a title that did not match the database record. The paper presents these results as evidence that factual consistency in domain-specific scientific prompting is a critical unsolved problem.
Load-bearing premise
The Mixed Titles task assumes a word salad made from two titles has exactly one right answer—those two papers with two DOIs—so any other response is scored as hallucination; if models instead treat the question as ill-posed, the headline rates measure prompt compliance rather than hallucination.
Editorial extensions
If this is right
- Users cannot safely copy bibliographic references from these models without checking every DOI against a database.
- A real DOI is not proof of a reliable citation, because the accompanying title was wrong in every verified case.
- The two-task protocol gives a reusable way to compare future open models on scientific retrieval and generation fidelity.
- Larger parameter counts did not reliably improve performance, so scaling alone is unlikely to fix citation hallucination.
Reading between the lines
- The authors do not test fuzzy title matching; a rerun that allows paraphrases or near-matches could show how much of the 0% title accuracy is exact-match strictness rather than fabrication.
- Because the Mixed Titles prompt asks models to answer a word salad as though it names two papers, low output counts and partial answers may partly reflect a model treating the query as ill-posed, which would make the 17.75% figure a measure of instruction-following as well as hallucination.
- The same benchmark could be run with retrieval augmentation or web search available to the model, which would separate failures of stored knowledge from failures of generation and citation formatting.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents ArxEval, a pipeline for evaluating hallucination in language models on scientific literature. It introduces two tasks: Jumbled Titles, where a model is given a shuffled paper title and asked to describe the paper, and Mixed Titles, where a model is given a shuffled concatenation of two titles and asked to return exactly two papers with their DOIs. Fifteen open-weight models are evaluated; the paper reports aggregate similarity scores for Jumbled Titles and DOI validity/title-match rates for Mixed Titles. The two central quantitative claims are that valid DOIs are generated only 17.75% of the time and that for every valid DOI retrieved, the associated title is incorrect 100% of the time across all models.
Significance. If the quantitative claims held, the paper would provide useful evidence on reference hallucination in LLMs, and the comparison across fifteen models could inform practical choices about model size. The work has genuine strengths: it validates DOIs against external repositories (Crossref, DataCite, UnPaywall, OpenAlex), covers a broad set of models, and openly discusses limitations such as quantization. However, the load-bearing conclusions currently rest on an unspecified title-comparison procedure and on inconsistent dataset and result numbers. The qualitative signal, that models invent plausible references from scrambled inputs, is credible and consistent with prior work, but the paper's specific quantitative claims are not supported as written.
major comments (4)
- [Section 5, Table 5] Section 5 states that for every valid DOI retrieved, the associated title was incorrect 100% of the time, and Table 5 reports 0.00% Matching Titles for all models. The paper never specifies how the model-generated title was compared with the API-retrieved title: no normalization, no string-matching method, no fuzzy matching, and no manual review are described. Because model outputs are free-text responses to a prompt asking to 'only mention the Title and the DOI,' parsed titles are likely to differ from official titles in superficial formatting such as case, punctuation, whitespace, or LaTeX escapes. Without a specified matching protocol, the 0.00% rates are uninterpretable, and the claim in Section 6 that every model 'completely failed to generate the corresponding DOI for the title they generated' is unsupported.
- [Section 3 vs. Section 4.1, Table 3] Section 3 states that 176 categories were used with 3 titles from each category, giving 528 titles; Section 4.1, however, says 'We select 5 titles from each category.' These assumptions produce different dataset sizes. They also do not match the Mixed Titles count: pairing 528 titles yields 264 mixed titles, not the 265 reported in Table 3 and used in Section 6 (where 265 mixed titles imply an expected output of 530 DOIs). The inconsistency must be resolved because the expected-output denominator drives the validity-rate interpretation.
- [Table 5, Section 6] Table 5 contains internally inconsistent rows. For Orca-2 (7B), Total DOIs is 176 but DOIs Found (20) plus DOIs Not Found (86) equals 106, and the reported 18.87% is 20/106 rather than 20/176. The Llama-3 row lists 8 DOIs not found although 112 total minus 9 found is 103. The Section 6 statement that 'valid DOIs were generated only 17.75% of the time' does not match any aggregate calculable from Table 5: summing the rows gives 497 valid out of 2,254 total (22.05%), while the unweighted mean of the per-model validity percentages is about 19.6%. The authors should report the exact formula used and correct the table.
- [Section 4.2, Section 6] The Mixed Titles task assumes that a shuffled concatenation of two titles has a determinate correct answer: exactly two real papers and their two DOIs. The prompt 'Tell me 2 papers related to this' on a word-salad input is ill-posed; a model that responds with one relevant paper or states that the input is not a valid title is penalized as hallucinating. This conflates prompt compliance and question clarity with hallucination, and it affects the benchmark denominator (530 expected DOIs). The authors should justify the determinate-answer assumption or, at minimum, provide a breakdown of responses by number of DOIs and analyze partial responses separately.
minor comments (5)
- [Section 1] The introduction says fifteen models are evaluated but lists only eleven; Table 4 includes fifteen. Complete the list or correct the count.
- [Section 3, Table 3] The Gunning Fog scores in the text (17 for Jumbled Titles, 19 for Mixed Titles) differ from Table 3 (17 and 20). Align the text and the table.
- [Table 5] The Llama-3 row shows '8 [91.96%]' for DOIs Not Found; this should be 103 [91.96%] to be consistent with the total and found counts.
- [Algorithm 1, Section 4.1] Algorithm 1 does not show how categories are selected or how many titles are taken per category, despite Section 4.1 stating '5 titles from each category'; clarify the sampling procedure.
- [Table 5] Percentages for DOIs Not Found use inconsistent denominators across rows; for Orca-2 (7B) the not-found percentage is based on 106 rather than the reported total of 176. Use a single consistent denominator.
Circularity Check
No significant circularity: the pipeline evaluates against external ground truth (arXiv abstracts and Crossref/DataCite/UnPaywall/OpenAlex), and no fitted parameter or self-citation is repackaged as a prediction.
full rationale
ArxEval's load-bearing quantities come from external repositories, not from the paper's own definitions. The Jumbled Titles task compares generated responses to the original arXiv abstract using cosine similarity, BERTScore, and STS; the ground truth is the abstract itself, independent of the model outputs. The Mixed Titles task validates generated DOIs via Crossref, DataCite, UnPaywall, and OpenAlex, and then retrieves official titles from those databases for comparison. None of these steps defines the target result in terms of the inputs or fits a parameter to a subset of the data and calls the remainder a prediction. The paper contains no load-bearing self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation. The most notable weakness is an evaluation-validity concern rather than a circularity concern: the title-comparison step leading to the '0.00% Matching Titles' claim is not described with normalization or matching criteria, so the 100% title-mismatch result could be an artifact of exact-string comparison. However, that would be a measurement artifact, not circular reasoning, because the comparison target is still externally sourced. Similarly, the 'expected output of 530 DOIs' assumption reflects prompt-compliance scoring, not a definitional equivalence. Therefore the paper's derivation chain is self-contained with respect to external evidence, and the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (1)
- Random seed for title shuffling =
not reported
assumptions (3)
- domain assumption Semantic similarity between model output and the original abstract measures faithful retrieval of the paper from a jumbled title.
- domain assumption A mixed title has a determinate expected answer of exactly two real papers with two DOIs, and every other output is hallucination.
- domain assumption DOI validity and title matching through Crossref, DataCite, UnPaywall, and OpenAlex provide complete and reliable ground truth.
Cite this review
Pith. "Pith review of ArxEval: Evaluating Retrieval and Generation in Language Models for Scientific Literature." pith.science (2026). https://pith.science/paper/AS3BGQCP
@misc{pith2026250110483,
author = {Pith},
title = {Pith review of: ArxEval: Evaluating Retrieval and Generation in Language Models for Scientific Literature},
year = {2026},
howpublished = {\url{https://pith.science/paper/AS3BGQCP}},
note = {Machine review of arXiv:2501.10483}
}
read the original abstract
Language Models [LMs] are now playing an increasingly large role in information generation and synthesis; the representation of scientific knowledge in these systems needs to be highly accurate. A prime challenge is hallucination; that is, generating apparently plausible but actually false information, including invented citations and nonexistent research papers. This kind of inaccuracy is dangerous in all the domains that require high levels of factual correctness, such as academia and education. This work presents a pipeline for evaluating the frequency with which language models hallucinate in generating responses in the scientific literature. We propose ArxEval, an evaluation pipeline with two tasks using ArXiv as a repository: Jumbled Titles and Mixed Titles. Our evaluation includes fifteen widely used language models and provides comparative insights into their reliability in handling scientific literature.
Figures
Reference graph
Works this paper leans on
-
[1]
A comprehensive overview of large language models, 2024
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. A comprehensive overview of large language models, 2024
work page 2024
-
[2]
Large language models: A survey, 2024
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. Large language models: A survey, 2024
2024
-
[3]
Evaluating large language models: A comprehensive survey, 2023
Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Supryadi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, and Deyi Xiong. Evaluating large language models: A comprehensive survey, 2023
2023
-
[4]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, November 2024
work page 2024
-
[5]
A comprehensive survey of hallucination in large language, image, video and audio foundation models, 2024
Pranab Sahoo, Prabhash Meharia, Akash Ghosh, Sriparna Saha, Vinija Jain, and Aman Chadha. A comprehensive survey of hallucination in large language, image, video and audio foundation models, 2024
2024
-
[6]
M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das
S. M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. A comprehensive survey of hallucination mitigation techniques in large language models, 2024
2024
-
[7]
Vipula Rawte, Swagata Chakraborty, Agnibh Pathak, Anubhav Sarkar, S.M Towhidul Islam Tonmoy, Aman Chadha, Amit Sheth, and Amitava Das. The troubling emergence of hallucination in large language models - an extensive definition, quantification, and prescriptive remediations. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Con...
work page 2023
-
[8]
Elijah Berberette, Jack Hutchins, and Amir Sadovnik. Redefining "hallucination" in llms: Towards a psychology- informed framework for mitigating misinformation, 2024
work page 2024
Show all 40 references
-
[9]
Qwen2 technical report
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng...
2024 arXiv
-
[10]
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan...
2024
-
[11]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024
-
[12]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Wei...
2024
-
[13]
Orca 2: Teaching small language models how to reason, 2023
Arindam Mitra, Luciano Del Corro, Shweti Mahajan, Andres Codas, Clarisse Simoes, Sahaj Agarwal, Xuxi Chen, Anastasia Razdaibiedina, Erik Jones, Kriti Aggarwal, Hamid Palangi, Guoqing Zheng, Corby Rosset, Hamed Khanpour, and Ahmed Awadallah. Orca 2: Teaching small language mode...
2023
-
[14]
mistralai/mistral-7b-instruct-v0.3 · hugging face, December 2024
Mistral AI Team. mistralai/mistral-7b-instruct-v0.3 · hugging face, December 2024. [Online; accessed 2024-12- 16]
2024
-
[15]
DeepSeek-AI, :, Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, Huazuo Gao, Kaige Gao, Wenjun Gao, Ruiqi Ge, Kang Guan, Daya Guo, Jianzhong Guo, Guangbo Hao, Zhewen Hao, Ying He, Wenjie Hu, Panpan Huang, Er...
2024
-
[16]
Team OLMo, Pete Walsh, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Shane Arora, Akshita Bhagia, Yuling Gu, Shengyi Huang, Matt Jordan, Nathan Lambert, Dustin Schwenk, Oyvind Tafjord, Taira Anderson, David Atkinson, Faeze Brahman, Christopher Clark, Pradeep Dasigi, Nouha Dziri, Mi...
2024
-
[17]
mistralai/mistral-nemo-instruct-2407 · hugging face
Mistral Team. mistralai/mistral-nemo-instruct-2407 · hugging face. [Online; accessed 2025-01-15]
2025
-
[18]
Free process rewards without process labels
Lifan Yuan, Wendi Li, Huayu Chen, Ganqu Cui, Ning Ding, Kaiyan Zhang, Bowen Zhou, Zhiyuan Liu, and Hao Peng. Free process rewards without process labels. arXiv preprint arXiv:2412.01981, 2024
2024 arXiv
-
[19]
upstage/solar-pro-preview-instruct · hugging face, 11 2024
upstage. upstage/solar-pro-preview-instruct · hugging face, 11 2024. [Online; accessed 2025-01-15]
2024
-
[20]
Clement, Matthew Bierbaum, Kevin P
Colin B. Clement, Matthew Bierbaum, Kevin P. O’Keeffe, and Alexander A. Alemi. On the use of arxiv as a dataset, 2019
2019
-
[21]
On large language models’ hallucination with regard to known facts
Che Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Yang Cheng, Fandong Meng, Mo Yu, Bowen Zhou, and Jie Zhou. On large language models’ hallucination with regard to known facts. arXiv preprint arXiv:2403.20009, 2024
2024 arXiv
-
[22]
Llms will always hallucinate, and we need to live with this
Sourav Banerjee, Ayushi Agarwal, and Saloni Singla. Llms will always hallucinate, and we need to live with this. arXiv preprint arXiv:2409.05746, 2024
2024 arXiv
-
[23]
Nouha Dziri, Sivan Milton, Mo Yu, Osmar Zaiane, and Siva Reddy. On the origin of hallucinations in con- versational models: Is it the datasets or the models? In Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz, editors, Proceedings of the 2022 Conferenc...
2022
-
[24]
Confabulation: The surprising value of large language model hallucinations
Peiqi Sui, Eamon Duede, Sophie Wu, and Richard Jean So. Confabulation: The surprising value of large language model hallucinations. arXiv preprint arXiv:2406.04175, 2024
2024 arXiv
-
[25]
Delucionqa: Detecting hallucinations in domain-specific question answering, 2023
Mobashir Sadat, Zhengyu Zhou, Lukas Lange, Jun Araki, Arsalan Gundroo, Bingqing Wang, Rakesh R Menon, Md Rizwan Parvez, and Zhe Feng. Delucionqa: Detecting hallucinations in domain-specific question answering, 2023
2023
-
[26]
Dahl: Domain-specific automated hallucination evaluation of long-form text through a benchmark dataset in biomedicine, 2024
Jean Seo, Jongwon Lim, Dongjun Jang, and Hyopil Shin. Dahl: Domain-specific automated hallucination evaluation of long-form text through a benchmark dataset in biomedicine, 2024
2024
-
[27]
Hallucina- tion of multimodal large language models: A survey, 2024
Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, and Mike Zheng Shou. Hallucina- tion of multimodal large language models: A survey, 2024
2024
-
[28]
Vidhalluc: Evaluating temporal hallucinations in multimodal large language models for video understanding, 2024
Chaoyu Li, Eun Woo Im, and Pooyan Fazli. Vidhalluc: Evaluating temporal hallucinations in multimodal large language models for video understanding, 2024
2024
-
[29]
Unified hallucination detection for multimodal large language models
Xiang Chen, Chenxi Wang, Yida Xue, Ningyu Zhang, Xiaoyan Yang, Qiang Li, Yue Shen, Lei Liang, Jinjie Gu, and Huajun Chen. Unified hallucination detection for multimodal large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Proceedings of the 62nd An...
2024
-
[30]
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, March 2023
2023
-
[31]
Ayush Agrawal, Mirac Suzgun, Lester Mackey, and Adam Kalai. Do language models know when they’re hallucinating references? In Yvette Graham and Matthew Purver, editors, Findings of the Association for Computational Linguistics: EACL 2024 , pages 912–928, St. Julian’s, Malta, M...
2024
-
[32]
Flesch–kincaid readability tests - wikipedia, 7 2004
Contributors to Wikimedia projects. Flesch–kincaid readability tests - wikipedia, 7 2004. [Online; accessed 2025-01-17]
2004
-
[33]
Gunning fog index - wikipedia, 10 2004
Contributors to Wikimedia projects. Gunning fog index - wikipedia, 10 2004. [Online; accessed 2025-01-17]
2004
-
[34]
Sentence-bert: Sentence embeddings using siamese bert-networks, 2019
Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks, 2019
2019
-
[35]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert, 2020
2020
-
[36]
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fan...
2019
-
[37]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[38]
bitsandbytes-foundation/bitsandbytes: Accessible large language models via k-bit quantization for pytorch., Decemeber 2024
bitsandbytes foundation. bitsandbytes-foundation/bitsandbytes: Accessible large language models via k-bit quantization for pytorch., Decemeber 2024. [Online; accessed 2024-12-16]
2024
-
[39]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023
-
[40]
A comprehensive evaluation of quantization strategies for large language models, 2024
Renren Jin, Jiangcun Du, Wuwei Huang, Wei Liu, Jian Luan, Bin Wang, and Deyi Xiong. A comprehensive evaluation of quantization strategies for large language models, 2024. 13
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.