ArxEval reports that 15 small language models hallucinate frequently on jumbled and mixed arXiv titles, but internal data errors and missing baselines prevent the quantitative rankings from being trusted.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
ArxEval: Evaluating Retrieval and Generation in Language Models for Scientific Literature
ArxEval reports that 15 small language models hallucinate frequently on jumbled and mixed arXiv titles, but internal data errors and missing baselines prevent the quantitative rankings from being trusted.