ArxEval reports that 15 small language models hallucinate frequently on jumbled and mixed arXiv titles, but internal data errors and missing baselines prevent the quantitative rankings from being trusted.
A comprehensive overview of large language models, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
ArxEval: Evaluating Retrieval and Generation in Language Models for Scientific Literature
ArxEval reports that 15 small language models hallucinate frequently on jumbled and mixed arXiv titles, but internal data errors and missing baselines prevent the quantitative rankings from being trusted.