RelationalFactQA shows that LLMs are much worse at retrieving facts as multi-record tables than as single answers, with the best model reaching only 24.7% tuple similarity.
Ai hallucination report 2025, 2025
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
RelationalFactQA: A Benchmark for Evaluating Tabular Fact Retrieval from Large Language Models
RelationalFactQA shows that LLMs are much worse at retrieving facts as multi-record tables than as single answers, with the best model reaching only 24.7% tuple similarity.