REVIEW 3 cited by
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While hallucinations of large language models (LLMs) prevail as a major challenge, existing evaluation benchmarks on factuality do not cover the diverse domains of knowledge that the real-world users of LLMs seek information about. To bridge this gap, we introduce WildHallucinations, a benchmark that evaluates factuality. It does so by prompting LLMs to generate information about entities mined from user-chatbot conversations in the wild. These generations are then automatically fact-checked against a systematically curated knowledge source collected from web search. Notably, half of these real-world entities do not have associated Wikipedia pages. We evaluate 118,785 generations from 15 LLMs on 7,919 entities. We find that LLMs consistently hallucinate more on entities without Wikipedia pages and exhibit varying hallucination rates across different domains. Finally, given the same base models, adding a retrieval component only slightly reduces hallucinations but does not eliminate hallucinations.
Forward citations
Cited by 3 Pith papers
-
U-Lens: Supporting User Uncertainty Management in Long-Form LLM Responses
U-Lens organizes long-form LLM uncertainty into prioritized multi-granular targets with evaluative explanations and response guidance, improving limited-budget verification over a confidence-cue baseline.
-
Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models
Hallucination in LLMs is driven by an oracle-invisible “decoding risk” term that grows with scale and causally compounds errors within a response.
-
A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs
Metalinguistic disagreements, where the dispute is over word meaning rather than facts, appear in LLM fact-checking against knowledge graphs, based on a 250-triple pilot study.
Discussion (0). Continue with ORCID to comment.