Pith. sign in

REVIEW 1 cited by

Language Models Hallucinate, but May Excel at Fact Verification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.14564 v2 pith:72RXMWJE submitted 2023-10-23 cs.CL

classification cs.CL
keywords llmsfactlanguagemodelsevenfactualhallucinatehuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent progress in natural language processing (NLP) owes much to remarkable advances in large language models (LLMs). Nevertheless, LLMs frequently "hallucinate," resulting in non-factual outputs. Our carefully-designed human evaluation substantiates the serious hallucination issue, revealing that even GPT-3.5 produces factual outputs less than 25% of the time. This underscores the importance of fact verifiers in order to measure and incentivize progress. Our systematic investigation affirms that LLMs can be repurposed as effective fact verifiers with strong correlations with human judgments. Surprisingly, FLAN-T5-11B, the least factual generator in our study, performs the best as a fact verifier, even outperforming more capable LLMs like GPT3.5 and ChatGPT. Delving deeper, we analyze the reliance of these LLMs on high-quality evidence, as well as their deficiencies in robustness and generalization ability. Our study presents insights for developing trustworthy generation models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mobilizing Waldo: Evaluating Multimodal AI for Public Mobilization

    cs.HC 2024-12 conditional novelty 4.0 of 10

    GPT-4o can describe and invent creative pitches for Where's Waldo scenes but cannot reliably locate Waldo or identify individual characters in dense illustrations.

Pith tools