Pith. sign in

REVIEW 1 cited by

TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13249 v2 pith:6MVAWZAO submitted 2024-02-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords factualllmssummarizationdialogueerrorevaluationanalysisbinary
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Single document news summarization has seen substantial progress on faithfulness in recent years, driven by research on the evaluation of factual consistency, or hallucinations. We ask whether these advances carry over to other text summarization domains. We propose a new evaluation benchmark on topic-focused dialogue summarization, generated by LLMs of varying sizes. We provide binary sentence-level human annotations of the factual consistency of these summaries along with detailed explanations of factually inconsistent sentences. Our analysis shows that existing LLMs hallucinate significant amounts of factual errors in the dialogue domain, regardless of the model's size. On the other hand, when LLMs, including GPT-4, serve as binary factual evaluators, they perform poorly and can be outperformed by prevailing state-of-the-art specialized factuality evaluation metrics. Finally, we conducted an analysis of hallucination types with a curated error taxonomy. We find that there are diverse errors and error distributions in model-generated summaries and that non-LLM based metrics can capture all error types better than LLM-based evaluators.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Party Conversational Agents: A Survey

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A survey of multi-party conversational AI that organizes tasks into state-of-mind modeling, semantic understanding, and action modeling, and argues that theory of mind is the key missing ingredient.

Pith tools