Pith. sign in

REVIEW 1 cited by

Metric Ensembles For Hallucination Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10495 v1 pith:F6JAZLTT submitted 2023-10-16 cs.CL

classification cs.CL
keywords metricsensemblehallucinationabstractiveconsistencydetectionevaluationsfind
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Abstractive text summarization has garnered increased interest as of late, in part due to the proliferation of large language models (LLMs). One of the most pressing problems related to generation of abstractive summaries is the need to reduce "hallucinations," information that was not included in the document being summarized, and which may be wholly incorrect. Due to this need, a wide array of metrics estimating consistency with the text being summarized have been proposed. We examine in particular a suite of unsupervised metrics for summary consistency, and measure their correlations with each other and with human evaluation scores in the wiki_bio_gpt3_hallucination dataset. We then compare these evaluations to models made from a simple linear ensemble of these metrics. We find that LLM-based methods outperform other unsupervised metrics for hallucination detection. We also find that ensemble methods can improve these scores even further, provided that the metrics in the ensemble have sufficiently similar and uncorrelated error rates. Finally, we present an ensemble method for LLM-based evaluations that we show improves over this previous SOTA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transparent NLP: Using RAG and LLM Alignment for Privacy Q&A

    cs.CL 2025-02 conditional novelty 5.0 of 10

    RAG systems with RAIN or MultiRAIN alignment outperform vanilla RAG on most privacy Q&A evaluation metrics, but none reach human expert quality and the approach is not yet practical.

Pith tools