Pith. sign in

REVIEW 1 cited by

Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.11648 v2 pith:IF7N7FI4 submitted 2023-10-18 cs.CL

classification cs.CL
keywords faithfulnessfflmlanguagemodelchatgptevaluationfoundationimprovements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite tremendous improvements in natural language generation, summarization models still suffer from the unfaithfulness issue. Previous work evaluates faithfulness either using models trained on the other tasks or in-domain synthetic data, or prompting a large model such as ChatGPT. This paper proposes to do zero-shot faithfulness evaluation simply with a moderately-sized foundation language model. We introduce a new metric FFLM, which is a combination of probability changes based on the intuition that prefixing a piece of text that is consistent with the output will increase the probability of predicting the output. Experiments show that FFLM performs competitively with or even outperforms ChatGPT on both inconsistency detection and faithfulness rating with 24x fewer parameters. FFLM also achieves improvements over other strong baselines.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A hybrid PubMed-KG RAG system improves NLI faithfulness of multi-stakeholder clinical-trial summaries by ~0.013 (p<0.0001) across GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash, while Unsupported Claims remains the d...

Pith tools