Pith. sign in

REVIEW 13 cited by

Evaluating the Factual Consistency of Abstractive Text Summarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.12840 v1 pith:LHQQIGUL submitted 2019-10-28 cs.CL

Evaluating the Factual Consistency of Abstractive Text Summarization

classification cs.CL
keywords consistencydocumentsfactualsourcegeneratedspanapproachconsistent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Currently used metrics for assessing summarization algorithms do not account for whether summaries are factually consistent with source documents. We propose a weakly-supervised, model-based approach for verifying factual consistency and identifying conflicts between source documents and a generated summary. Training data is generated by applying a series of rule-based transformations to the sentences of source documents. The factual consistency model is then trained jointly for three tasks: 1) identify whether sentences remain factually consistent after transformation, 2) extract a span in the source documents to support the consistency prediction, 3) extract a span in the summary sentence that is inconsistent if one exists. Transferring this model to summaries generated by several state-of-the art models reveals that this highly scalable approach substantially outperforms previous models, including those trained with strong supervision using standard datasets for natural language inference and fact checking. Additionally, human evaluation shows that the auxiliary span extraction tasks provide useful assistance in the process of verifying factual consistency.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Layer-Resolved Optimal Transport for Hallucination Detection in NMT and Abstractive Summarization

    cs.CL 2026-06 unverdicted novelty 6.0

    Layer-resolved OT detects source-disengagement hallucinations in NMT but achieves only 57% balanced accuracy on summarization because content misrepresentation can occur with correct attention.

  2. Detect, Remask, Repair: Diffusion Editing for Faithful Summarization of Evolving Contexts

    cs.CL 2026-06 unverdicted novelty 6.0

    Diffusion-based localized editing framework for faithful summarization of evolving contexts, introducing the StreamSum benchmark and showing tradeoffs in faithfulness, speed, and preservation.

  3. ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop

    cs.AI 2026-06 conditional novelty 6.0

    ProcAgent demonstrates a complete step-by-step assembly assistant that runs on a single edge device by pairing cheap continuous perception with on-demand vision-language verification, a task graph, and human confirmation.

  4. Constrained Paraphrase Consistency for LLM Hallucination Detection

    cs.CL 2026-06 unverdicted novelty 6.0

    CCHD formulates hallucination detector training as constrained optimization with paraphrase-consistency and label-preservation rules solved via gradient descent-ascent, outperforming baselines on factuality benchmarks.

  5. Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives

    cs.CL 2026-04 unverdicted novelty 6.0

    A proposed pipeline shows LLMs introduce detectable race and gender biases when summarizing life narratives, creating potential for representational harm in research.

  6. No-Worse Context-Aware Decoding: Preventing Neutral Regression in Context-Conditioned Generation

    cs.CL 2026-04 unverdicted novelty 6.0

    NWCAD uses a two-stream setup with a two-stage gate to prevent accuracy drops on baseline-correct items under non-informative contexts while retaining gains from helpful contexts.

  7. A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

    cs.CL 2026-06 conditional novelty 5.0

    A multi-role red-teaming framework with attacker, target, and jury LLMs measures faithfulness in English and Arabic, finding false-premise prompts and length limits change unfaithfulness rates.

  8. Cross Paraphrastic Invariance Learning for Hallucination Detection

    cs.CL 2026-06 unverdicted novelty 5.0

    CPIL is a contrastive two-stage method that enforces paraphrase invariance on limited labeled data to outperform baselines in hallucination detection across 11 tasks.

  9. HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs

    cs.CL 2026-05 unverdicted novelty 5.0

    HalluScan benchmark tests hallucination detectors on LLMs, identifies NLI Verification as top performer with 0.88 AUROC, and introduces HalluScore (r=0.41 with humans) plus a routing method for 2x cost savings.

  10. A Stepwise Questioning Expert-Editor Multi-Agent Framework for Long-Document Summarization

    cs.CL 2026-07 conditional novelty 4.0

    An expert-editor stepwise-questioning multi-agent pipeline improves ROUGE/BERTScore/FactCC for long scientific summarization on two datasets relative to direct generation and HERA.

  11. A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

    cs.CL 2026-06 unverdicted novelty 4.0

    Introduces a multi-role red teaming framework using attacker and jury models that increases attack success rates by up to 7.9% on LLM faithfulness in question-answering tasks.

  12. HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs

    cs.CL 2026-05 unverdicted novelty 4.0

    HalluScan benchmark evaluates hallucination detection in LLMs, reporting NLI Verification at AUROC 0.88 and introducing HalluScore (r=0.41 with humans) plus Adaptive Detection Routing for 2x cost savings.

  13. A Community-Based Approach for Stance Distribution and Argument Organization

    cs.CL 2026-04 unverdicted novelty 4.0

    Unsupervised graph community detection organizes arguments to reveal stance distributions in debates.