Pith. sign in

REVIEW 3 cited by

Entity-level Factual Consistency of Abstractive Text Summarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.09130 v1 pith:PCFIT7NM submitted 2021-02-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords entityconsistencyfactualabstractivedocumententity-levelgeneratedhallucination
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A key challenge for abstractive summarization is ensuring factual consistency of the generated summary with respect to the original document. For example, state-of-the-art models trained on existing datasets exhibit entity hallucination, generating names of entities that are not present in the source document. We propose a set of new metrics to quantify the entity-level factual consistency of generated summaries and we show that the entity hallucination problem can be alleviated by simply filtering the training data. In addition, we propose a summary-worthy entity classification task to the training process as well as a joint entity and summary generation approach, which yield further improvements in entity level metrics.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A single linear direction in an observer model's residual stream detects contextual hallucinations, transfers across models and datasets, and causally steers generation hallucination rates.

  2. A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Introduces a multi-role red teaming framework using attacker and jury models that increases attack success rates by up to 7.9% on LLM faithfulness in question-answering tasks.

  3. Improving Factuality for Dialogue Response Generation via Graph-Based Knowledge Augmentation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The paper proposes TG-DRG and GA-DRG, two graph-augmented frameworks that combine coreference resolution, knowledge selection, and graph encoding to improve factuality of dialogue responses, evaluated with a newly pro...

Pith tools