Pith. sign in

REVIEW 4 cited by

QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.08542 v2 pith:MHCGHQCY submitted 2021-12-16 cs.CL

classification cs.CL
keywords qa-basedmetricmetricsconsistencyentailment-basedfactualperformanceqafacteval
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Factual consistency is an essential quality of text summarization models in practical settings. Existing work in evaluating this dimension can be broadly categorized into two lines of research, entailment-based and question answering (QA)-based metrics, and different experimental setups often lead to contrasting conclusions as to which paradigm performs the best. In this work, we conduct an extensive comparison of entailment and QA-based metrics, demonstrating that carefully choosing the components of a QA-based metric, especially question generation and answerability classification, is critical to performance. Building on those insights, we propose an optimized metric, which we call QAFactEval, that leads to a 14% average improvement over previous QA-based metrics on the SummaC factual consistency benchmark, and also outperforms the best-performing entailment-based metric. Moreover, we find that QA-based and entailment-based metrics can offer complementary signals and be combined into a single metric for a further performance boost.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new benchmark built with executable phrase-level edits shows that most LLMs detect and explain factual inconsistencies in summaries only weakly, with the best model scoring 0.49 on the joint task.

  2. LLM-as-a-Verifier: A General-Purpose Verification Framework

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Expecting over scoring-token logits yields continuous, scalable verification that improves agent trajectory selection and dense RL rewards across coding, robotics, and medical benchmarks.

  3. Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly

    cs.CV 2024-12 conditional novelty 4.0 of 10

    ECVA is a benchmark of 2,240 long real-world videos with human-written cause, effect, and severity annotations, plus a baseline VLM and a GPT-based metric.

  4. Human-Calibrated Automated Testing and Validation of Generative Language Models

    cs.CL 2024-11 conditional novelty 4.0 of 10

    The paper proposes HCAT, a human-calibrated evaluation framework for RAG models, combining automated tests, embedding metrics, and conformal prediction.

Pith tools