Pith. sign in

REVIEW 1 cited by

Towards Question-Answering as an Automatic Metric for Evaluating the Content Quality of a Summary

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.00490 v3 pith:WZ2422MT submitted 2020-10-01 cs.CL

classification cs.CL
keywords metricssummarymetriccontentoverlapqaevalqualityanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A desirable property of a reference-based evaluation metric that measures the content quality of a summary is that it should estimate how much information that summary has in common with a reference. Traditional text overlap based metrics such as ROUGE fail to achieve this because they are limited to matching tokens, either lexically or via embeddings. In this work, we propose a metric to evaluate the content quality of a summary using question-answering (QA). QA-based methods directly measure a summary's information overlap with a reference, making them fundamentally different than text overlap metrics. We demonstrate the experimental benefits of QA-based metrics through an analysis of our proposed metric, QAEval. QAEval out-performs current state-of-the-art metrics on most evaluations using benchmark datasets, while being competitive on others due to limitations of state-of-the-art models. Through a careful analysis of each component of QAEval, we identify its performance bottlenecks and estimate that its potential upper-bound performance surpasses all other automatic metrics, approaching that of the gold-standard Pyramid Method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Using maps and metadata as extra context for GPT-4o caption generation yields a richer remote sensing dataset, fMoW-mm, with claimed lower hallucination rates and better few-shot detection than prior datasets.

Pith tools