REVIEW 4 cited by
QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Factual consistency is an essential quality of text summarization models in practical settings. Existing work in evaluating this dimension can be broadly categorized into two lines of research, entailment-based and question answering (QA)-based metrics, and different experimental setups often lead to contrasting conclusions as to which paradigm performs the best. In this work, we conduct an extensive comparison of entailment and QA-based metrics, demonstrating that carefully choosing the components of a QA-based metric, especially question generation and answerability classification, is critical to performance. Building on those insights, we propose an optimized metric, which we call QAFactEval, that leads to a 14% average improvement over previous QA-based metrics on the SummaC factual consistency benchmark, and also outperforms the best-performing entailment-based metric. Moreover, we find that QA-based and entailment-based metrics can offer complementary signals and be combined into a single metric for a further performance boost.
Forward citations
Cited by 4 Pith papers
-
SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits
A new benchmark built with executable phrase-level edits shows that most LLMs detect and explain factual inconsistencies in summaries only weakly, with the best model scoring 0.49 on the joint task.
-
LLM-as-a-Verifier: A General-Purpose Verification Framework
Expecting over scoring-token logits yields continuous, scalable verification that improves agent trajectory selection and dense RL rewards across coding, robotics, and medical benchmarks.
-
Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly
ECVA is a benchmark of 2,240 long real-world videos with human-written cause, effect, and severity annotations, plus a baseline VLM and a GPT-based metric.
-
Human-Calibrated Automated Testing and Validation of Generative Language Models
The paper proposes HCAT, a human-calibrated evaluation framework for RAG models, combining automated tests, embedding metrics, and conformal prediction.
Discussion (0). Continue with ORCID to comment.