Pith. sign in

REVIEW

Calibrating Likelihoods towards Consistency in Summarization Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08764 v1 pith:A467ZGKE submitted 2023-10-12 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelssummarizationconsistencysequencesbettercalibratinglikelihoodsummaries
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the recent advances in abstractive text summarization, current summarization models still suffer from generating factually inconsistent summaries, reducing their utility for real-world application. We argue that the main reason for such behavior is that the summarization models trained with maximum likelihood objective assign high probability to plausible sequences given the context, but they often do not accurately rank sequences by their consistency. In this work, we solve this problem by calibrating the likelihood of model generated sequences to better align with a consistency metric measured by natural language inference (NLI) models. The human evaluation study and automatic metrics show that the calibrated models generate more consistent and higher-quality summaries. We also show that the models trained using our method return probabilities that are better aligned with the NLI scores, which significantly increase reliability of summarization models.

Discussion (0). Continue with ORCID to comment.

Pith tools