Pith. sign in

REVIEW 1 cited by

On the Evaluation Metrics for Paraphrase Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.08479 v2 pith:ZFMFKQQL submitted 2022-02-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords metricsevaluationparaphraseanalysesexperimentsfindingsgenerationparascore
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we revisit automatic metrics for paraphrase evaluation and obtain two findings that disobey conventional wisdom: (1) Reference-free metrics achieve better performance than their reference-based counterparts. (2) Most commonly used metrics do not align well with human annotation. Underlying reasons behind the above findings are explored through additional experiments and in-depth analyses. Based on the experiments and analyses, we propose ParaScore, a new evaluation metric for paraphrase generation. It possesses the merits of reference-based and reference-free metrics and explicitly models lexical divergence. Experimental results demonstrate that ParaScore significantly outperforms existing metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Style Extraction on Text Embeddings Using VAE and Parallel Dataset

    cs.CL 2025-02 reject novelty 3.0 of 10

    A VAE (actually an autoencoder) trained on KJV minus ASV embedding differences separates ASV from other Bible translations by reconstruction error, with an unstable claimed accuracy around 84 percent.

Pith tools