REVIEW 1 cited by
On the Evaluation Metrics for Paraphrase Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper we revisit automatic metrics for paraphrase evaluation and obtain two findings that disobey conventional wisdom: (1) Reference-free metrics achieve better performance than their reference-based counterparts. (2) Most commonly used metrics do not align well with human annotation. Underlying reasons behind the above findings are explored through additional experiments and in-depth analyses. Based on the experiments and analyses, we propose ParaScore, a new evaluation metric for paraphrase generation. It possesses the merits of reference-based and reference-free metrics and explicitly models lexical divergence. Experimental results demonstrate that ParaScore significantly outperforms existing metrics.
Forward citations
Cited by 1 Pith paper
-
Style Extraction on Text Embeddings Using VAE and Parallel Dataset
A VAE (actually an autoencoder) trained on KJV minus ASV embedding differences separates ASV from other Bible translations by reconstruction error, with an unstable claimed accuracy around 84 percent.
Discussion (0). Continue with ORCID to comment.