REVIEW 5 cited by
MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A robust evaluation metric has a profound impact on the development of text generation systems. A desirable metric compares system output against references based on their semantics rather than surface forms. In this paper we investigate strategies to encode system and reference texts to devise a metric that shows a high correlation with human judgment of text quality. We validate our new metric, namely MoverScore, on a number of text generation tasks including summarization, machine translation, image captioning, and data-to-text generation, where the outputs are produced by a variety of neural and non-neural systems. Our findings suggest that metrics combining contextualized representations with a distance measure perform the best. Such metrics also demonstrate strong generalization capability across tasks. For ease-of-use we make our metrics available as web service.
Forward citations
Cited by 5 Pith papers
-
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization
Across eight languages, n-gram metrics such as ROUGE correlate less with human ratings in fusional languages than in isolating and agglutinative ones, while the neural metric COMET correlates better, especially in low...
-
AllSummedUp: un framework open-source pour comparer les metriques d'evaluation de resume
On SummEval, LLM-based summary evaluators are expensive and unstable, and several published correlations did not reproduce when using open-weight models.
-
Statistical Hypothesis Testing for Auditing Robustness in Language Models
A permutation-based hypothesis test on pairwise semantic similarities detects whether LLM outputs shift under arbitrary input or model perturbations.
-
Objective Metrics for Evaluating Large Language Models Using External Data Sources
The submission is unverifiable because its abstract and full text describe two entirely different papers.
-
BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining
A proposed Persian biomedical LLM, BioPars, is evaluated on medical QA datasets and reported to beat GPT-4 on a self-built Persian QA benchmark, but the training setup is not described.
Discussion (0). Sign in to comment.