Pith. sign in

REVIEW 10 cited by

COMET: A Neural Framework for MT Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.09025 v2 pith:CVWKL3KC submitted 2020-09-18 cs.CL

COMET: A Neural Framework for MT Evaluation

classification cs.CL
keywords frameworkmodelsevaluationtranslationcomethumanjudgementsmetrics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present COMET, a neural framework for training multilingual machine translation evaluation models which obtains new state-of-the-art levels of correlation with human judgements. Our framework leverages recent breakthroughs in cross-lingual pretrained language modeling resulting in highly multilingual and adaptable MT evaluation models that exploit information from both the source input and a target-language reference translation in order to more accurately predict MT quality. To showcase our framework, we train three models with different types of human judgements: Direct Assessments, Human-mediated Translation Edit Rate and Multidimensional Quality Metrics. Our models achieve new state-of-the-art performance on the WMT 2019 Metrics shared task and demonstrate robustness to high-performing systems.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

    eess.AS 2026-06 unverdicted novelty 7.0

    GigaSpeechBench is a new 680-hour in-the-wild multilingual ASR/AST benchmark with five modules for low-resource languages, Chinese dialects, English accents, domain terminology, and age-varied speech, showing model pe...

  2. GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

    eess.AS 2026-06 conditional novelty 7.0

    GigaSpeechBench, a new 680-hour human-annotated in-the-wild ASR/AST benchmark, shows current ASR systems degrade sharply on low-resource languages, dialects, accents, terminology-dense domains, and child/older-adult speech.

  3. Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models

    cs.CL 2026-05 unverdicted novelty 7.0

    Introduces the AG-MG Parallel Corpus of 132k aligned pairs and benchmarks fine-tuning of NLLB, M2M100, and Llama-Krikri-8B models, reporting up to +10.3 BLEU improvement with a peak score of 13.16.

  4. A framework for analyzing concept representations in neural models

    cs.CL 2026-05 unverdicted novelty 7.0

    A new framework shows concept subspaces are not unique, estimator choice affects containment and disentanglement, LEACE works well but generalizes poorly, and HuBERT encodes phone info as contained and disentangled fr...

  5. GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

    eess.AS 2026-06 conditional novelty 6.0

    A 680-hour in-the-wild multilingual ASR/AST benchmark exposes large, consistent degradation of leading commercial and open-source systems on dialects, accents, terminology, age, and low-resource speech.

  6. Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains

    cs.CL 2026-04 unverdicted novelty 6.0

    Automatic translation metrics show lower agreement with humans on unseen technical domains than humans show with each other, and their robustness claims weaken when benchmarked against inter-annotator agreement instea...

  7. PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing

    eess.AS 2026-04 unverdicted novelty 6.0

    PS-TTS and PS-Comet TTS use isochrony via language model paraphrasing plus phonetic synchronization with DTW on vowel distances to achieve better lip-sync and semantic preservation in automated dubbing than standard T...

  8. Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning

    cs.CL 2026-07 conditional novelty 5.0

    On Swiss legal translation, reinforcement learning with a ChrF reward improves small open models more than supervised fine-tuning, but frontier reasoning models still score higher.

  9. An Explainable Approach to Document-level Translation Evaluation with Topic Modeling

    cs.CE 2026-04 unverdicted novelty 5.0

    A topic-modeling framework measures document-level thematic consistency in translations by aligning key tokens across languages with a bilingual dictionary and scoring via cosine similarity, providing explainable insi...

  10. Enhancing Scientific Discourse: Machine Translation for the Scientific Domain

    cs.CL 2026-05 conditional novelty 4.0

    Development of domain-specific scientific corpora for English-Spanish, English-French, and English-Portuguese and their application to fine-tuning NMT models.