Pith. sign in

REVIEW 12 cited by

xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10482 v1 pith:CYPUBE3A submitted 2023-10-16 cs.CL

classification cs.CL
keywords evaluationtranslationerrorerrorsxcometdetectionsentence-levellearned
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Widely used learned metrics for machine translation evaluation, such as COMET and BLEURT, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight into translation errors (e.g., what are the errors and what is their severity). On the other hand, generative large language models (LLMs) are amplifying the adoption of more granular strategies to evaluation, attempting to detail and categorize translation errors. In this work, we introduce xCOMET, an open-source learned metric designed to bridge the gap between these approaches. xCOMET integrates both sentence-level evaluation and error span detection capabilities, exhibiting state-of-the-art performance across all types of evaluation (sentence-level, system-level, and error span detection). Moreover, it does so while highlighting and categorizing error spans, thus enriching the quality assessment. We also provide a robustness analysis with stress tests, and show that xCOMET is largely capable of identifying localized critical errors and hallucinations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 16 citations worldwide. Full citation record

  1. From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set

    cs.CL 2024-11 conditional novelty 7.0 of 10

    Using per-example in-context demonstrations built from historical same-source human MQM ratings makes an LLM judge dramatically better at fine-grained MT evaluation on WMT'23 and WMT'24.

  2. Evaluating LLMs on Chinese Idiom Translation

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Across 900 annotated translation pairs from nine MT systems, the best system still mistranslates Chinese idioms in 28% of cases, and standard metrics miss these errors (Pearson correlation below 0.48).

  3. CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A confidence-reward score for selecting preference pairs improves DPO-based machine translation fine-tuning over reward-only selection methods on ALMA-7B and NLLB-1.3B.

  4. PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics

    cs.CL 2024-12 conditional novelty 6.0 of 10

    PromptOptMe compresses the inputs of the GEMBA-MQM MT evaluation prompt with a two-stage fine-tuned LLaMA 3.2 model, achieving a 2.37x token reduction without quality loss in the headline GPT-4o configuration.

  5. Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation

    cs.CL 2026-08 conditional novelty 5.0 of 10

    Segment-level QE scores are not reliable enough to gate translation review, according to a new 104k-segment evaluation and a synthesis of prior critiques.

  6. Hunyuan-MT Technical Report

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Hunyuan-MT and Chimera, a 7B open-source translation model and its multi-candidate fusion variant, claim state-of-the-art multilingual translation including Mandarin to minority languages, with open weights.

  7. Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A 7B open-weight translation model matches or outperforms far larger commercial systems across 28 languages in automatic and human evaluations.

  8. Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A new data-mixing recipe (Parallel-First Monolingual-Second) and a 9B model, GemmaX2-28, achieve translation quality competitive with Google Translate and GPT-4 across 28 languages.

  9. MT-LENS: An all-in-one Toolkit for Better Machine Translation Evaluation

    cs.CL 2024-12 conditional novelty 5.0 of 10

    MT-LENS is an open-source extension of LM-eval-harness that bundles MT quality, gender bias, added toxicity, and perturbation-robustness evaluations into one command-line and web-based toolkit.

  10. A Measure of the System Dependence of Automated Metrics

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A new metric, SysDep, quantifies how much an MT metric depends on the system being scored, and XCOMET's system dependence is large enough to change system rankings.

  11. An Interdisciplinary Approach to Human-Centered Machine Translation

    cs.CL 2025-06 accept novelty 4.0 of 10

    A position survey calling for human-centered machine translation, synthesizing translation studies and HCI to broaden MT evaluation and design beyond benchmark quality.

  12. Reference-free Evaluation Metrics for Text Generation: A Survey

    cs.CL 2025-01 conditional novelty 4.0 of 10

    The survey classifies reference-free NLG evaluation metrics into learning from human judgments, pseudo-judgments, context-hypothesis correspondence, and peer evaluation.

Pith tools