Pith. sign in

REVIEW 7 cited by

WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12404 v1 pith:Z5OY7BJP submitted 2025-02-18 cs.CL

classification cs.CL
keywords languagesdatasetwmt24benchmarkdialectslanguagellmspost-edits
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As large language models (LLM) become more and more capable in languages other than English, it is important to collect benchmark datasets in order to evaluate their multilingual performance, including on tasks like machine translation (MT). In this work, we extend the WMT24 dataset to cover 55 languages by collecting new human-written references and post-edits for 46 new languages and dialects in addition to post-edits of the references in 8 out of 9 languages in the original WMT24 dataset. The dataset covers four domains: literary, news, social, and speech. We benchmark a variety of MT providers and LLMs on the collected dataset using automatic metrics and find that LLMs are the best-performing MT systems in all 55 languages. These results should be confirmed using a human-based evaluation, which we leave for future work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation

    cs.CL 2025-02 conditional novelty 7.0 of 10

    BOUQuET is a handcrafted, multicentric, paragraph-level machine translation evaluation dataset in 8 non-English pivot languages, designed to be community-extendable.

  2. Disentangling Language Modeling and Boundaries

    cs.CL 2026-08 reject novelty 6.0 of 10

    The paper hypothesizes that next-byte and boundary distributions in byte-level LMs can be disentangled, proposes two experiments to test it, but provides no experimental results.

  3. Enhancing Robustness of Autoregressive Language Models against Orthographic Attacks via Pixel-based Approach

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A word-as-image pixel language model trained with next-token prediction reports lower perplexity than a token-embedding LLaMA on noisy and non-Latin-script text, though its noise evaluation holds tokenization fixed.

  4. Hunyuan-MT Technical Report

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Hunyuan-MT and Chimera, a 7B open-source translation model and its multi-candidate fusion variant, claim state-of-the-art multilingual translation including Mandarin to minority languages, with open weights.

  5. Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model

    cs.CL 2025-05 reject novelty 5.0 of 10

    A compact 1.5B Arabic-English model beats GPT-4o mini only on the authors' own Tarjama-25 benchmark, while trailing large models on standard WMT24++ and IWSLT2017 tests.

  6. TACTIC: Translation Agents with Cognitive-Theoretic Interactive Collaboration

    cs.CL 2025-06 conditional novelty 4.0 of 10

    TACTIC, a cognitive-inspired six-agent workflow, improves LLM translation quality over direct prompting on FLORES-200 and WMT24, with the best DeepSeek-V3 setup reaching 96.19 XCOMET on English-to-X.

  7. TransBench: Benchmarking Machine Translation for Industrial-Scale Applications

    cs.CL 2025-05 reject novelty 2.0 of 10

    TransBench is a proposed e-commerce MT benchmark with a three-level evaluation framework and a fine-tuned quality-scoring model, but the paper contains no results and no released data or code.

Pith tools