Pith. sign in

REVIEW 10 cited by

WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12404 v1 pith:Z5OY7BJP submitted 2025-02-18 cs.CL

WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects

classification cs.CL
keywords languagesdatasetwmt24benchmarkdialectslanguagellmspost-edits
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

As large language models (LLM) become more and more capable in languages other than English, it is important to collect benchmark datasets in order to evaluate their multilingual performance, including on tasks like machine translation (MT). In this work, we extend the WMT24 dataset to cover 55 languages by collecting new human-written references and post-edits for 46 new languages and dialects in addition to post-edits of the references in 8 out of 9 languages in the original WMT24 dataset. The dataset covers four domains: literary, news, social, and speech. We benchmark a variety of MT providers and LLMs on the collected dataset using automatic metrics and find that LLMs are the best-performing MT systems in all 55 languages. These results should be confirmed using a human-based evaluation, which we leave for future work.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation

    cs.CL 2026-05 unverdicted novelty 7.0

    dGRPO merges outcome-based policy optimization with dense teacher guidance from on-policy distillation, yielding more stable long-context reasoning on the new LongBlocks synthetic dataset.

  2. VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models

    cs.CL 2025-08 unverdicted novelty 7.0

    VocabTailor introduces a decoupled dynamic vocabulary selection framework that reduces vocabulary-related memory in SLMs by up to 99% with minimal task performance loss.

  3. HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains

    cs.CL 2026-05 unverdicted novelty 6.0

    HardMTBench is a difficulty-aware benchmark of 20,000 directional test items across 12 domains that widens GEMBA score ranges by a factor of two and reveals domain-specific weaknesses in 22 MT systems.

  4. RouteLMT: Learned Sample Routing for Hybrid LLM Translation Deployment

    cs.CL 2026-04 unverdicted novelty 6.0

    RouteLMT learns to route MT requests to large or small LLMs by predicting marginal quality gain from small-model token representations, yielding a better quality-budget Pareto frontier than baselines.

  5. CHORUS: Effort-Aware Multi-Agent Human-AI Collaboration for Professional Translation

    cs.HC 2026-02 unverdicted novelty 6.0

    CHORUS multi-agent system reduced professional translation time by 33.8% while lowering cognitive effort and raising BLEU/COMET scores in a 30-participant within-subject study.

  6. A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation

    cs.CL 2026-05 unverdicted novelty 5.0

    Combines GRPO with teacher-guided on-policy distillation and introduces LongBlocks dataset to yield more stable long-context reasoning than either method alone.

  7. Hunyuan-MT Technical Report

    cs.CL 2025-09 conditional novelty 5.0

    Hunyuan-MT and Chimera, a 7B open-source translation model and its multi-candidate fusion variant, claim state-of-the-art multilingual translation including Mandarin to minority languages, with open weights.

  8. Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild

    cs.CL 2026-05 unverdicted novelty 4.0

    Hy-MT2 is a new family of fast multilingual translation models that claim to outperform several open-source LLMs and commercial APIs across diverse evaluation settings while supporting efficient on-device deployment.

  9. Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild

    cs.CL 2026-05 unverdicted novelty 3.0

    Hy-MT2 presents three new multilingual translation models that claim to outperform listed open-source and commercial systems on diverse tasks while enabling low-storage on-device use.

  10. EXAONE 4.5 Technical Report

    cs.CL 2026-04 unverdicted novelty 2.0

    EXAONE 4.5 is a new open-weight multimodal model that matches general benchmarks and outperforms similar-scale models on document understanding and Korean contextual reasoning.