Pith. sign in

REVIEW 31 cited by

A Call for Clarity in Reporting BLEU Scores

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1804.08771 v2 pith:N47GMDVJ submitted 2018-04-23 cs.CL

classification cs.CL
keywords bleumachinescorestranslationmetricparametersreferencereporting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The field of machine translation faces an under-recognized problem because of inconsistency in the reporting of scores from its dominant metric. Although people refer to "the" BLEU score, BLEU is in fact a parameterized metric whose values can vary wildly with changes to these parameters. These parameters are often not reported or are hard to find, and consequently, BLEU scores between papers cannot be directly compared. I quantify this variation, finding differences as high as 1.8 between commonly used configurations. The main culprit is different tokenization and normalization schemes applied to the reference. Pointing to the success of the parsing community, I suggest machine translation researchers settle upon the BLEU scheme used by the annual Conference on Machine Translation (WMT), which does not allow for user-supplied reference processing, and provide a new tool, SacreBLEU, to facilitate this.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research

    cs.CE 2026-06 unverdicted novelty 7.0 of 10

    Introduces the Matter to Mechanism benchmark of 2,645 structured instances and a composite metric suite for evaluating AI co-scientists on problem-to-hypothesis reasoning in battery materials research.

  2. Word Level Timestamp Generation for Automatic Speech Recognition and Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The paper teaches the Canary ASR and speech-translation model to output word-level start and end timestamps directly using forced-alignment teacher labels.

  3. Knowledge-Enhanced Program Repair for Data Science Code

    cs.SE 2025-02 conditional novelty 6.0 of 10

    DSrepair combines a knowledge graph of data science APIs with AST-level bug localization to repair LLM-generated code, fixing more DS-1000 tasks than five baseline repair methods.

  4. Towards AI-driven Sign Language Generation with Non-manual Markers

    cs.HC 2025-02 conditional novelty 6.0 of 10

    The authors combine an LLM, motion matching, and a pose-to-video model to generate ASL videos with non-manual markers, reporting a BLEU-4 of 0.276 for text-to-gloss and a user study where DHH participants rated genera...

  5. Improving FIM Code Completions via Context & Curriculum Based Learning

    cs.IR 2024-12 conditional novelty 6.0 of 10

    Fine-tuning FIM code models on curriculum examples with retrieved context improves completion quality and live acceptance, with the largest gains for small models.

  6. Widening the Representation Bottleneck in Neural Machine Translation with Lexical Shortcuts

    cs.CL 2019-06 conditional novelty 6.0 of 10

    Gated lexical shortcut connections added to the transformer yield 0.9 BLEU average gains on five WMT directions while lowering the lexical content stored in hidden states.

  7. Parameter-Efficient Multi-Task Fine-Tuning in Code-Related Tasks

    cs.SE 2026-01 conditional novelty 5.0 of 10

    Multi-task QLoRA on Qwen2.5-Coder matches or beats single-task QLoRA and full fine-tuning for code generation and Python summarization, but lags in Java-to-C# translation.

  8. FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation

    eess.AS 2026-01 conditional novelty 5.0 of 10

    A hierarchical Q-Former compresses speech to about 1.67 tokens/sec, enabling hour-long audio processing with near-linear memory scaling and competitive benchmark scores.

  9. Are Today's LLMs Ready to Explain Well-Being Concepts?

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    AI judges can score explanations of well-being concepts, and small models fine-tuned with preference data score better than larger models, although judges and explainers are all AIs.

  10. AccessGuru: Leveraging LLMs to Detect and Correct Web Accessibility Violations in HTML Code

    cs.SE 2025-07 conditional novelty 5.0 of 10

    AccessGuru combines accessibility testing tools and LLM prompting to correct syntactic, semantic, and layout HTML accessibility violations, reporting up to 84% average violation score decrease on a new benchmark.

  11. The first open machine translation system for the Chechen language

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A 171K-pair Chechen-Russian parallel corpus plus a fine-tuned NLLB-200 model are released, giving the first open Chechen-Russian translation system with human-evaluated quality near Google Translate.

  12. Extend Adversarial Policy Against Neural Machine Translation via Unknown Token

    cs.CL 2025-01 conditional novelty 5.0 of 10

    DexChar adds UNK-mediated character perturbations and noisy discriminator augmentation to produce semantic-preserving adversarial examples for subword NMT.

  13. Optimizing Speech Multi-View Feature Fusion through Conditional Computation

    eess.AS 2025-01 conditional novelty 5.0 of 10

    A gradient-sensitive gating network plus multi-stage dropout fuses FBanks and HuBERT features, matching BLEU while cutting MuST-C training epochs by roughly 1.24x.

  14. AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues

    cs.CV 2024-12 conditional novelty 5.0 of 10

    AV-EmoDialog uses speech and face encoders with a large language model to generate emotion-aware dialogue responses from audio-visual input, reporting better emotional alignment than the compared baselines.

  15. M$^{3}$-20M: A Large-Scale Multi-Modal Molecule Dataset for AI-driven Drug Design and Discovery

    q-bio.QM 2024-12 conditional novelty 5.0 of 10

    M3-20M is a new multi-modal molecular dataset with over 20 million molecules, and the paper reports improved molecule generation and property prediction with it.

  16. From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Dual-LoRA plus Visual Cue Enhancement improves efficient visual instruction tuning over LoRA and LoRA-MoE baselines with near-vanilla-LoRA inference time.

  17. Root Mean Square Layer Normalization

    cs.LG 2019-10 conditional novelty 5.0 of 10

    RMSNorm delivers re-scaling invariance and comparable accuracy to LayerNorm while cutting computation by skipping mean subtraction, yielding 7-64% runtime reductions across tested models.

  18. Low-Resource Corpus Filtering using Multilingual Sentence Embeddings

    cs.CL 2019-06 unverdicted novelty 5.0 of 10

    LASER sentence embeddings are applied directly to filter parallel corpora, achieving the best BLEU scores in the WMT19 low-resource tasks for Nepali-English and Sinhala-English by margins of 1.3 and 1.4.

  19. Enhancing Scientific Discourse: Machine Translation for the Scientific Domain

    cs.CL 2026-05 conditional novelty 4.0 of 10

    Development of domain-specific scientific corpora for English-Spanish, English-French, and English-Portuguese and their application to fine-tuning NMT models.

  20. Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Skin-SOAP is a weakly supervised multimodal system that turns a skin lesion image and sparse clinical text into structured SOAP notes, evaluated with two new metrics.

  21. TACTIC: Translation Agents with Cognitive-Theoretic Interactive Collaboration

    cs.CL 2025-06 conditional novelty 4.0 of 10

    TACTIC, a cognitive-inspired six-agent workflow, improves LLM translation quality over direct prompting on FLORES-200 and WMT24, with the best DeepSeek-V3 setup reaching 96.19 XCOMET on English-to-X.

  22. BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System

    cs.CL 2025-05 conditional novelty 4.0 of 10

    BeaverTalk combines VAD segmentation, Whisper ASR, and a LoRA-fine-tuned Gemma 3 with a single-sentence memory bank to achieve BLEU 24.64 to 37.23 on ACL 60/60 across two language pairs and two latency regimes.

  23. On VLMs for Diverse Tasks in Multimodal Meme Classification

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A VLM-exclamation-to-LLM distillation pipeline (CoVExFiL) improves meme classification over prompting and LoRA fine-tuning, especially for sentiment.

  24. Pivot Language for Low-Resource Machine Translation

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Hindi-pivot transfer gives a 14.2 SacreBLEU on Nepali-English devtest, beating the fully supervised direct baseline by 6.6 points, but with no code, no error bars, and a non-controlled baseline.

  25. A comparison of data filtering techniques for English-Polish LLM-based machine translation in the biomedical domain

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Filtering an English-Polish biomedical corpus with LASER embeddings lets mBART50 match full-corpus BLEU while using 60 percent of the training data.

  26. IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A machine-translated version of MMLU-Pro in nine Indic languages is released as a benchmark, with baseline accuracy scores for multilingual LLMs.

  27. Trustformer: A Trusted Federated Transformer

    cs.LG 2025-01 reject novelty 4.0 of 10

    A federated Transformer training method that transmits k-means centroids instead of weights, but whose convergence proof is flawed and privacy claim is unsupported.

  28. A Practical Guide for Evaluating LLMs and LLM-Reliant Systems

    cs.AI 2025-06 conditional novelty 3.0 of 10

    A guide that organizes LLM evaluation into three pillars (datasets, metrics, and methodology) and introduces a '5 D's' checklist for building evaluation datasets.

  29. PIER: A Novel Metric for Evaluating What Matters in Code-Switching

    cs.CL 2025-01 reject novelty 3.0 of 10

    PIER is a WER variant restricted to tagged points of interest and is proposed as a more honest evaluation of code-switched ASR.

  30. Robust Machine Translation with Domain Sensitive Pseudo-Sources: Baidu-OSU WMT19 MT Robustness Shared Task System Report

    cs.CL 2019-06 unverdicted novelty 3.0 of 10

    Baidu-OSU WMT19 system achieves >10 BLEU gain on En-Fr and Fr-En social media translation via domain sensitive training and pseudo noisy sources.

  31. A Contemporary Survey of Large Language Model Assisted Program Analysis

    cs.SE 2025-02 conditional novelty 1.0 of 10

    A review that catalogs how large language models are used in static, dynamic, and hybrid program analysis, and outlines open challenges.

Pith tools