REVIEW 31 cited by
A Call for Clarity in Reporting BLEU Scores
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The field of machine translation faces an under-recognized problem because of inconsistency in the reporting of scores from its dominant metric. Although people refer to "the" BLEU score, BLEU is in fact a parameterized metric whose values can vary wildly with changes to these parameters. These parameters are often not reported or are hard to find, and consequently, BLEU scores between papers cannot be directly compared. I quantify this variation, finding differences as high as 1.8 between commonly used configurations. The main culprit is different tokenization and normalization schemes applied to the reference. Pointing to the success of the parsing community, I suggest machine translation researchers settle upon the BLEU scheme used by the annual Conference on Machine Translation (WMT), which does not allow for user-supplied reference processing, and provide a new tool, SacreBLEU, to facilitate this.
Forward citations
Cited by 31 Pith papers
-
Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research
Introduces the Matter to Mechanism benchmark of 2,645 structured instances and a composite metric suite for evaluating AI co-scientists on problem-to-hypothesis reasoning in battery materials research.
-
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
The paper teaches the Canary ASR and speech-translation model to output word-level start and end timestamps directly using forced-alignment teacher labels.
-
Knowledge-Enhanced Program Repair for Data Science Code
DSrepair combines a knowledge graph of data science APIs with AST-level bug localization to repair LLM-generated code, fixing more DS-1000 tasks than five baseline repair methods.
-
Towards AI-driven Sign Language Generation with Non-manual Markers
The authors combine an LLM, motion matching, and a pose-to-video model to generate ASL videos with non-manual markers, reporting a BLEU-4 of 0.276 for text-to-gloss and a user study where DHH participants rated genera...
-
Improving FIM Code Completions via Context & Curriculum Based Learning
Fine-tuning FIM code models on curriculum examples with retrieved context improves completion quality and live acceptance, with the largest gains for small models.
-
Widening the Representation Bottleneck in Neural Machine Translation with Lexical Shortcuts
Gated lexical shortcut connections added to the transformer yield 0.9 BLEU average gains on five WMT directions while lowering the lexical content stored in hidden states.
-
Parameter-Efficient Multi-Task Fine-Tuning in Code-Related Tasks
Multi-task QLoRA on Qwen2.5-Coder matches or beats single-task QLoRA and full fine-tuning for code generation and Python summarization, but lags in Java-to-C# translation.
-
FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
A hierarchical Q-Former compresses speech to about 1.67 tokens/sec, enabling hour-long audio processing with near-linear memory scaling and competitive benchmark scores.
-
Are Today's LLMs Ready to Explain Well-Being Concepts?
AI judges can score explanations of well-being concepts, and small models fine-tuned with preference data score better than larger models, although judges and explainers are all AIs.
-
AccessGuru: Leveraging LLMs to Detect and Correct Web Accessibility Violations in HTML Code
AccessGuru combines accessibility testing tools and LLM prompting to correct syntactic, semantic, and layout HTML accessibility violations, reporting up to 84% average violation score decrease on a new benchmark.
-
The first open machine translation system for the Chechen language
A 171K-pair Chechen-Russian parallel corpus plus a fine-tuned NLLB-200 model are released, giving the first open Chechen-Russian translation system with human-evaluated quality near Google Translate.
-
Extend Adversarial Policy Against Neural Machine Translation via Unknown Token
DexChar adds UNK-mediated character perturbations and noisy discriminator augmentation to produce semantic-preserving adversarial examples for subword NMT.
-
Optimizing Speech Multi-View Feature Fusion through Conditional Computation
A gradient-sensitive gating network plus multi-stage dropout fuses FBanks and HuBERT features, matching BLEU while cutting MuST-C training epochs by roughly 1.24x.
-
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
AV-EmoDialog uses speech and face encoders with a large language model to generate emotion-aware dialogue responses from audio-visual input, reporting better emotional alignment than the compared baselines.
-
M$^{3}$-20M: A Large-Scale Multi-Modal Molecule Dataset for AI-driven Drug Design and Discovery
M3-20M is a new multi-modal molecular dataset with over 20 million molecules, and the paper reports improved molecule generation and property prediction with it.
-
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
Dual-LoRA plus Visual Cue Enhancement improves efficient visual instruction tuning over LoRA and LoRA-MoE baselines with near-vanilla-LoRA inference time.
-
Root Mean Square Layer Normalization
RMSNorm delivers re-scaling invariance and comparable accuracy to LayerNorm while cutting computation by skipping mean subtraction, yielding 7-64% runtime reductions across tested models.
-
Low-Resource Corpus Filtering using Multilingual Sentence Embeddings
LASER sentence embeddings are applied directly to filter parallel corpora, achieving the best BLEU scores in the WMT19 low-resource tasks for Nepali-English and Sinhala-English by margins of 1.3 and 1.4.
-
Enhancing Scientific Discourse: Machine Translation for the Scientific Domain
Development of domain-specific scientific corpora for English-Spanish, English-French, and English-Portuguese and their application to fine-tuning NMT models.
-
Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes
Skin-SOAP is a weakly supervised multimodal system that turns a skin lesion image and sparse clinical text into structured SOAP notes, evaluated with two new metrics.
-
TACTIC: Translation Agents with Cognitive-Theoretic Interactive Collaboration
TACTIC, a cognitive-inspired six-agent workflow, improves LLM translation quality over direct prompting on FLORES-200 and WMT24, with the best DeepSeek-V3 setup reaching 96.19 XCOMET on English-to-X.
-
BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System
BeaverTalk combines VAD segmentation, Whisper ASR, and a LoRA-fine-tuned Gemma 3 with a single-sentence memory bank to achieve BLEU 24.64 to 37.23 on ACL 60/60 across two language pairs and two latency regimes.
-
On VLMs for Diverse Tasks in Multimodal Meme Classification
A VLM-exclamation-to-LLM distillation pipeline (CoVExFiL) improves meme classification over prompting and LoRA fine-tuning, especially for sentiment.
-
Pivot Language for Low-Resource Machine Translation
Hindi-pivot transfer gives a 14.2 SacreBLEU on Nepali-English devtest, beating the fully supervised direct baseline by 6.6 points, but with no code, no error bars, and a non-controlled baseline.
-
A comparison of data filtering techniques for English-Polish LLM-based machine translation in the biomedical domain
Filtering an English-Polish biomedical corpus with LASER embeddings lets mBART50 match full-corpus BLEU while using 60 percent of the training data.
-
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
A machine-translated version of MMLU-Pro in nine Indic languages is released as a benchmark, with baseline accuracy scores for multilingual LLMs.
-
Trustformer: A Trusted Federated Transformer
A federated Transformer training method that transmits k-means centroids instead of weights, but whose convergence proof is flawed and privacy claim is unsupported.
-
A Practical Guide for Evaluating LLMs and LLM-Reliant Systems
A guide that organizes LLM evaluation into three pillars (datasets, metrics, and methodology) and introduces a '5 D's' checklist for building evaluation datasets.
-
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
PIER is a WER variant restricted to tagged points of interest and is proposed as a more honest evaluation of code-switched ASR.
-
Robust Machine Translation with Domain Sensitive Pseudo-Sources: Baidu-OSU WMT19 MT Robustness Shared Task System Report
Baidu-OSU WMT19 system achieves >10 BLEU gain on En-Fr and Fr-En social media translation via domain sensitive training and pseudo noisy sources.
-
A Contemporary Survey of Large Language Model Assisted Program Analysis
A review that catalogs how large language models are used in static, dynamic, and hybrid program analysis, and outlines open challenges.
Discussion (0). Continue with ORCID to comment.