REVIEW 3 major objections 5 minor 69 references
Multi-Hypothesis Distillation of Multilingual Neural Translation Models for Low-Resource Languages
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Training a small translation model on several translations of each source sentence, rather than the teacher's single best beam-search output, improves quality for low-resource languages and also softens two known side effects of…
desk verdict A careful, reproducible empirical study showing that training students on multiple sampled teacher translations beats single-beam KD for low-resource pairs; the main claims hold, though the abstract oversells the corpus-size result and the hallucination analysis is thin. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the teacher's output distribution, represented in MHD by $M$ decoded hypotheses per source instead of a single mode. The training signal is $$\mathcal{L}_{\mathrm{MHD}} = -\sum_{i=1}^{N}\sum_{m=1}^{M}\sum_{t=1}^{T_i} \log P(\tilde{y}_{i,m,t} \mid \tilde{y}_{i,m,<t}, x_i; \theta_S),$$ which is simply the standard sequence-level KD loss on a corpus where each source sentence appears $M$ times. The decoding method is what shapes that corpus: beam search and diverse beam search return ranked high-probability lists, giving low variability and, for poorly fitted languages, increasingly improbable continuations as $M$ grows; top-$p$ and top-$k$ sample independently, giving high lexical variability and stable probabilities; MBR reranks epsilon-sampled candidates by expected chrF, filtering out bad translations. The paper uses these properties to explain when MHD helps: sampling wins where the teacher is weak or the corpus is small, while high-quality ranked outputs regain the edge when monolingual data are abundant.
What would settle it
Have human translators rank the FLORES+ devtest outputs of the D1-BS, D10-top-p, and D10-MBR students for eng-ibo and bam-swh; if the sampling- or MBR-trained students do not beat the beam-KD student, the central claim fails.
Extended reading notes
Core claim
MHD is a sequence-level knowledge distillation method: a teacher translation model (NLLB-200 in the 1.3B and 3.3B sizes) decodes $M$ hypotheses $\tilde{y}_{i,1}, \ldots, \tilde{y}_{i,M}$ for each source sentence $x_i$ using one of five decoding strategies, and the student is trained with the standard cross-entropy objective on the repeated dataset, so every source appears $M$ times paired with a different target. The paper's central empirical claim is that, for low-resource directions, this beats the standard sequence-level KD baseline $D^1_{BS}$ in which each source has only its beam-search output. Gains are largest when the hypotheses are sampled (top-$p$ and top-$k$) rather than ranked (beam search and diverse beam search), with MBR decoding best for the least-resourced Bambara pairs. The paper also reports that multi-hypothesis training reduces gender-bias amplification as measured by contrastive conditioning, and reduces hallucinations in most settings, although multiple beam-search hypotheses from a weak teacher can increase them for Bambara.
Load-bearing premise
The load-bearing premise is that chrF++ on the FLORES+ devtest set judges translation quality faithfully for all seven language pairs; the paper validates against COMET only for English-Swahili and Swahili-English, so a metric artifact for Igbo or Bambara would undermine the ranking at the center of the claim.
Editorial extensions
If this is right
- MHD reaches results comparable to standard sequence-level KD while using a much smaller monolingual corpus, so the method lowers the data requirement for distilling a competitive student.
- For the lowest-resource directions (involving Bambara and the into-English pairs), sampling-based hypotheses give the strongest students; top-$p$ is nearly as effective as MBR and much faster.
- With one million source sentences, the advantage of sampling over beam search narrows, and diverse beam search can even lead for Swahili-English; the right decoding choice depends on corpus size and teacher quality.
- Where the teacher is poorly calibrated, multiple beam-search hypotheses degrade student quality, while multiple sampled hypotheses keep improving it, which is direct evidence that the mode is not a good summary of the teacher's distribution.
- Training on $M=10$ hypotheses systematically mitigates gender-bias amplification compared to a single beam hypothesis, with sampling methods reducing it most.
Reading between the lines
- Beyond the paper: if the gains come from diversity rather than from the identity of any single good translation, then controlling diversity directly, for example by tuning top-$p$ while filtering repeated sentences, could further improve student models at lower cost than MBR.
- Beyond the paper: the paper's vocabulary-coverage curves suggest a practical recipe: with scarce monolingual data, spend distillation effort on covering the target vocabulary with several sampled hypotheses first, then switch to high-quality single hypotheses once coverage saturates.
- Beyond the paper: because MHD needs only monolingual text and access to the teacher's outputs, the same recipe could be applied to even lower-resource languages or to API-only teachers, and could be combined with back-translation to grow the source side as well.
- Beyond the paper: the multi-hypothesis idea is not specific to translation; other sequence-generation tasks where the teacher's mode is unrepresentative might benefit from the same keep-several-outputs training signal, but that is an extrapolation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Multi-Hypothesis Distillation (MHD), a sequence-level knowledge distillation method that trains a compact student translation model on multiple teacher-generated translations per source sentence. The teacher is an NLLB-200 model and the students are 65M-parameter Transformers. Experiments cover seven low-resource directions (eng-swh, eng-ibo, eng-bam, swh-eng, ibo-eng, bam-eng, bam-swh), two teacher sizes (1.3B and 3.3B), several decoding methods (beam search, diverse beam search, top-p, top-k, MBR), sweeps over the number of hypotheses M, corpus-size sweeps, and analyses of gender bias, hallucinations, vocabulary coverage, and decoding-parameter sensitivity. The central claim is that MHD with M=10 hypotheses, especially with sampling-based decoding, improves student performance over standard single-hypothesis beam-search KD while also reducing gender-bias amplification and hallucinations.
Significance. If the claims hold, the paper makes a useful practical contribution: it shows that a black-box multilingual teacher accessed only through decoding can be distilled into a much smaller bilingual student using monolingual data, and that sampling-based hypothesis generation can beat beam-search KD in low-resource settings. The study is unusually thorough: 482 trained models, two teacher scales, significance testing with paired approximate randomization, and public code. The vocabulary-coverage and corpus-size analyses (Figures 7-9) are informative and could guide practitioners. The claims are falsifiable and the experimental protocol is mostly reproducible from the description.
major comments (3)
- [Section 5.1, Eq. (3), Fig. 2] The comparison between D^10_Z and D^1_BS confounds the number of hypotheses per source with the total number of training examples and optimization steps. A 100k-source D^10_Z corpus contains 1M target sentences, while D^1_BS contains 100k target sentences. The 'best translation per source' control reported in Section 5.1 rules out a single lucky translation but does not control for data quantity. To attribute the gains to diversity rather than to tenfold more training data, please add a control in which the single D^1_BS translation is repeated ten times per source sentence, or otherwise match the total number of target sentences across conditions.
- [Section 5.1, Tables 8-9] The zero-shot direction bam-swh shows a clear metric disagreement: BLEU ranks D10_BS below D1_BS (1.2 vs 2.1), while chrF++ ranks it above (20.4 vs 8.7). The conclusion that MHD with beam search helps in the zero-shot scenario therefore rests entirely on chrF++, which is validated with COMET only for eng-swh and swh-eng in Section 4.3. Please add a neural metric or human evaluation for at least bam-swh, or restrict the claim to sampling-based MHD, for which BLEU and chrF++ agree in direction.
- [Section 5.4, Table 4] The gender-bias reductions reported in Table 4 are small (e.g., eng-swh D10_BS 51.0 vs D1_BS 49.2; eng-bam D10_BS 50.3 vs 50.8) and no significance testing, confidence intervals, or run-level variance is reported. Since bias mitigation is part of the headline contribution, please provide uncertainty estimates or significance tests for these differences, or soften the corresponding conclusion.
minor comments (5)
- [Table 1 caption] The caption reads 'ChrfF++ scores'; the metric should be written 'chrF++'.
- [Section 2.1] The text spells 'Kullback-Leiber'; the correct name is 'Kullback-Leibler'.
- [Figure 8 caption] The caption contains an unmatched parenthesis in 'the swh-eng) training corpus'; please fix the parenthetical.
- [Appendix D.4, Tables 8-9] The captions refer to underlined and bolded values, but these visual markers are not described in the text; please state explicitly in the caption which comparison each marker refers to and ensure the markers are visible in the published PDF.
- [Template front matter] The received/revised dates in the JAIR template ('Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009') appear to be leftover template text and should be updated or removed.
Circularity Check
No significant circularity: the MHD objective is a standard cross-entropy loss over teacher-generated hypotheses, and the central comparison against single-beam KD is evaluated on external FLORES+ benchmarks without fitting to the test set.
full rationale
The paper's derivation chain is self-contained and empirical. Equation 3 defines MHD simply as sequence-level cross-entropy over M teacher-generated translations per source sentence; no parameter of the method is fitted to FLORES+ devtest, and the student models are evaluated with external metrics (chrF++, BLEU, and COMET for eng-swh and swh-eng) against held-out references. The central claim that sampling-based MHD with M=10 outperforms standard D1_BS sequence-level KD is not defined in terms of its own output: the teacher's synthetic corpora are produced by fixed decoding methods (beam search, diverse beam search, top-p, top-k, MBR) with standard hyperparameters, and student quality is measured independently on FLORES+. The paper even runs a control experiment selecting only the best COMET-without-reference translation per source for eng-swh D10_top-p, obtaining performance similar to D1_top-p, which directly addresses the concern that a single good translation drives the result. The only metric-alignment caveat is the acknowledged use of fastChrF as the MBR utility function while chrF++ is the primary evaluation metric; the authors explicitly state 'we cannot rule out a metric bias introduced by using the same type of metric to rank the MBR candidates and for evaluation', and they report BLEU rankings showing MBR is not most effective under BLEU. This is an honest limitation affecting the MBR-specific comparison, not the main sampling-versus-beam-search claim, and it is not a hidden reduction of the method to its evaluation. Self-citations [17] and [18] are descriptive references to the authors' prior NAACL Findings paper and related work on low-resource NMT; they are not invoked as uniqueness theorems or as substitutes for the experimental evidence, so they are not load-bearing circularity. No fitted parameter is renamed as a prediction, no known result is repackaged under new coordinates, and no claim reduces by construction to an input of the paper.
Assumptions & free parameters
free parameters (5)
- M (number of hypotheses per source sentence) =
10
- p (top-p sampling threshold) =
0.7
- k (top-k sampling size) =
10
- epsilon (MBR candidate sampling threshold) =
0.02
- beam width and DBS diversity penalty =
n=10, lambda=0.5
assumptions (5)
- domain assumption The teacher model's output distribution beyond the mode contains transferable knowledge for the student.
- domain assumption chrF++ on FLORES+ devtest is a valid proxy for translation quality in all seven language pairs.
- domain assumption Contrastive conditioning with NLLB-200 1.3B as evaluator measures gender bias amplification in the target languages.
- domain assumption SONAR embedding cosine similarity identifies hallucinations.
- domain assumption The chosen monolingual corpora (OSCAR, ParaCrawl, bayelemabaga, etc.) are representative of the target languages.
Cite this review
Pith. "Pith review of Multi-Hypothesis Distillation of Multilingual Neural Translation Models for Low-Resource Languages." pith.science (2026). https://pith.science/paper/XLIOGDKK
@misc{pith2026250721568,
author = {Pith},
title = {Pith review of: Multi-Hypothesis Distillation of Multilingual Neural Translation Models for Low-Resource Languages},
year = {2026},
howpublished = {\url{https://pith.science/paper/XLIOGDKK}},
note = {Machine review of arXiv:2507.21568}
}
abstract
This paper explores sequence-level knowledge distillation (KD) of multilingual pre-trained encoder-decoder translation models. We argue that the teacher model's output distribution holds valuable insights for the student, beyond the approximated mode obtained through beam search (the standard decoding method), and present Multi-Hypothesis Distillation (MHD), a sequence-level KD method that generates multiple translations for each source sentence. This provides a larger representation of the teacher model distribution and exposes the student model to a wider range of target-side prefixes. We leverage $n$-best lists from beam search to guide the student's learning and examine alternative decoding methods to address issues like low variability and the under-representation of infrequent tokens. For low-resource languages, our research shows that while sampling methods may slightly compromise translation quality compared to beam search based approaches, they enhance the generated corpora with greater variability and lexical richness. This ultimately improves student model performance and mitigates the gender bias amplification often associated with KD.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
David Adelani, Md Mahfuz Ibn Alam, Antonios Anastasopoulos, Akshita Bhagia, Marta R. Costa-jussà, Jesse Dodge, Fahim Faisal, Christian Federmann, Natalia Fedorova, Francisco Guzmán, Sergey Koshelev, Jean Maillard, Vukosi Marivate, Jonathan Mbuya, Alexandre Mourachko, Safiyyah Saleem, Holger Schwenk, and Guillaume Wenzek. 2022. Findings of the WMT’22 Share...
work page 2022
-
[2]
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. 2024. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=3zKtaqxLhW
work page 2024
-
[3]
Jaimeen Ahn, Hwaran Lee, Jinhwa Kim, and Alice Oh. 2022. Why Knowledge Distillation Amplifies Gender Bias and How to Mitigate from the Perspective of DistilBERT. InProceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP). Association for Computational Linguistics, Seattle, Washington, 266–272. https://doi.org/10.18653/v1/2022...
- [4]
-
[5]
Laurie Burchell, Alexandra Birch, and Kenneth Heafield. 2022. Exploring diversity in back translation for low-resource machine translation. InProceedings of the Third Workshop on Deep Learning for Low-Resource Natural Language Processing. Association for Computational Linguistics, Hybrid, 67–79. https://doi.org/10.18653/v1/2022.deeplo-1.8
-
[6]
David Dale, Elena Voita, Loic Barrault, and Marta R. Costa-jussà. 2023. Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, a...
-
[7]
Ona De Gibert, Raúl Vázquez, Mikko Aulamo, Yves Scherrer, Sami Virpioja, and Jörg Tiedemann. 2023. Four Approaches to Low-Resource Multilingual NMT: The Helsinki Submission to the AmericasNLP 2023 Shared Task. InProceedings of the Workshop on Natural Language Processing for Indigenous Languages of the Americas (AmericasNLP), Manuel Mager, Abteen Ebrahimi,...
work page 2023
-
[8]
Alexandra DeLucia, Aaron Mueller, Xiang Lisa Li, and João Sedoc. 2021. Decoding Methods for Neural Narrative Generation. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021). Association for Computational Linguistics, Online, 166–185. https://doi.org/10.18653/v1/2021.gem-1.16 Submited to JAIR on July 2025. ...
Show all 69 references
-
[9]
Heejin Do and Gary Geunbae Lee. 2023. Target-Oriented Knowledge Distillation with Language-Family-Based Grouping for Multilingual NMT.ACM Trans. Asian Low-Resour. Lang. Inf. Process.22, 2, Article 42 (mar 2023), 18 pages. https://doi.org/10.1145/3546067
2023 doi
-
[10]
Paul-Ambroise Duquenne, Holger Schwenk, and Benoît Sagot. 2023. SONAR: Sentence-Level Multimodal and Language-Agnostic Representations. arXiv:2308.11466 [cs.CL] https://arxiv.org/abs/2308.11466
2023 arXiv
-
[11]
Bryan Eikema and Wilker Aziz. 2020. Is MAP Decoding All You Need? The Inadequacy of the Mode in Neural Machine Translation. In Proceedings of the 28th International Conference on Computational Linguistics. International Committee on Computational Linguistics, Barcelona, Spain ...
2020 doi
-
[12]
Bryan Eikema and Wilker Aziz. 2022. Sampling-Based Approximations to Minimum Bayes Risk Decoding for Neural Machine Translation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). As...
2022 doi
-
[13]
Maxim Enis and Mark Hopkins. 2024. From LLM to NMT: Advancing Low-Resource Machine Translation with Claude. arXiv:2404.13813 [cs.CL] https://arxiv.org/abs/2404.13813
2024 arXiv
-
[14]
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Michael Auli, and Armand Joulin. 2021. Beyond English-Centric M...
2021
-
[15]
Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical Neural Story Generation. arXiv:1805.04833 [cs.CL]
2018 arXiv
-
[16]
Mara Finkelstein and Markus Freitag. 2024. MBR and QE Finetuning: Training-time Distillation of the Best and Most Expensive Decoding Methods. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=bkNx3O0sND
2024
-
[17]
Sánchez-Cartagena
Aarón Galiano-Jiménez, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, and Víctor M. Sánchez-Cartagena. 2025. Beyond the Mode: Sequence-Level Distillation of Multilingual Translation Models for Low-Resource Language Pairs. InFindings of the Association for Computational Lin...
2025 doi
-
[18]
Sánchez-Cartagena, and Juan Antonio Pérez-Ortiz
Aarón Galiano-Jiménez, Felipe Sánchez-Martínez, Víctor M. Sánchez-Cartagena, and Juan Antonio Pérez-Ortiz. 2023. Exploiting large pre-trained models for low-resource neural machine translation. InProceedings of the 24th Annual Conference of the European Association for Machine...
2023
-
[19]
Vikrant Goyal, Sourav Kumar, and Dipti Misra Sharma. 2020. Efficient Neural Machine Translation for Low-Resource Languages via Exploiting Related Languages. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop. As...
2020 doi
-
[20]
Miguel Graça, Yunsu Kim, Julian Schamper, Shahram Khadivi, and Hermann Ney. 2019. Generalizing Back-Translation in Neural Machine Translation. InProceedings of the Fourth Conference on Machine Translation (Volume 1: Research Papers). Association for Computational Linguistics, ...
2019 doi
-
[21]
Alex Graves. 2012. Sequence Transduction with Recurrent Neural Networks. arXiv:1211.3711 [cs.NE]
2012 arXiv
-
[22]
Guerreiro, Elena Voita, and André Martins
Nuno M. Guerreiro, Elena Voita, and André Martins. 2023. Looking for a Needle in a Haystack: A Comprehensive Study of Hallucinations in Neural Machine Translation. InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, An...
2023 doi
-
[23]
Varun Gumma, Raj Dabre, and Pratyush Kumar. 2023. An Empirical Study of Leveraging Knowledge Distillation for Compressing Multilingual Neural Machine Translation Models. InProceedings of the 24th Annual Conference of the European Association for Ma- chine Translation, Mary Nur...
2023
-
[24]
John Hewitt, Christopher Manning, and Percy Liang. 2022. Truncation Sampling as Language Model Desmoothing. InFindings of the Association for Computational Linguistics: EMNLP 2022, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistic...
2022 doi
-
[25]
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the Knowledge in a Neural Network. arXiv:1503.02531 [stat.ML]
2015 arXiv
-
[26]
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curious Case of Neural Text Degeneration. arXiv:1904.09751 [cs.CL]
2020 arXiv
-
[27]
Vivek Iyer, Bhavitvya Malik, Pavel Stepachev, Pinzhen Chen, Barry Haddow, and Alexandra Birch. 2024. Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation. InProceedings of the Ninth Conference on Submited to JAIR on Ju...
2024 doi
-
[28]
Yoon Kim and Alexander M. Rush. 2016. Sequence-Level Knowledge Distillation. InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Austin, Texas, 1317–1327. https://doi.org/10.18653/v1/D16- 1139
2016 doi
-
[29]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In3rd International Conference on Learning Representations, ICLR 2015, Conference Track Proc.http://arxiv.org/abs/1412.6980
2015 arXiv
-
[30]
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, Ma...
2024
-
[31]
Geza Kovacs, Daniel Deutsch, and Markus Freitag. 2024. Mitigating Metric Bias in Minimum Bayes Risk Decoding. InProceedings of the Ninth Conference on Machine Translation, Barry Haddow, Tom Kocmi, Philipp Koehn, and Christof Monz (Eds.). Association for Computational Linguisti...
2024 doi
-
[32]
Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for ...
2018 doi
-
[33]
Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat
Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A. Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2023. MADLAD-400: A Multilingual And Document-Level Large Audited Dataset. arXiv:2309.04662 [cs.CL]
2023 arXiv
-
[34]
Ilia Kulikov, Alexander Miller, Kyunghyun Cho, and Jason Weston. 2019. Importance of Search and Evaluation Strategies in Neural Dialogue Modeling. InProceedings of the 12th International Conference on Natural Language Generation. Association for Computational Linguistics, Toky...
2019 doi
-
[35]
Solomon Kullback and Richard A Leibler. 1951. On Information and Sufficiency.The Annals of Mathematical Statistics22, 1 (1951), 79–86
1951
-
[36]
Shankar Kumar and William Byrne. 2004. Minimum Bayes-Risk Decoding for Statistical Machine Translation. InProceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL
2004
-
[37]
Wen Lai, Jindřich Libovický, and Alexander Fraser. 2021. The LMU Munich System for the WMT 2021 Large-Scale Multilingual Machine Translation Shared Task. InProceedings of the Sixth Conference on Machine Translation. Association for Computational Linguistics, Online, 412–417. h...
2021
-
[38]
Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, and David Ifeoluwa Adelani. 2025. SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced Afric...
2025
-
[39]
Mathias Müller and Rico Sennrich. 2021. Understanding the Properties of Minimum Bayes Risk Decoding in Neural Machine Translation. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural L...
2021
- [40]
-
[41]
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021. MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers. arXiv:2102.01454 [cs.CL]
2021 arXiv
-
[42]
Maja Popović. 2017. chrF++: words helping character n-grams. InProceedings of the Second Conference on Machine Translation. Association for Computational Linguistics, Copenhagen, Denmark, 612–618. https://doi.org/10.18653/v1/W17-4770
2017 doi
-
[43]
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2015. Sequence Level Training with Recurrent Neural Networks.CoRRabs/1511.06732 (2015). https://api.semanticscholar.org/CorpusID:7147309
2015 arXiv
-
[44]
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A Neural Framework for MT Evaluation. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, Online, 2685–2702. https:/...
2020 doi
-
[45]
Stefan Riezler and John T. Maxwell. 2005. On Some Pitfalls in Automatic Evaluation and Significance Testing for MT. InProceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization. Association for Computational Ling...
2005
-
[46]
Sánchez-Cartagena, Marta Bañón, Sergio Ortiz-Rojas, and Gema Ramírez-Sánchez
Víctor M. Sánchez-Cartagena, Marta Bañón, Sergio Ortiz-Rojas, and Gema Ramírez-Sánchez. 2018. Prompsit’s submission to WMT 2018 Parallel Corpus Filtering shared task. InProceedings of the Third Conference on Machine Translation, Volume 2: Shared Task Papers. Association for Co...
2018
-
[47]
Barbara Scalvini, Iben Nyholm Debess, Annika Simonsen, and Hafsteinn Einarsson. 2025. Rethinking Low-Resource MT: The Surprising Effectiveness of Fine-Tuned Multilingual Models in the LLM Age. InProceedings of the Joint 25th Nordic Conference on Computational Linguistics and 1...
2025
-
[48]
Chufan Shi, Haoran Yang, Deng Cai, Zhisong Zhang, Yifan Wang, Yujiu Yang, and Wai Lam. 2024. A Thorough Examination of Decoding Methods in the Era of LLMs. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal,...
2024 doi
-
[49]
Yewei Song, Saad Ezzini, Jacques Klein, Tegawende Bissyande, Clément Lefebvre, and Anne Goujon. 2023. Letz Translate: Low-Resource Machine Translation for Luxembourgish. arXiv:2303.01347 [cs.CL]
2023 arXiv
-
[50]
Smith, and Luke Zettlemoyer
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019. Evaluating Gender Bias in Machine Translation. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 1679–1684. https:...
2019 doi
-
[51]
Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. 2022. A Contrastive Framework for Neural Text Generation. arXiv:2202.06417 [cs.CL]
2022 arXiv
-
[52]
Xu Tan, Yi Ren, Di He, Tao Qin, and Tie-Yan Liu. 2019. Multilingual Neural Machine Translation with Knowledge Distillation. InSeventh International Conference on Learning Representations. https://openreview.net/forum?id=S1gUsoR9YX
2019
-
[53]
Chau Tran, Shruti Bhosale, James Cross, Philipp Koehn, Sergey Edunov, and Angela Fan. 2021. Facebook AI WMT21 News Translation Task Submission. InProc. of the Sixth Conference on Machine Translation (WMT). 205–215
2021
-
[54]
Jannis Vamvas and Rico Sennrich. 2021. Contrastive Conditioning for Assessing Disambiguation in MT: A Case Study of Distilled Bias. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Online and P...
2021 doi
-
[55]
Jannis Vamvas and Rico Sennrich. 2024. Linear-time Minimum Bayes Risk Decoding with Reference Aggregation. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). ...
2024 doi
-
[56]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InProceedings of the 31st International Conference on Neural Information Processing Systems(Long Beach, California, US...
2017
-
[57]
Ashwin Vijayakumar, Michael Cogswell, Ramprasaath Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2018. Diverse Beam Search for Improved Description of Complex Scenes.Proceedings of the AAAI Conference on Artificial Intelligence32, 1 (Apr. 2018). https://doi....
2018 doi
-
[58]
Jiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Anuoluw...
2024
-
[59]
Jun Wang, Eleftheria Briakou, Hamid Dadkhahi, Rishabh Agarwal, Colin Cherry, and Trevor Cohn. 2024. Don’t Throw Away Data: Better Sequence Knowledge Distillation. arXiv:2407.10456 [cs.CL] https://arxiv.org/abs/2407.10456
2024 arXiv
-
[60]
Gian Wiher, Clara Meister, and Ryan Cotterell. 2022. On Decoding Strategies for Neural Text Generators. arXiv:2203.15721 [cs.CL]
2022 arXiv
-
[61]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[62]
Martindale, and Marine Carpuat
Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, and Marine Carpuat. 2023. Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection.Transactions of the Association for Computational Linguistics11 (2023), 546–564. htt...
2023 doi
-
[63]
Zhengzhe Yu, Daimeng Wei, Zongyao Li, Hengchao Shang, Xiaoyu Chen, Zhanglin Wu, Jiaxin Guo, Minghan Wang, Lizhi Lei, Min Zhang, Hao Yang, and Ying Qin. 2021. HW-TSC’s Participation in the WMT 2021 Large-Scale Multilingual Translation Task. InProceedings of the Sixth Conference...
2021
-
[64]
Hugh Zhang, Daniel Duckworth, Daphne Ippolito, and Arvind Neelakantan. 2021. Trading Off Diversity and Quality in Natural Language Generation. InProceedings of the Workshop on Human Evaluation of NLP Systems (HumEval). Association for Computational Linguistics, Online, 25–33. ...
2021
-
[65]
Songming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen, Wenjuan Han, Jian Liu, and Jinan Xu. 2023. Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation. InProceedings of the 61st Annual Meeting of the Association for Computational Linguis...
2023 doi
-
[66]
Yuhao Zhang, Ziyang Wang, Runzhe Cao, Binghao Wei, Weiqiao Shan, Shuhan Zhou, Abudurexiti Reheman, Tao Zhou, Xin Zeng, Laohu Wang, et al. 2020. The niutrans machine translation systems for wmt20. InProceedings of the Fifth Conference on Machine Translation. 338–345
2020
-
[67]
Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2024. Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis. InFindings of the Association for Computational Linguistics: NAACL 2024,...
2024 doi
-
[68]
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018. Texygen: A Benchmarking Platform for Text Generation Models. InThe 41st International ACM SIGIR Conference on Research & Development in Information Retrieval(Ann Arbor, MI, USA)(SIGIR ’18)...
2018
-
[2004]
https://aclanthology.org/N04-1022
Association for Computational Linguistics, Boston, Massachusetts, USA, 169–176. https://aclanthology.org/N04-1022
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.