REVIEW 2 major objections 6 minor 37 references
An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages
T0 review · 2 major / 6 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read BM25 retrieval makes many-shot translation into low-resource languages far more data-efficient: 50 retrieved examples roughly match 250 random ones.
desk verdict Solid empirical map of many-shot ICL for truly low-resource MT, with a usable BM25 efficiency result (50≈250, 250≈1000) that holds across models and is honestly costed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Many-shot ICL with English-side BM25 retrieval: for each test sentence the system ranks a fixed ~1,012-sentence parallel pool by BM25 similarity on the English source, inserts the top-k parallel pairs into a long prompt, and lets the model generate the translation in one pass. That retrieval step is what produces the reported efficiency gains over random many-shot selection.
What would settle it
Native-speaker Direct Assessment or adequacy ratings on held-out sentences for the same ten languages showing that 50 BM25 examples systematically underperform 250 random examples, or that 250 BM25 fails to match 1,000 random ones, would falsify the data-efficiency claim.
Extended reading notes
Core claim
Many-shot in-context learning improves English-to-low-resource machine translation roughly log-linearly as the number of parallel examples grows from 0 to 1,000. The decisive result is that BM25 retrieval of more informative English-side examples substantially raises data efficiency: on average, 50 BM25 examples roughly match 250 random many-shot examples, and 250 BM25 examples perform similarly to 1,000 random ones. In-domain examples scale more reliably than Bible data for most languages, length-based ordering has little effect, and ICL continues to improve results even after LoRA fine-tuning of an open model.
Load-bearing premise
The claim treats automatic chrF++ and spBLEU scores on FLORES-style Wikipedia sentences, using a fixed English-side pool of about a thousand examples, as a reliable enough proxy for real translation quality and example usefulness.
Editorial extensions
If this is right
- Low-resource communities can approach 1,000-shot quality with roughly 50–250 retrieved examples, cutting inference cost and context length.
- When in-domain data exist they should be preferred; out-of-domain religious text still helps some languages but plateaus or degrades as shots increase.
- Open-weight models can match full-pool fine-tuning quality with only 10–25 retrieved ICL examples, and adding ICL on top of fine-tuning still helps.
- Gains are larger for translation into the low-resource language than for translation into English.
- For proprietary long-context models, many-shot ICL remains usable even when fine-tuning is unavailable.
Reading between the lines
- Lexical BM25 may be sufficient for many low-resource MT settings, reducing the need for dense embedding retrieval or elaborate selection pipelines.
- The large 0-to-1-shot jump for Tifinagh suggests that under-represented scripts can be unlocked by a single correct-script example, after which further gains come mainly from lexical and syntactic coverage.
- If the same efficiency ratio holds for other generation tasks into low-resource languages, BM25 many-shot ICL could become a default cheap adaptation method when fine-tuning is impractical.
- Human adequacy may lag automatic metrics at low shot counts, so deployment thresholds should be set by native-speaker judgments rather than chrF++ plateaus alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a systematic empirical study of many-shot in-context learning for English↔X machine translation into ten extremely low-resource languages recently added to FLORES+. Across Gemini 2.5 Flash, GPT-4.1, Llama 3.3 70B, and Gemma 3 27B, and shot counts from 0 to 1,000, the authors show that performance (chrF++/spBLEU) generally improves with more in-context examples, with larger gains for eng→X than X→eng. The central practical result is that English-side BM25 retrieval substantially improves sample efficiency: on Gemini eng→X, 50 BM25 examples roughly match 250 random examples and 250 BM25 examples roughly match 1,000 random examples (Table 1). Ablations cover out-of-domain Bible examples, length-based ordering, and LoRA fine-tuning of Gemma 3 27B, where ICL remains complementary to fine-tuning.
Significance. If the efficiency and complementarity findings hold under the stated evaluation protocol, the paper is a useful contribution for low-resource MT practice and for the many-shot ICL literature. The multi-model, multi-language shot grids, BM25-vs-random comparison, domain and ordering ablations, and LoRA+ICL case study are more thorough than typical few-shot LRL MT papers. Explicit cost accounting (>$30k API spend; per-direction cost estimates) and the decision to release results are strengths that improve reproducibility for communities that cannot re-run the full grid. The work is primarily empirical rather than theoretical; its value is in careful measurement and a concrete, low-overhead retrieval recipe (BM25 on English) that reduces inference cost while preserving most of the many-shot gain.
major comments (2)
- Appendix I (native-speaker DA on one eng→Oro sentence) shows that low-shot outputs can score non-trivially on chrF++ while receiving DA=0, with adequacy only emerging at k≥100. The abstract and §3.2 efficiency claim (50 BM25 ≈ 250 random; 250 BM25 ≈ 1,000 random) is stated in automatic-metric terms, but the introduction frames the result as improving practical usability for LRL communities. Please either (i) qualify the efficiency claim more carefully in the abstract/conclusion as metric-level sample efficiency, or (ii) add a small systematic human evaluation (even a few dozen items on 2–3 languages) so that the practical-quality interpretation is not resting on a single illustrative sentence.
- Table 1 and Figure 1 show that BM25 gains and many-shot scaling are highly language-dependent (large for Oro/Tamazight/Ladin; near-flat for Sudanese Arabic and Mauritian Creole on some settings), and open-weight models degrade or hit context limits at k=500–1000 (§3.1). The abstract’s averaged Gemini eng→X statement is accurate for the reported mean, but the main text should more explicitly bound the claim: efficiency gains are strongest for proprietary long-context models and for languages with large headroom, not a uniform law across all ten languages and all four models.
minor comments (6)
- Figure 1 is dense (20 panels). Consider moving some languages to the appendix or using a shared y-axis range per direction so cross-language scaling is easier to compare.
- §2 states ICL examples are sampled from the FLORES+ devtest split and evaluation on dev; please restate this once in the main results section so readers do not miss the pool/test split when interpreting BM25 retrieval.
- Table 2 reports eng→Efik only in the main text; eng→Ibibio is deferred to Appendix G. A one-sentence quantitative summary for Ibibio in §3.2 would make the complementarity claim easier to verify without the appendix.
- Limitations correctly notes English-centric pairs, fixed prompts, and lack of COMET/MetricX. A brief forward pointer from §3.2 to Limitations when discussing “practical usability” would help readers weight the automatic metrics appropriately.
- Minor consistency: abstract says “ten truly low-resource languages recently added to FLORES+”; footnote 1 notes Tamazight/Quechua were in FLORES-200 with community improvements—align the wording so “recently added” is not overstated.
- In Appendix H, dictionary results are negative at matched budget; a short sentence in the main text (or Related Work) noting that dictionary augmentation did not help under realistic LRL dictionary quality would prevent readers from assuming it was untested.
Circularity Check
No significant circularity: efficiency ratios and scaling claims are measured outcomes on held-out FLORES+ queries, not quantities forced by definition or self-citation.
full rationale
This is an empirical comparison paper, not a first-principles derivation. The load-bearing claims (many-shot scaling; 50 BM25 ≈ 250 random and 250 BM25 ≈ 1,000 random on avg eng→X chrF++; domain and ordering ablations; ICL complementary to LoRA) are obtained by constructing prompts from a fixed ~1,012-sentence pool, generating translations with external LLMs, and scoring against FLORES+ references with chrF++/spBLEU. BM25 is the standard Robertson–Zaragoza ranker applied to English sources; it is not defined in terms of the reported efficiency ratios, and those ratios are post-hoc averages over languages (Table 1, Figure 1). Self-citations (e.g., Adelani et al. 2022 on African MT data needs; dataset papers for Ibom-NLP / Sudanese-Flores / FLORES+ extensions) supply background and data provenance, not uniqueness theorems or ansatzes that force the measured deltas. Fine-tuning comparisons use the same pool as ICL but report complementary gains rather than tautological reuse. Automatic-metric reliability for extremely LRLs is a validity concern already noted in Limitations; it does not make the reported numbers equal their inputs by construction. No self-definitional loop, fitted-input-as-prediction, or load-bearing self-citation chain is present.
Assumptions & free parameters
free parameters (3)
- shot counts k
- BM25 default parameters and English tokenization
- LoRA fine-tuning hyperparameters and subset sizes (250 vs 1012)
assumptions (4)
- domain assumption FLORES+/Ibom-NLP parallel sentences are sufficiently clean and representative that more in-domain examples improve generation rather than inject systematic noise.
- domain assumption chrF++ and spBLEU on FLORES splits are adequate primary quality measures for these languages (COMET/MetricX unavailable).
- domain assumption English-side BM25 is a valid proxy for example usefulness for eng→X ICL.
- ad hoc to paper Fixed single prompt template without CoT or heavy prompt engineering is representative of practical many-shot MT.
Cite this review
Pith. "Pith review of An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages." pith.science (2026). https://pith.science/paper/KXMSVFRR
@misc{pith2026260402596,
author = {Pith},
title = {Pith review of: An Empirical Study of Many-Shot In-Context Learning for Machine Translation of Low-Resource Languages},
year = {2026},
howpublished = {\url{https://pith.science/paper/KXMSVFRR}},
note = {Machine review of arXiv:2604.02596}
}
read the original abstract
In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks from a few examples, making it promising for languages underrepresented in pre-training. Recent work on many-shot ICL suggests that modern LLMs can further benefit from larger ICL examples enabled by their long context windows. However, such gains depend on careful example selection, and the inference cost can be prohibitive for low-resource language communities. In this paper, we present an empirical study of many-shot ICL for machine translation from English into ten truly low-resource languages recently added to FLORES+. We analyze the effects of retrieving more informative examples, using out-of-domain data, and ordering examples by length. Our findings show that many-shot ICL becomes more effective as the number of examples increases. More importantly, we show that BM25-based retrieval substantially improves data efficiency: 50 retrieved examples roughly match 250 many-shot examples, while 250 retrieved examples perform similarly to 1,000 many-shot examples. We further show that ICL provides additional gains on top of fine-tuning.
Figures
Reference graph
Works this paper leans on
-
[1]
David Ifeoluwa Adelani, Jesujoba Oluwadara Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen H. Muhammad, Guyo D. Jarso, Oreen Yousuf, and 26 others. 2022. https://doi.org...
-
[2]
Zhang, Bernd Bohnet, Luis Rosias, Stephanie C
Rishabh Agarwal, Avi Singh, Lei M. Zhang, Bernd Bohnet, Luis Rosias, Stephanie C. Y. Chan, Biao Zhang, Ankesh Anand, Zaheer Abbas, Azade Nova, John D. Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle. 2024. https://openreview.net/forum?id=AB6XpMzvqH Many- Shot In - Context Learning . In The Thirty -eighth Annual Conference on Ne...
2024
-
[3]
Sweta Agrawal, Chunting Zhou, Mike Lewis, Luke Zettlemoyer, and Marjan Ghazvininejad. 2023. https://doi.org/10.18653/v1/2023.findings-acl.564 In-context examples selection for machine translation . In Findings of the Association for Computational Linguistics: ACL 2023, pages 8857--8873, Toronto, Canada. Association for Computational Linguistics
-
[4]
Felermino Dario Mario Ali, Henrique Lopes Cardoso, and Rui Sousa-Silva. 2024. https://doi.org/10.18653/v1/2024.wmt-1.45 Expanding FLORES + benchmark for more low-resource settings: P ortuguese-emakhuwa machine translation evaluation . In Proceedings of the Ninth Conference on Machine Translation, pages 579--592, Miami, Florida, USA. Association for Comput...
-
[5]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and 1 others. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877--1901
2020
-
[6]
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobsson, Idan Szpektor, Nan-Jiang Jiang, and 3416 others. 2025. https://arxiv.org/abs/2507.06261 Gemini 2.5: Pus...
arXiv 2025
-
[7]
Sara Court and Micha Elsner. 2024. https://doi.org/10.18653/v1/2024.wmt-1.125 Shortcomings of LLM s for low-resource translation: Retrieval and understanding are both the problem . In Proceedings of the Ninth Conference on Machine Translation, pages 1332--1354, Miami, Florida, USA. Association for Computational Linguistics
-
[8]
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, and Zhifang Sui. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.64 A survey on in-context learning . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1107--1128, Miami, Florid...
Show all 37 references
-
[9]
Andrew Drozdov, Honglei Zhuang, Zhuyun Dai, Zhen Qin, Razieh Rahimi, Xuanhui Wang, Dana Alon, Mohit Iyyer, Andrew McCallum, Donald Metzler, and Kai Hui. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.950 P a R a D e: Passage ranking using demonstrations with LLM s . In ...
2023 doi
-
[10]
Samuel Frontull, Thomas Str \"o hle, Carlo Zoli, Werner Pescosta, Ulrike Frenademez, Matteo Ruggeri, Daria Valentin, Karin Comploj, Gabriel Perathoner, Silvia Liotto, and Paolo Anvidalfarei. 2025. https://aclanthology.org/2025.wmt-1.81/ Bringing L adin to FLORES + . In Proceed...
2025
-
[11]
Gemini-Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, and 1 others. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530
2024 arXiv
-
[12]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[13]
Juraj Juraska, Mara Finkelstein, Daniel Deutsch, Aditya Siddhant, Mehdi Mirzazadeh, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.wmt-1.63 M etric X -23: The G oogle submission to the WMT 2023 metrics shared task . In Proceedings of the Eighth Conference on Machin...
2023 doi
-
[14]
Oluwadara Kalejaiye, Luel Hagos Beyene, David Ifeoluwa Adelani, Mmekut-mfon Gabriel Edet, Aniefon Daniel Akpan, Eno-Abasi Urua, and Anietie Andy. 2025. https://aclanthology.org/2025.ijcnlp-long.22/ Ibom NLP : A step toward inclusive natural language processing for N igeria ' s...
2025
-
[15]
Xiaonan Li and Xipeng Qiu. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.411 Finding support examples for in-context learning . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 6219--6235, Singapore. Association for Computational Linguistics
2023 doi
-
[16]
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O ' Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, M...
2022 doi
-
[17]
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. https://doi.org/10.18653/v1/2022.acl-long.556 Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity . In Proceedings of the 60th Annual Meeting of th...
2022 doi
-
[18]
Man Luo, Xin Xu, Yue Liu, Panupong Pasupat, and Mehran Kazemi. 2024. https://openreview.net/forum?id=NQPo8ZhQPa In-context learning with retrieved demonstrations for language models: A survey . Transactions on Machine Learning Research. Survey Certification
2024
-
[19]
Ali Marashian, Enora Rice, Luke Gessler, Alexis Palmer, and Katharina von der Wense. 2025. https://aclanthology.org/2025.coling-main.472/ From priest to doctor: Domain adaptation for low-resource neural machine translation . In Proceedings of the 31st International Conference ...
2025
-
[20]
NLLB Team , Marta R. Costa-juss \`a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez,...
2024 doi
-
[21]
Alp Oktem, Mohamed Aymane Farhi, Brahim Essaidi, Naceur Jabouja, and Farida Boudichat. 2025. https://aclanthology.org/2025.wmt-1.82/ Correcting the tamazight portions of FLORES + and OLDI seed datasets . In Proceedings of the Tenth Conference on Machine Translation, pages 1072...
2025
-
[22]
Renhao Pei, Yihong Liu, Peiqin Lin, Fran c ois Yvon, and Hinrich Schuetze. 2025. https://doi.org/10.18653/v1/2025.acl-long.429 Understanding in-context machine translation for low-resource languages: A case study on M anchu . In Proceedings of the 63rd Annual Meeting of the As...
2025 doi
-
[23]
Maja Popovi \'c . 2017. https://doi.org/10.18653/v1/W17-4770 chr F ++: words helping character n-grams . In Proceedings of the Second Conference on Machine Translation, pages 612--618, Copenhagen, Denmark. Association for Computational Linguistics
2017 doi
-
[24]
Xiao Pu, Mingqi Gao, and Xiaojun Wan. 2023. Summarization is (almost) dead. arXiv preprint arXiv:2309.09558
2023 arXiv
-
[25]
Yush Rajcoomar. 2025. https://doi.org/10.18653/v1/2025.wmt-1.92 K oz K reol MRU WMT 2025 C reole MT system description: Koz kreol: Multi-stage training for E nglish -- mauritian creole MT . In Proceedings of the Tenth Conference on Machine Translation, pages 1183--1190, Suzhou...
2025 doi
-
[26]
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.213 COMET : A neural framework for MT evaluation . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685--2702, ...
2020 doi
-
[27]
Stephen Robertson and Hugo Zaragoza. 2009. https://doi.org/10.1561/1500000019 The probabilistic relevance framework: Bm25 and beyond . Found. Trends Inf. Retr., 3(4):333–389
2009 doi
-
[28]
Luis Frentzen Salim, Esteban Carlin, Alexandre Morinvil, Xi Ai, and Lun-Wei Ku. 2026. Beyond many-shot translation: Scaling in-context demonstrations for low-resource machine translation. arXiv preprint arXiv:2602.04764
2026
-
[29]
Hadia Mohmmedosman Ahmed Samil and David Ifeoluwa Adelani. 2026. https://openreview.net/forum?id=uLKTcetdkB Sudanese-flores: Extending FLORES + to sudanese arabic dialect . In 7th Workshop on African Natural Language Processing
2026
-
[30]
Garrett Tanzer, Mirac Suzgun, Eline Visser, Dan Jurafsky, and Luke Melas-Kyriazi. 2024. https://openreview.net/forum?id=tbVWug9f2h A benchmark for learning to translate a new language from one grammar book . In The Twelfth International Conference on Learning Representations
2024
-
[31]
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, and et al. 2025. https://arxiv.org/abs/2503.19786 Gemma 3 technical report . Preprint, arXiv:2503.19786
2025 arXiv
-
[32]
Inacio Vieira, Will Allred, S \'e amus Lankford, Sheila Castilho, and Andy Way. 2024. https://aclanthology.org/2024.amta-research.20/ How much data is enough data? fine-tuning large language models for in-house translation: Performance evaluation across multiple dataset sizes ...
2024
-
[33]
Biao Zhang, Barry Haddow, and Alexandra Birch. 2023. Prompting large language model for machine translation: A case study. In International conference on machine learning, pages 41092--41110. PMLR
2023
-
[34]
Chen Zhang, Xiao Liu, Jiuheng Lin, and Yansong Feng. 2024. https://doi.org/10.18653/v1/2024.findings-acl.519 Teaching Large Language Models an Unseen Language on the Fly . In Findings of the Association for Computational Linguistics : ACL 2024 , pages 8783--8800, Bangkok, Thai...
2024 doi
-
[35]
Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2024. https://doi.org/10.18653/v1/2024.findings-naacl.176 Multilingual machine translation with large language models: Empirical results and analysis . In Findings of the ...
2024 doi
-
[36]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[37]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.