REVIEW 4 major objections 6 minor 36 references
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
T0 review · 4 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read BOUQuET argues that machine translation evaluation should rest on originally written, multi-domain paragraphs in eight non-English pivot languages, extendable by crowd translation to every written language.
desk verdict A useful, well-released dataset whose main 'multicentric' claim is undercut by the fact that 87.5% of each pivot language is translated from English, not handcrafted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machine that carries the argument is Source-BOUQuET itself: a handcrafted parallel corpus in which each of eight linguist teams wrote 250 original sentences following explicit guidelines, so every text exists first in a non-English language and English is always a translation, not the origin. Each sentence belongs to a paragraph of three to six sentences and to one of eight domains (narration, dialogue, social-media posts, social-media comments, tutorials and how-to material, miscellaneous website content, reflection and opinion pieces, and a miscellaneous category), and is annotated with register features (connectedness, preparedness, social differential), linguistic coverage requirements (word order, verb morphology, agreement, gender, case, named entities, numbers, slang, emojis), and contextual notes about gender and formality. Those annotations let translators working from English recover information English does not encode, and they let the benchmark be analysed by domain and register using SONAR embeddings and Wasserstein distances. The open-initiative layer, consisting of contribution guidelines, an annotation tool, and a repository where any written language can be added by translating from one of the pivot languages or from English, is what turns the fixed dataset into a dynamic, community-extensible benchmark.
What would settle it
Ask independent native speakers to read unlabelled BOUQuET sentences and assign each to one of the eight domains; if their labels do not match the dataset's assigned domains at high agreement, the claim that the data represents natural register and domain use collapses.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that a dataset can be built as a multi-way parallel benchmark without starting from English or from scraped text: 250 sentences authored in each of eight pivot languages (Egyptian Arabic and Modern Standard Arabic, Mandarin Chinese, German, French, Hindi, Indonesian, Russian, Spanish), plus their English translations, expanded to 2,000 sentences per language by translating from English with paragraph-level context and additional annotations for grammatical gender, formality, and markedness. The result, the paper reports, ranks systems differently at paragraph level than at sentence level and gives higher automatic scores across 14 models, which the authors interpret as evidence that BOUQuET is easier for non-experts while still posing distinct challenges. The strongest claim is the contamination argument: because Source-BOUQuET was written from scratch, not mined, it is free from contamination in each initial state, and keeping one split hidden protects it after release.
Load-bearing premise
BOUQuET's value rests on the premise that handcrafted sentences written by eight linguists from guidelines genuinely represent natural language use across the eight domains and registers in every language; the paper does not externally validate that the elicited texts match real-world distributions of register or domain.
Editorial extensions
If this is right
- If the contamination claim is right, researchers can evaluate future MT and LLM systems on the hidden split with confidence that pre-training leakage has not inflated scores, and the same template could be applied to build contamination-free benchmarks in other NLP tasks.
- If the broader-domain claim is right, MT rankings on BOUQuET should generalise better to everyday translation use, including conversations, tutorials, social media, and website content, than rankings measured only on news or Wikipedia test sets.
- If the paragraph-level finding holds, translation quality must be judged with discourse context rather than sentence by sentence, since model rankings change when evaluation moves to paragraphs.
- If the pivot-language premise is right, translating BOUQuET into low-resource languages from Hindi, Spanish, Russian, French, Arabic, Indonesian, or Mandarin should be easier and more culturally apt than translating from English alone.
- The 55 completed languages at submission show that the commissioned pipeline is viable, and the open call is designed to extend that coverage to any written language.
Reading between the lines
- An unstated consequence is that BOUQuET-style handcrafting raises the bar for contamination audits: any benchmark could claim a clean initial state only if it publishes the original creation process, not just a hidden split.
- The paper does not analyse whether the eight linguists' output matches real corpora; a natural test is to ask independent native speakers to classify BOUQuET sentences into domains and compare their labels to the assigned ones, since author idiolect could otherwise masquerade as domain variation.
- The register annotations could support a diagnostic the paper does not run: measuring whether particular MT systems collapse informal registers such as dialogues and social comments into neutral formal prose, which would show up as systematic metric gaps within those domains.
- A testable extension would couple BOUQuET with human quality judgements on the same sentences, especially for pivot-to-low-resource pairs, to check whether the 'easier to translate' property measured by CometKiwi and MetricX matches human ease ratings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Source-BOUQuET, a manually constructed multi-way parallel machine-translation evaluation dataset in 8 non-English pivot languages plus English, organized into paragraphs spanning 8 domains and multiple registers, with a hidden evaluation split and an accompanying open crowdsourced translation initiative. The authors compare BOUQuET's domain coverage against FLORES+, NTREX-128, and NLLB-MD using SONAR embeddings and Wasserstein distances, and they provide a preliminary MT benchmark of 14 open-weight systems evaluated with CometKiwi and MetricX. The central claims are that BOUQuET offers broader domain/register representation, a non-English-centric and multicentric construction, contamination-free initial states, and a lower translation difficulty that makes it suitable for community extension.
Significance. If the central claims are validated, BOUQuET would be a valuable complement to FLORES/NTREX-type benchmarks, particularly for evaluating MT across everyday registers and for using non-English pivot languages in dataset expansion. The paper has concrete strengths: the dataset is released on Hugging Face, a test split is kept hidden for contamination control, detailed linguistic coverage requirements and paragraph-level quality checks are described, and the open-initiative infrastructure is genuinely useful. The original 250 sentences per pivot language, the English translations, and the QA documentation are all positive contributions. However, the advertised non-English-centric property is only established for a small subset of the content per language, and the quantitative evidence for broader domain coverage and lower translation difficulty lacks uncertainty quantification. These issues affect the paper's central differentiators and need to be addressed before the claims can be accepted as stated.
major comments (4)
- [Section 3.4 vs. Abstract and Section 3.1] The abstract states that BOUQuET is "handcrafted in 8 non-English languages," and Section 3.1 describes a "non-English-centric focus" in which the dataset is "handcrafted by proficient speakers" of those languages. Section 3.4, however, says that only 250 sentences per pivot language are originally authored in that language and that the remaining 1,750 sentences per pivot language are translated from English. This means 87.5% of each pivot language's content is not originally authored in that language and is therefore exposed to translationese effects such as calqued word order, over-explicit pronouns, and reduced target-specific register features. The "multicentric" characterization is fully supported only for the 250-sentence seed subset per language, not for the full 2,000-sentence dataset. The authors should either revise the wording to distinguish the 250-sentence seed set from the translated expansion, or provide per-language validation that the 1,750 translated sentences are representative of natural register and domain usage in each pivot language; the current Section 4 domain analysis, which compares against English-domain reference corpora, does not provide that validation.
- [Section 4, Figure 3] The claim that BOUQuET has "a broader representation of domains" relative to FLORES+, NTREX-128, and NLLB-MD rests on Wasserstein distances computed after downsampling each set to 2,000 sentences. The paper does not report confidence intervals, repeated sampling, or statistical tests, even though the domain reference sets are large and the sample size is finite. A single point estimate per domain cannot establish the claim of "lowest consistent results" in Figure 3. Please add bootstrap or permutation confidence intervals, or report repeated-sample ranges, and state how many random draws were used for the 2,000-sentence subsamples.
- [Section 4, Table 4] The conclusion that BOUQuET is "easier to translate" than FLORES+ and NTREX-128 is based on averaged CometKiwi and MetricX scores in Table 4 without standard errors, confidence intervals, or significance tests. The paper itself calls the MT benchmark preliminary, so the text should frame the difficulty difference as a hypothesis rather than a demonstrated property. In addition, the ranking-swap analysis uses only CometKiwi and paragraph-level results are reported without explaining how sentence-level scores are aggregated into paragraph scores; this should be clarified so the swaps and Pearson correlations are reproducible.
- [Section 5 and Table 6] The paper states that BOUQuET "includes 55 multi-way parallel completed languages" at the time of submission, but Section 3 describes the construction methodology only for the 9-language Source-BOUQuET (8 pivots plus English). It is unclear whether all 55 languages in Table 6 have the full 2,000-sentence parallel coverage, which pivot language served as the source for each, and whether the QA strategies in Appendix C were applied to all 55. Since the multi-way parallel coverage is a central quantitative claim of the paper, the status of these 55 languages and their release metadata should be specified explicitly.
minor comments (6)
- [Section 3.4] The paragraph explaining that "English isn't an ideal source language" sits awkwardly next to the statement that 1,750 of 2,000 sentences per pivot language are translated from English; the paper should address this tension directly.
- [Section 3.5 and Table 1] Section 3.5 reports the evaluation set as 632 sentences / 144 paragraphs, while Table 1 reports the BOUQuET Eval split as 144 paragraphs and 628 sentences; these numbers should be reconciled.
- [Table 4] The column header "BOUQUET P N TREX P" appears garbled and should be reformatted to clearly separate the paragraph-level BOUQuET, FLORES, and NTREX columns.
- [Figure 6 caption] The caption says "from top to down" and should be "from top to bottom."
- [Section 4] The phrase "Rankings is computed" should be "Rankings are computed," and the paper should specify whether the Pearson correlations are computed on sentence-level or paragraph-level CometKiwi scores.
- [General] The dataset name is spelled inconsistently as BOUQuET and BOUQUET across the title, Table 1, and several section headings; please standardize the spelling.
Circularity Check
No load-bearing circularity: dataset construction is described in-paper, MT and domain benchmarks use external metrics, and self-cited baselines (FLORES+, SONAR) are not used to force the central claims.
full rationale
The paper's central claims—that Source-BOUQuET is handcrafted, multicentric, contamination-free, broader in domain coverage, and easier to translate—are supported by the in-paper description of the creation process (Section 3.2, Appendix A), not by a derivation from a fitted model. The MT benchmark (Section 4, Tables 4 and 9) uses the external metrics CometKiwi and MetricX, which were not involved in dataset construction, and the ranking comparisons against FLORES+ and NTREX-128 are independently computable, so no prediction is forced by construction. The domain-representation analysis (Section 4) measures Wasserstein distance between SONAR embeddings of BOUQuET and public domain corpora; those same corpora were used as statistical guidance for sentence/paragraph length and CEFR complexity before creation (Section 3.2), so the low WD values partly reflect design intent rather than surprise, but the measured quantity (embedding-space proximity) is distinct from the guided quantities (length statistics, CEFR levels), and the comparison is against fixed external alternatives, so the demonstration is not circular. Self-citations (FLORES-101/200, FLORES+, NLLB-MD, 2M-FLORES, SONAR) serve as baselines and analysis tools with overlapping authorship, but no central premise depends on an unverified self-citation; these prior datasets are independently released and externally reproducible. One non-circular but notable concern: Section 3.4 states that only 250 of 2,000 sentences per pivot language were originally authored in that language, with 1,750 translated from English, which qualifies (but does not falsify) the abstract's 'handcrafted in 8 non-English languages' and 'multicentric' phrasing; this is a claim-adequacy/correctness risk, not a circular derivation. No equation-level reduction of a prediction to its inputs was found, so the circularity score is low.
Assumptions & free parameters
assumptions (3)
- domain assumption SONAR embeddings capture domain-relevant semantic similarity
- domain assumption CEFR labels from a SONAR-based model are a valid proxy for linguistic complexity
- domain assumption The eight selected domains cover the practical space of translation needs
Cite this review
Pith. "Pith review of BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation." pith.science (2026). https://pith.science/paper/DIS4KF3B
@misc{pith2026250204314,
author = {Pith},
title = {Pith review of: BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation},
year = {2026},
howpublished = {\url{https://pith.science/paper/DIS4KF3B}},
note = {Machine review of arXiv:2502.04314}
}
read the original abstract
BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages. Each of these source languages are representative of the most widely spoken ones and therefore they have the potential to serve as pivot languages that will enable more accurate translations. The dataset is multicentric to enforce representation of multilingual language features. In addition, the dataset goes beyond the sentence level, as it is organized in paragraphs of various lengths. Compared with related machine translation datasets, we show that BOUQuET has a broader representation of domains while simplifying the translation task for non-experts. Therefore, BOUQuET is specially suitable for crowd-source extension for which we are launching a call aiming at collecting a multi-way parallel corpus covering any written language.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng X...
arXiv 2023
-
[4]
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020. https://arxiv.org/abs/2001.08435 The pushshift reddit dataset . CoRR, abs/2001.08435
arXiv 2020
-
[5]
Yulong Chen, Yang Liu, and Yue Zhang. 2021. https://aclanthology.org/2021.inlg-1.33 D ialog S um challenge: Summarizing real-life scenario dialogues . In Proceedings of the 14th International Conference on Natural Language Generation, pages 308--313, Aberdeen, Scotland, UK. Association for Computational Linguistics
work page 2021
-
[6]
Everlyn Chimoto and Bruce Bassett. 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.348 COMET - QE and active learning for low-resource machine translation . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 4735--4740, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics
-
[7]
Marta R. Costa-jussà, Bokai Yu, Pierre Andrews, Belen Alastruey, Necati Cihan Camgoz, Joe Chuang, Jean Maillard, Christophe Ropers, Arina Turkantenko, and Carleigh Wood. 2024. https://arxiv.org/abs/2412.08274 2m-belebele: Highly multilingual speech and american sign language comprehension dataset . Preprint, arXiv:2412.08274
arXiv 2024
-
[8]
Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A Rossi, and Thien Huu Nguyen. 2023. Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback. arXiv e-prints, pages arXiv--2307
2023
Show all 36 references
-
[9]
John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, Sandra Kublik, Meor Amer, Viraat Aryabumi, Jon Ander Campos, Yi-Chern Tan, Tom Kocmi, Florian Strub, Nathan Grinsztaj...
2024 arXiv
-
[10]
Daniel Deutsch, Eleftheria Briakou, Isaac Caswell, Mara Finkelstein, Rebecca Galor, Juraj Juraska, Geza Kovacs, Alison Lui, Ricardo Rei, Jason Riesa, Shruti Rijhwani, Parker Riley, Elizabeth Salesky, Firas Trabelsi, Stephanie Winkler, Biao Zhang, and Markus Freitag. 2025. http...
2025 arXiv
-
[11]
Paul-Ambroise Duquenne, Holger Schwenk, and Benoît Sagot. 2023. https://arxiv.org/abs/2308.11466 Sonar: Sentence-level multimodal and language-agnostic representations . Preprint, arXiv:2308.11466
2023 arXiv
-
[12]
Eberhard, Gary F
David M. Eberhard, Gary F. Simons, and Charles D. Fennig. 2024. https://www.ethnologue.com/insights/how-many-languages/ Ethnologue: Languages of the world. twenty-seventh edition . https://www.ethnologue.com/. Last accessed on 2025-02-03
2024
-
[13]
Christian Federmann, Tom Kocmi, and Ying Xin. 2022. https://doi.org/10.18653/v1/2022.sumeval-1.4 NTREX -128 -- news test references for MT evaluation of 128 languages . In Proceedings of the First Workshop on Scaling Up Multilingual Evaluation, pages 21--24, Online. Associatio...
2022 doi
-
[14]
Martin Gerlach and Francesc Font - Clos. 2018. https://arxiv.org/abs/1812.08092 A standardized project gutenberg corpus for statistical analysis of natural language and quantitative linguistics . CoRR, abs/1812.08092
2018 arXiv
-
[15]
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc ' Aurelio Ranzato, Francisco Guzm \'a n, and Angela Fan. 2022. https://doi.org/10.1162/tacl_a_00474 The F lores-101 evaluation benchmark for low-resource and multilingua...
2022 doi
-
[16]
M. A. K. Halliday and C. M. I. Matthiessen. 2004. An Introduction to Functional Grammar. Routledge
2004
-
[17]
Juraj Juraska, Daniel Deutsch, Mara Finkelstein, and Markus Freitag. 2024. https://doi.org/10.18653/v1/2024.wmt-1.35 M etric X -24: The G oogle submission to the WMT 2024 metrics shared task . In Proceedings of the Ninth Conference on Machine Translation, pages 492--504, Miami...
2024 doi
-
[18]
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, ...
2024 doi
-
[19]
Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat
Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A. Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2023. https://arxiv.org/abs/2309.04662 Madlad-400: A multilingual and document-level large audited ...
2023 arXiv
-
[20]
William Labov. 1991. Sociolinguistic patterns. University of Pennsylvania Press
1991
-
[21]
Wing Lian, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet Vong, and "Teknium". 2023. Openorca: An open dataset of gpt augmented flan reasoning traces. https://https://huggingface.co/Open-Orca/OpenOrca
2023
-
[22]
Jean Maillard, Laurie Burchell, Antonios Anastasopoulos, Christian Federmann, Philipp Koehn, and Skyler Wang. 2024. https://doi.org/10.18653/v1/2024.wmt-1.4 Findings of the WMT 2024 shared task of the open language data initiative . In Proceedings of the Ninth Conference on Ma...
2024 doi
-
[23]
Shervin Malmasi, Anjie Fang, Besnik Fetahu, Sudipta Kar, and Oleg Rokhlenko. 2022. https://aclanthology.org/2022.coling-1.334 M ulti C o NER : A large-scale multilingual dataset for complex named entity recognition . In Proceedings of the 29th International Conference on Compu...
2022
-
[24]
Guerreiro, Ricardo Rei, Duarte M
Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, and André F. T. Martins. 2024. https://a...
2024 arXiv
-
[25]
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, C a g lar Gu l c ehre, and Bing Xiang. 2016. https://doi.org/10.18653/v1/K16-1028 Abstractive text summarization using sequence-to-sequence RNN s and beyond . In Proceedings of the 20th SIGNLL Conference on Computational Natural...
2016 doi
-
[26]
Cohen, and Mirella Lapata
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. https://doi.org/10.18653/v1/D18-1206 Don`t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization . In Proceedings of the 2018 Conference on Empirical Methods in Natura...
2018 doi
-
[27]
NLLBTeam. 2024. https://doi.org/10.1038/s41586-024-07335-x Scaling neural machine translation to 200 languages . Nature, 630:841–846
2024 doi
-
[28]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(1)
2020
-
[29]
Guerreiro, Jo \ a o Alves, Pedro Henrique Martins, Patrick Fernandes, Helena Wu, Tania Vaz, Duarte Alves, Amin Farajian, Sweta Agrawal, Antonio Farinhas, Jos \'e G
Ricardo Rei, Jose Pombal, Nuno M. Guerreiro, Jo \ a o Alves, Pedro Henrique Martins, Patrick Fernandes, Helena Wu, Tania Vaz, Duarte Alves, Amin Farajian, Sweta Agrawal, Antonio Farinhas, Jos \'e G. C. De Souza, and Andr \'e Martins. 2024. https://doi.org/10.18653/v1/2024.wmt-...
2024 doi
-
[30]
Oscar Sainz, Jon Campos, Iker Garc \'i a-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.722 NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark . In Findings of the ...
2023 doi
-
[31]
Shivalika Singh, Freddie Vargus, Daniel D ' souza, B \"o rje Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O ' Mahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Moura, Dominik Krzemi \'n ski, Haki...
2024
-
[32]
Shuo Sun and Kevin Duh. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.340 CLIRM atrix: A massively large collection of bilingual and multilingual datasets for cross-lingual information retrieval . In Proceedings of the 2020 Conference on Empirical Methods in Natural Langua...
2020 doi
-
[33]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://arxiv.org/abs/2302.13971 Llama:...
2023 arXiv
-
[34]
Minghao Wu, Weixuan Wang, Sinuo Liu, Huifeng Yin, Xintong Wang, Yu Zhao, Chenyang Lyu, Longyue Wang, Weihua Luo, and Kaifu Zhang. 2025. https://arxiv.org/abs/2504.15521 The bitter lesson learned from 2,000+ multilingual benchmarks . Preprint, arXiv:2504.15521
2025 arXiv
-
[35]
Xinyan Yu, Trina Chatterjee, Akari Asai, Junjie Hu, and Eunsol Choi. 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.273 Beyond counting datasets: A survey of multilingual dataset construction and necessary resources . In Findings of the Association for Computational Lin...
2022 doi
-
[36]
Yiran Zhao, Chaoqun Liu, Yue Deng, Jiahao Ying, Mahani Aljunied, Zhaodonghui Li, Lidong Bing, Hou Pong Chan, Yu Rong, Deli Zhao, and Wenxuan Zhang. 2025. https://arxiv.org/abs/2503.00865 Babel: Open multilingual large language models serving over 90\ Preprint, arXiv:2503.00865
2025 arXiv
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.