REVIEW 3 major objections 5 minor 63 references
ChiKhaPo uses eight word-level subtasks to test LLM comprehension and generation in more than 2,700 languages, and finds that current models have near-zero lexical competence in most of them.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
ChiKhaPo is an 8-subtask benchmark that measures word-level comprehension and generation in 2,700+ languages and shows state-of-the-art models perform poorly on low-resource languages.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection Massive coverage, genuine resource; but unvalidated PanLex gold labels and per-model prompts make scores conditional as a measurement instrument. the 3 major comments →
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that word-level lexical competence—understanding and producing individual words—can be measured on a massively multilingual scale using existing resources, and that doing so exposes a large gap in current LLMs. ChiKhaPo scores every target-language word as 0 or 1 (or a probability in the conditioned-language-modeling task) in four task settings and two directions. Evaluated on six open-weight models, the benchmark finds that comprehension of a word in a low-resource language is consistently easier for models than generating that word, that scores rise roughly logarithmically with resource availability, and that word-translation scores track sentence-level MT scor
What carries the argument
The central object is a per-word score rather than a sentence-level metric. A bilingual lexicon supplies the reference translation or translations; a model output is credited if it matches a reference through exact match, inflection, substring, or synonymy heuristics. Four tasks elicit that match in different settings: direct word translation, word translation with sentence context, next-word generation probability in a translation-conditioned language model, and word presence in a sentence-level translation. This design lets the benchmark cover languages that only have word lists, while the sentence-level tasks add harder, resource-dependent layers.
Load-bearing premise
The load-bearing premise is that the word-list translations used as correct answers are accurate ground truth for thousands of low-resource languages—if many entries are wrong, low scores may reflect bad labels rather than missing competence, a limitation the paper states explicitly.
What would settle it
Have native speakers verify a random sample of the gold translations used in the word-translation task for low-resource languages; if a substantial fraction are wrong, the benchmark's absolute scores are not valid measures of lexical competence.
If this is right
- With 2,700+ languages covered in the word-translation task, basic lexical competence can now be tracked for nearly all written languages, not just a few dozen.
- Word-translation scores correlate strongly with sentence-level MT scores, so ChiKhaPo can serve as a cheap evaluation proxy in the absence of translation data.
- The consistent comprehension-over-generation gap identifies generation into low-resource languages as the more urgent target for improvement.
- The roughly logarithmic relationship between resourcedness and score suggests that mid-resource languages may improve quickly with modest data, while very low-resource languages need a different approach.
- The soft translation-conditioned language-modeling task gives model developers a per-checkpoint signal during training, though the paper cautions it is not comparable across models.
Where Pith is reading between the lines
- A speaker-verified subset of gold translations would test whether low scores reflect model failure or label noise; the paper's own limitation section says benchmark quality is bounded by lexicon annotations.
- The strong word-level/MT correlation suggests a testable hypothesis: improving word-level lexical recall in low-resource languages should transfer to sentence-level MT quality.
- The benchmark's four tasks could be extended with morphological inflection and word-order checks once resources exist, directly addressing the out-of-scope skills the paper names.
- Practitioners could use WT scores to prioritize which of the languages without translation benchmarks deserve parallel-data collection or human evaluation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ChiKhaPo, a multilingual benchmark for word-level lexical comprehension and generation, with 8 subtasks across four task families (word translation, word translation with context, translation-conditioned language modeling, and bag-of-words MT). The benchmark is assembled from existing resources: GATITOS, IDS, and PanLex lexicons; GLOTLID monolingual data; and FLORES+ bitext. The authors report coverage of 2,746 languages for word translation and 525 for word translation with context, evaluate six open-weight LLMs, and analyze score variation by language family, resource level, and evaluation direction. They also report a correlation between WT scores and FLORES+ BLEU scores, proposing WT as a cheap proxy for MT quality in languages without bitext.
Significance. If the benchmark is valid, it fills a real gap: no existing evaluation covers lexical competence at this language scale. The resource is useful in principle, and the authors have released dataset, code, and a Python package, which strengthens reproducibility. The finding that models score near zero on many low-resource languages is important if trustworthy. However, the benchmark's value depends on the quality of the underlying lexicons, especially PanLex, which provides most WT entries. The paper's own limitations section acknowledges that lexicon quality constrains the benchmark, and the manual evaluation in Appendix E.2 does not validate the lexicon labels themselves. The central claim to 'measure basic lexical comprehension' is therefore currently supported only conditionally.
major comments (3)
- [§3.1.1, §3.2, Appendix D (Table 8), §7] The WT task, which underlies the headline 2,746-language coverage and the correlation analysis in §5, scores a model only if its output matches some entry in Ξ(w). The vast majority of entries come from PanLex (5,731 language pairs vs. 177 for GATITOS and 240 for IDS). PanLex is an automatically harvested, unvalidated resource, and §7 concedes that 'the quality of the benchmark is restricted by the available annotations in the lexicons' and that valid variants may be missing. The manual evaluation in Appendix E.2 (Table 13) reports low false-positive and false-negative rates, but that check only measures agreement between the string-matching heuristics and the lexicon; it does not validate that the lexicon entries are correct translations. Consequently, low model scores on low-resource languages may reflect erroneous or incomplete gold labels rather than a lack of lexical competence. The
- [§3.1.2, §7] WTWC is designed to test word comprehension in context, but the accepted answer set Ξ(w) is not sense-annotated. The evaluation averages a binary correctness variable over occurrences of w (defined in §3.1.2); if the same word appears in contexts requiring different senses, any lexicon equivalent that is correct for one sense is accepted for all contexts. The paper itself states in §7 that 'our lexicons also do not annotate word sense.' This makes WTWC scores a mixture of lexical access and sense disambiguation, with no way to separate the two. The authors should quantify sense ambiguity in the WTWC data: how many target words have multiple senses in the sampled contexts, and how often a contextually wrong but lexically plausible answer is scored correct. Without such numbers, the WTWC scores are hard to interpret as context-sensitive lexical competence.
- [§3.1.3, TCLM X→model] The TCLM X→model score averages over F, the union of the lexicon equivalents Ξ(wX(m)) and the FastAlign alignments A(wX(m)). If a lexicon contains many equivalents, or if alignments add extra English words, the denominator |F| grows and the score for a given target word is diluted even when the model assigns high probability to the aligned English word. This makes the TCLM X→model score dependent on the completeness of the same unvalidated lexicons that drive the WT concern. The paper should report the distribution of |F| and consider a variant that takes the maximum probability over F rather than the mean, or otherwise demonstrate that the averaging is not responsible for the observed language-level patterns.
minor comments (5)
- [Appendix E.2 / Table 13] The manual evaluation is extremely small relative to the scale of the benchmark: 283 WT samples, 121 WTWC samples, and 229 BOW MT samples, with 'at least 10 responses' per model-direction pair. The paper should state the number of languages and models actually annotated and give per-language breakdowns, since a single accidentally easy or hard language can dominate the reported rates.
- [§5, Figure 5] The text reports correlations of 0.873 and 0.769 between WT and BLEU but does not say whether these are Pearson r or Spearman rho, nor does it report confidence intervals or a fitted equation. Given that the regression is used to justify WT as a proxy for MT, the details should be specified.
- [Appendix F] Prompt exploration was carried out in Spanish only, and different prompts were assigned to different models. Since the benchmark spans 2,700+ languages, the prompt choices may not be equally appropriate across scripts, formality levels, or instruction-following behaviors. A brief discussion of this limitation would be helpful.
- [Appendix G.1] Typo: 'We trained a language a decision tree regressor' should be 'We trained a decision tree regressor.'
- [Appendix E.1.1] The text reads 'fuzzywuzzy 5 similarity score' — the stray '5' appears to be a typo, and the threshold value (75) should be stated clearly.
Circularity Check
No circularity: the benchmark is assembled from external resources and model scores are empirical measurements; the WT–BLEU correlation is a post-hoc fit, not fed back into the benchmark.
full rationale
I traced the paper's derivation chain. ChiKhaPo is constructed from external resources: lexicons (GATITOS, IDS, PanLex), monolingual data (GLOTLID), and bitext (FLORES+). Model scores are computed by matching model outputs against lexicon entries; no model output, fitted parameter, or measured score is fed back into the construction of the benchmark or into the definition of the scores. The claim that six models 'struggle' on low-resource languages is an empirical measurement, not a consequence of the benchmark's definitional setup: scores vary substantially across languages and models (e.g., Spanish WT model→X reaches 68.3% for bloomz-7b1-mt, while Mossi stays near 0%). The WT–BLEU correlation in Section 5 is a post-hoc regression over independently computed FLORES+ BLEU scores and ChiKhaPo WT scores; it is not used to define ChiKhaPo scores and therefore is not a fitted-input-called-prediction. The decision-tree feature importance analysis is descriptive rather than a load-bearing derivation. No self-citation chain is invoked as evidence; the paper's supporting references are external resources. The acknowledged limitation in Section 7 — 'The quality of the benchmark is restricted by the available annotations in the lexicons we work with' — is a validity threat about noisy gold labels, not a circularity: an inaccurate external lexicon would bias model scores but would not make the benchmark's score definitions reduce to their inputs. For the same reason, the Appendix E manual evaluation, which checks the string-matching heuristics against the lexicon, does not by itself validate lexicon accuracy; that is a quality concern, not a circular derivation. Overall, the benchmark's central claims are empirical and externally grounded, so the circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (2)
- fuzzywuzzy inflection threshold =
75 (similarity score)
- sampling caps =
300 words (WT/WTWC), 30% of parallel data (TCLM/BOW MT), min 100 words/language
axioms (4)
- domain assumption Ability to translate a word into English is a valid proxy for lexical comprehension of that word.
- domain assumption Lexicon entries from PanLex/IDS/GATITOS are correct ground-truth translations.
- domain assumption FLORES+ devtest parallel sentences are valid stimuli for word-level lexical probing.
- domain assumption English WordNet synonyms are acceptable correct answers in the X→model direction.
Cite this review
Pith. "Pith review of ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models." pith.science (2026). https://pith.science/paper/R2WUHK7Z
@misc{pith2026251016928,
author = {Pith},
title = {Pith review of: ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/R2WUHK7Z}},
note = {Machine review of arXiv:2510.16928}
}
read the original abstract
Existing benchmarks for large language models (LLMs) are largely restricted to high- or mid-resource languages, and often evaluate performance on higher-order tasks in reasoning and generation. However, plenty of evidence points to the fact that LLMs lack basic linguistic competence in the vast majority of the world's 3800+ written languages. We introduce ChiKhaPo, consisting of 8 subtasks of varying difficulty designed to evaluate the lexical comprehension and generation abilities of generative models. ChiKhaPo draws on existing lexicons, monolingual data, and bitext, and provides coverage for 2700+ languages for 2 subtasks, surpassing any existing benchmark in terms of language coverage. We further show that 6 SOTA models struggle on our benchmark, and discuss the factors contributing to performance scores, including language family, language resourcedness, task, and comprehension versus generation directions. With ChiKhaPo, we hope to enable and encourage the massively multilingual benchmarking of LLMs.
Figures
Reference graph
Works this paper leans on
-
[1]
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Maxamed Axmed, and 1 others. 2023. Mega: Multilingual evaluation of generative ai. arXiv preprint arXiv:2303.12528
Pith/arXiv arXiv 2023
-
[2]
Alham Fikri Aji, Genta Indra Winata, Fajri Koto, Samuel Cahyawijaya, Ade Romadhony, Rahmad Mahendra, Kemal Kurniawan, David Moeljadi, Radityo Eko Prasojo, Timothy Baldwin, Jey Han Lau, and Sebastian Ruder. 2022. https://doi.org/10.18653/v1/2022.acl-long.500 One country, 700+ languages: NLP challenges for underrepresented languages and dialects in I ndones...
-
[3]
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. https://arxiv.org/abs/2311.16867 The falcon series of open language models . Preprint, arXiv:2...
Pith/arXiv arXiv 2023
-
[4]
Sotiris Anagnostidis and Jannis Bulian. 2024. https://arxiv.org/abs/2408.11865 How susceptible are llms to influence in prompts? Preprint, arXiv:2408.11865
Pith/arXiv arXiv 2024
-
[5]
Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, Kelly Marchisio, Max Bartolo, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Aidan Gomez, Phil Blunsom, Marzieh Fadaee, and 2 others. 2024. https://arxiv.org/abs/2405.15032 Aya 23: Open weigh...
Pith/arXiv arXiv 2024
-
[6]
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. https://aclanthology.org/2024.acl-long.44 The belebele benchmark: a parallel reading comprehension dataset in 122 language variants . In Proceedings of the 62nd Annual Meeting of the ...
2024
-
[7]
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017. What do neural machine translation models learn about morphology? arXiv preprint arXiv:1704.03471
Pith/arXiv arXiv 2017
-
[8]
Akhiad Bercovich, Itay Levy, Izik Golan, Mohammad Dabbah, Ran El-Yaniv, Omri Puny, Ido Galil, Zach Moshe, Tomer Ronen, Najeeb Nabwani, and 1 others. 2025. Llama-nemotron: Efficient reasoning models. arXiv preprint arXiv:2505.00949
arXiv 2025
-
[9]
intercontinental dictionary series
Hans-Jörg Bibiko. 2023. Cldf dataset derived from key and comrie's "intercontinental dictionary series" from 2023
2023
-
[10]
Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen
Mingyang Chen, Tianpeng Li, Haoze Sun, Yijie Zhou, Chenzheng Zhu, Haofen Wang, Jeff Z. Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen. 2025. https://arxiv.org/abs/2503.19470 Research: Learning to reason with search for llms via reinforcement learning . Preprint, arXiv:2503.19470
Pith/arXiv arXiv 2025
-
[11]
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. https://doi.org/10.1162/tacl_a_00317 T y D i QA : A benchmark for information-seeking question answering in typologically diverse languages . Transactions of the Association for Computational Linguistics, 8:454--470
-
[12]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168
Pith/arXiv arXiv 2021
-
[13]
Bowman, Holger Schwenk, and Veselin Stoyanov
Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. https://arxiv.org/abs/1809.05053 Xnli: Evaluating cross-lingual sentence representations . Preprint, arXiv:1809.05053
Pith/arXiv arXiv 2018
-
[14]
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm Free dolly: Introducing the world's first truly open instruction-tuned llm . Accessed: 2023-06-30
2023
-
[15]
John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, Sandra Kublik, Meor Amer, Viraat Aryabumi, Jon Ander Campos, Yi-Chern Tan, Tom Kocmi, Florian Strub, Nathan Grinsztajn, Yannis Flet-Berliac, and 26 others. 2024. https://arxiv.org/abs/2412.04261 Ay...
Pith/arXiv arXiv 2024
-
[16]
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, and 181 others. 2025. https://arxiv.org/abs/2501.12948 Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement lea...
Pith/arXiv arXiv 2025
-
[17]
Chris Dyer, Victor Chahuneau, and Noah A Smith. 2013. A simple, fast, and effective reparameterization of ibm model 2. In Proceedings of the 2013 conference of the North American chapter of the association for computational linguistics: human language technologies, pages 644--648
2013
-
[18]
Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Gim \'e nez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, and Katharina Kann. 2022. https://doi.org/10.18653/v1/2022.acl-long.435 A mericas NLI : Evalu...
-
[19]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
Pith/arXiv arXiv 2024
-
[20]
Harald Hammarström, Robert Forkel, Martin Haspelmath, and Sebastian Bank. 2025. https://doi.org/10.5281/zenodo.15525265 Glottolog 5.2 . Available online at http://glottolog.org, Accessed on 2025-09-25
-
[21]
Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. https://aclanthology.org/2021.findings-acl.413 XL -sum: Large-scale multilingual abstractive summarization for 44 languages . In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693...
2021
-
[22]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300
Pith/arXiv arXiv 2020
-
[23]
Ayyoob ImaniGooghari, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, Andr \'e Martins, Fran c ois Yvon, and Hinrich Sch \"u tze. 2023. https://aclanthology.org/2023.acl-long.61 Glot500: Scaling multilingual corpora and language models to 500 languages . In Proceedings of the 61st Annual Me...
2023
-
[24]
Vivek Iyer, Pinzhen Chen, and Alexandra Birch. 2023. Towards effective disambiguation for machine translation with large language models. arXiv preprint arXiv:2309.11668
Pith/arXiv arXiv 2023
-
[25]
Alexander Jones, Isaac Caswell, Orhan Firat, and Ishank Saxena. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.26 GATITOS : Using a new multilingual lexicon for low-resource machine translation . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 371--405, Singapore. Association for Computational Linguistics
-
[26]
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the nlp world. arXiv preprint arXiv:2004.09095
Pith/arXiv arXiv 2020
-
[27]
David Kamholz, Jonathan Pool, and Susan Colowick. 2014. http://www.lrec-conf.org/proceedings/lrec2014/pdf/1029_Paper.pdf P an L ex: Building a resource for panlingual lexical translation . In Proceedings of the Ninth International Conference on Language Resources and Evaluation ( LREC '14) , pages 3145--3150, Reykjavik, Iceland. European Language Resource...
2014
-
[28]
Akshara Kandimalla, Pintu Lohar, Souvik Kumar Maji, and Andy Way. 2022. Improving english-to-indian language neural machine translation systems. Information, 13(5):245
2022
-
[29]
Amir Kargaran, Ayyoob Imani, François Yvon, and Hinrich Schuetze. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.410 Glotlid: Language identification for low-resource languages . In Findings of the Association for Computational Linguistics: EMNLP 2023, page 6155–6218. Association for Computational Linguistics
-
[30]
Gerhard Kremer, Katrin Erk, Sebastian Pad \'o , and Stefan Thater. 2014. https://doi.org/10.3115/v1/E14-1057 What substitutes tell us - analysis of an all-words lexical substitution corpus . In Proceedings of the 14th Conference of the E uropean Chapter of the Association for Computational Linguistics , pages 540--549, Gothenburg, Sweden. Association for ...
-
[31]
Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.360 W iki L ingua: A new benchmark dataset for cross-lingual abstractive summarization . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4034--4048, Online. Association for Computational Linguistics
-
[32]
Mina Lee, Chris Donahue, Robin Jia, Alexander Iyabor, and Percy Liang. 2021. https://doi.org/10.18653/v1/2021.naacl-main.345 Swords: A benchmark for lexical substitution with improved data coverage and quality . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies...
-
[33]
Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. https://doi.org/10.18653/v1/2020.acl-main.653 MLQA : Evaluating cross-lingual extractive question answering . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7315--7330, Online. Association for Computational Linguistics
-
[34]
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cosgrove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew Arad Hudson, and 31 others. 2023. https://openreview.net/forum?i...
2023
-
[35]
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona T. Diab, and 2 others. 2021. https://arxiv.org/abs/2112.10668 Few-shot learnin...
Pith/arXiv arXiv 2021
-
[36]
Yile Liu, Ziwei Ma, Xiu Jiang, Jinglu Hu, Jing Chang, and Liang Li. 2025. Maxife: Multilingual and cross-lingual instruction following evaluation. arXiv preprint arXiv:2506.01776
Pith/arXiv arXiv 2025
-
[37]
Gonzalo Martínez, Javier Conde, Elena Merino-Gómez, Beatriz Bermúdez-Margaretto, José Alberto Hernández, Pedro Reviriego, and Marc Brysbaert. 2024. https://doi.org/10.1371/journal.pone.0308259 Establishing vocabulary tests as a benchmark for evaluating large language models . PLOS ONE, 19(12):1--17
-
[38]
Diana McCarthy. 2002. https://doi.org/10.3115/1118675.1118691 Lexical substitution as a task for WSD evaluation . In Proceedings of the ACL -02 Workshop on Word Sense Disambiguation: Recent Successes and Future Directions , pages 089--115. Association for Computational Linguistics
arXiv 2002
-
[39]
Diana McCarthy and Roberto Navigli. 2007. https://aclanthology.org/S07-1009/ S em E val-2007 task 10: E nglish lexical substitution task . In Proceedings of the Fourth International Workshop on Semantic Evaluations ( S em E val-2007) , pages 48--53, Prague, Czech Republic. Association for Computational Linguistics
2007
-
[40]
Rada Mihalcea, Ravi Sinha, and Diana McCarthy. 2010. https://aclanthology.org/S10-1002/ S em E val-2010 task 2: Cross-lingual lexical substitution . In Proceedings of the 5th International Workshop on Semantic Evaluation, pages 9--14, Uppsala, Sweden. Association for Computational Linguistics
2010
-
[41]
George A. Miller. 1994. https://aclanthology.org/H94-1111/ W ord N et: A lexical database for E nglish . In H uman L anguage T echnology: Proceedings of a Workshop held at P lainsboro, N ew J ersey, M arch 8-11, 1994
1994
-
[42]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023. https://arxiv.org/abs/2211.01786 Crosslingual gene...
Pith/arXiv arXiv 2023
-
[43]
NLLB Team , Marta R. Costa-juss \`a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, and 20 others. 2024. https://doi.org/10.1038/s41586-024-073...
-
[44]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311--318
2002
-
[45]
Bolette Pedersen, Nathalie S rensen, Sussi Olsen, Sanni Nimb, and Simon Gray. 2024. https://aclanthology.org/2024.lrec-main.1421/ Towards a D anish semantic reasoning benchmark - compiled from lexical-semantic resources for assessing selected language understanding capabilities of large language models . In Proceedings of the 2024 Joint International Conf...
2024
-
[46]
Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen
Edoardo M. Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020. https://ducdauge.github.io/files/xcopa.pdf XCOPA: A multilingual dataset for causal commonsense reasoning . arXiv preprint
2020
-
[47]
Maja Popovi \'c . 2015. chrf: character n-gram f-score for automatic mt evaluation. In Proceedings of the tenth workshop on statistical machine translation, pages 392--395
2015
-
[48]
James Pustejovsky. 2016. Lexical semanics, page 33–64. Cambridge Handbooks in Language and Linguistics. Cambridge University Press
2016
-
[49]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025. https://arxiv.org/abs/2412.15115 Qwen2.5 technical report . Preprint, arXiv:2412.15115
Pith/arXiv arXiv 2025
-
[50]
Alessandro Raganato, Yves Scherrer, and J \"o rg Tiedemann. 2019. https://doi.org/10.18653/v1/W19-5354 The M u C o W test suite at WMT 2019: Automatically harvested multilingual contrastive word sense disambiguation test sets for machine translation . In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pag...
-
[51]
Annette Rios, Mathias M \"u ller, and Rico Sennrich. 2018. https://doi.org/10.18653/v1/W18-6437 The word sense disambiguation test suite at WMT 18 . In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 588--596, Belgium, Brussels. Association for Computational Linguistics
-
[52]
Annette Rios Gonzales, Laura Mascarell, and Rico Sennrich. 2017. https://doi.org/10.18653/v1/W17-4702 Improving word sense disambiguation in neural machine translation with sense embeddings . In Proceedings of the Second Conference on Machine Translation, pages 11--19, Copenhagen, Denmark. Association for Computational Linguistics
-
[53]
Sebastian Ruder. 2021. Challenges and opportunities in nlp benchmarking
2021
-
[54]
Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzemiński, Hakimeh Fadaei, Irem Ergün, Ifeoma Okoh, and 14 others. 2024. https://arxiv.org/abs/2402....
Pith/arXiv arXiv 2024
-
[55]
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adri Garriga-Alonso, and 1 others. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on machine learning research
2023
-
[56]
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, and 179 others. 2024. https://arxiv.org/abs/2408.00118 Gemma 2: ...
Pith/arXiv arXiv 2024
-
[57]
Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West. 2024. https://arxiv.org/abs/2402.10588 Do llamas work in english? on the latent language of multilingual transformers . Preprint, arXiv:2402.10588
Pith/arXiv arXiv 2024
-
[58]
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019. Paws-x: A cross-lingual adversarial dataset for paraphrase identification. arXiv preprint arXiv:1908.11828
Pith/arXiv arXiv 2019
-
[59]
Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. 2023. https://arxiv.org/abs/2306.05179 M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models . Preprint, arXiv:2306.05179
Pith/arXiv arXiv 2023
-
[60]
Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Khai Hao, Xu Han, Zhen Leng Thai, Shuo Wang, Zhiyuan Liu, and Maosong Sun. 2024. https://arxiv.org/abs/2402.13718 bench: Extending long context evaluation beyond 100k tokens . Preprint, arXiv:2402.13718
Pith/arXiv arXiv 2024
-
[61]
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D'souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. 2024. https://arxiv.org/abs/2402.07827 Aya model: An instruction finetuned open-access multil...
Pith/arXiv arXiv 2024
-
[62]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[63]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.