Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs

T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read MultiNRC: native multilingual reasoning questions in French, Spanish, and Chinese show current LLMs all scoring below 50%.

desk verdict A genuinely new native-authored multilingual reasoning benchmark, but the headline 'no model above 50%' is uncalibrated because of the adversarial filter and missing human baseline. read the letter →

arxiv 2507.17476 v1 pith:2TYANLBB submitted 2025-07-23 cs.CL cs.AI

classification cs.CLcs.AI
keywords multilingualreasoningbenchmarknativespeakersculturallinguisticwordplayLLMevaluationEnglishtranslation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces MultiNRC, a benchmark of more than 1,000 reasoning questions written by native speakers in French, Spanish, and Chinese, spanning linguistic reasoning, wordplay and riddles, cultural and tradition reasoning, and math with cultural relevance. It claims this benchmark measures something existing multilingual benchmarks miss: reasoning that depends on language-specific and culturally grounded knowledge, not just translated English content. Across 14 leading LLMs, none scores above 50% on the original-language questions, with the best model at 49.00%. When the cultural and math questions are translated into English, models improve by about 10 points on average for math but show no meaningful gain on cultural reasoning, suggesting that language is not the only barrier.

What carries the argument

The load-bearing object is MultiNRC's construction protocol: native speakers author short-answer reasoning questions, and an item is admitted only if at least three of five frontier models fail it, with two separate native-speaker review layers checking ground-truth answers and difficulty. The English-equivalent subset is built by having annotators translate the cultural and math prompts into English while preserving structure and solvability. Evaluation uses an LLM-as-a-judge that compares model responses with the ground-truth short answers, which the paper reports aligns with human judgments in over 95% of cases.

What would settle it

Rescore the questions that were rejected by the 3-of-5 failure filter and compare model accuracy on them with the kept set; if top models score substantially above 50% on the rejected items, the benchmark's headline difficulty is an artifact of selection rather than evidence about multilingual reasoning in general.

Watch

Extended reading notes

Core claim

The paper's central claim is that MultiNRC is a hard, valid test of native multilingual reasoning, and that current LLMs still fail it at scale. The benchmark keeps only questions that at least three of five strong models answered incorrectly at creation time, and the resulting set—1,055 questions—yields a top score of 49.00% (o3-pro) and an average below one-third. On the question types that can be translated faithfully (cultural and math), models score roughly ten points higher when the same math problems are posed in English, but cultural reasoning does not improve with English wording. The paper interprets this as evidence that models can retrieve relevant cultural knowledge more reliably in English for math, while the cultural-tradition questions require knowledge that is missing in both languages.

Load-bearing premise

The benchmark's central claim depends on the assumption that keeping only questions that stump at least three of five strong models produces a representative sample of native multilingual reasoning, rather than a curated set of model-specific failures.

Editorial extensions

If this is right

  • If MultiNRC measures what it claims, current LLMs have a large, measurable gap in culturally grounded non-English reasoning that English benchmarks cannot reveal.
  • The roughly 10-point math improvement on English translations means part of the multilingual deficit is a knowledge-retrieval problem tied to language, not purely a reasoning problem.
  • The flat cultural-reasoning scores across languages imply that many cultural/tradition questions require specific knowledge absent from model training data regardless of language.
  • The taxonomy of four reasoning categories provides a way to track model weaknesses independently of overall accuracy, since models rank differently across categories and languages.
  • Releasing the dataset gives other researchers a direct way to measure progress on native multilingual reasoning as models improve.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication not pursued in the paper is that the 3-of-5 failure filter may create a selection effect: if the rejected questions are far easier, then the benchmark's top score below 50% could overstate the general difficulty of native multilingual reasoning.
  • Because the English-equivalent set excludes linguistic and wordplay categories by design, the paper cannot measure whether translation would help those categories; a version that adapts puns or grammatical puzzles into English analogues could test this separately.
  • The judge's high agreement with humans is reported for short ground-truth answers; a natural extension is to audit the judge on longer free-form answers or on multi-part questions such as the Spanish example with a list of dates, where human judgment may diverge more.
  • The finding that models show different rankings by language suggests that a single aggregate multilingual score can hide large within-family disparities, and benchmarks like this one may be useful for detecting language-specific training data gaps.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces MultiNRC, a benchmark of 1,055 native French, Spanish, and Chinese reasoning questions across four categories (linguistic reasoning, wordplay and riddles, cultural/tradition reasoning, and math reasoning with cultural relevance). The items were authored by native speakers and filtered by a rule that keeps only questions that at least 3 of 5 state-of-the-art LLMs fail. The paper evaluates 14 LLMs in the original language and on English-equivalent translations of the cultural and math categories, reporting that no model exceeds 50% accuracy on the full benchmark (best: o3-pro at 49.00%) and that models improve on average by about 10% on math when the same questions are presented in English. The authors conclude that current LLMs remain poor at native multilingual reasoning and that English translation helps math reasoning but not cultural reasoning.

Significance. If the claims are properly calibrated, MultiNRC fills a real gap: it evaluates culturally and linguistically grounded reasoning in native languages rather than translated English benchmarks, and the English-equivalent design is a useful tool for separating language effects from knowledge-retrieval effects. The paper's concrete strengths are the two-layer native-speaker review, the reported LLM-judge agreement of over 95% with a Scott's Pi of 0.88, the systematic evaluation of 14 models, the language-by-category breakdown, and the public dataset release. The main limitation is calibration: because the benchmark is adversarially filtered against LLM performance and no human native-speaker baseline is reported, the headline deficiency claim is currently under-supported. The central descriptive findings are likely sound as measurements of performance on this particular filtered benchmark, but the interpretive step from those numbers to a general statement about multilingual reasoning ability needs additional evidence.

major comments (4)
  1. [Section 3.3 and Abstract] The inclusion rule in Section 3.3 ('we only keep questions that 3 or more of the 5 models fail to correctly answer') selects the benchmark adversarially against LLM performance, so the headline 'none above 50%' in Table 3 is partly inherited from the construction rule rather than being a measured property of native multilingual reasoning in general. The abstract and Section 5 infer a general deficiency ('current LLMs are still not good at native multilingual reasoning'), but without a native-speaker accuracy floor on the retained items, a 49% score could be either far below or close to human performance. The paper should report human native-speaker accuracy on the benchmark (ideally also on the English-equivalent subset) or explicitly restrict the claim to 'no model exceeds 50% on an adversarially filtered hard benchmark'.
  2. [Section 3.2 and Table 5] The English-equivalent set is constructed by asking annotators to translate 'the logic, not literal words,' but the paper provides no validation that the translated items are semantically equivalent or of comparable difficulty to the originals. The +10% math advantage in Table 5 is computed by subtracting Original accuracy from English accuracy on this pair set, so a systematic difference in explicitness, numeric precision, or cultural-cue salience introduced during translation could produce the observed delta without reflecting a general 'English advantage' in math reasoning. The paper should report a human equivalence and difficulty rating (e.g., bilingual native annotators judging whether the translation preserves the required reasoning steps) and show that the effect persists on a human-calibrated subset.
  3. [Section 4 (Automatic Evaluation)] The LLM-judge validation is reported only as a single summary: over 95% agreement with human judgment and a Scott's Pi of 0.88. The paper omits the sample size of the validation set, how the validation items were selected, the distribution of disagreements (false positives vs. false negatives), and whether judge errors correlate with language or category. Since every accuracy number in Section 5 rests on this metric, the paper should provide these details, ideally with a per-language breakdown.
  4. [Section 5 and Table 5] The claim that 'most models perform substantially better in math reasoning in English compared to in original languages (+10%)' is reported as an average over the filtered subset, but the paper does not analyze how the Section 3.3 filter interacts with language-dependent difficulty. If the filter disproportionately removes items that are easy in both languages, or items where the English translation is easy but the original is hard, the delta is not representative of math reasoning in general. A difficulty-matched analysis, or a human-calibrated subset, is needed before this finding can be generalized to native multilingual math reasoning as a whole.
minor comments (4)
  1. [Section 3.3, Table 2] The sentence comparing dataset size to MGSM says 'more than the 250 per-language count of MGSM'; stating the comparison explicitly as per-language (325–392 per language for MultiNRC) would remove ambiguity.
  2. [Section 6] There is a duplicated word in 'it will be be important'; also, the related-work sentence 'MultiLoko Hupkes & Bogoychev (2025)' is missing a comma between the dataset name and the citation.
  3. [Table 4 and Table 5] The deltas in Table 5 are presented without confidence intervals or significance tests; given cell sizes of roughly 70–120 items per category/language, some of the reported differences are within sampling noise. Reporting per-cell item counts or confidence intervals would make the cross-model comparisons more interpretable.
  4. [Appendix Table 8] In the Chinese math example, the English ground-truth answer gives 'approximately 2.08 meters' and '18.04 kilograms' without the source-dependent range ('217 cm (according to sources, one chi is generally between 23.1 and 23.3 cm)') present in the Chinese answer. This illustrates that English-equivalent answers may differ in precision, which is exactly the kind of issue that the equivalence validation in Major Comment 2 should address.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the difficulty filter is transparent and the headline scores are empirical, not forced by construction.

full rationale

MultiNRC is a benchmark-construction and evaluation paper rather than a derivation from first principles. There are no fitted parameters that are later relabeled as predictions, no load-bearing self-citations, and no imported uniqueness theorem. The one self-referential design point is Section 3.3, which retains only items that at least 3 of 5 SOTA models fail. This makes the benchmark adversarially hard by construction and means that low scores on the filtering models partly restate the inclusion rule; the average accuracy of those five models on the final set is capped at 40% simply by counting failures. However, the paper's headline numbers are not forced by that rule: o3-pro was not among the five filtering models and its 49% score, the category-by-category rankings, the language gaps, and the +10% English-math delta are empirical results obtained by running 14 models on the released items. The two native-speaker review layers check prompt/answer quality and verify the 3-of-5 failure condition, and the LLM judge is calibrated against human judgments with 95% agreement. The absence of a native-speaker accuracy floor is a real calibration limitation for the broader claim that LLMs are 'not good' at native multilingual reasoning, but it is a validity concern rather than a circular derivation. Overall, the central contribution has independent content and no significant circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

Central claims rest on dataset construction assumptions rather than derivation. The only hand-set numeric parameter is the 3-of-5 difficulty filter; no fitted constants or invented entities are used. The benchmark's external validity depends on the four domain assumptions listed above.

free parameters (1)
  • Difficulty threshold (3 of 5 SOTA model failures) = 3 out of 5
    Hand-selected admission criterion in Section 3.3 that defines which items enter MultiNRC; directly shapes the reported accuracy ceiling and is not derived from any theory.
assumptions (4)
  • domain assumption Native-speaking annotators produce correct and unambiguous ground-truth answers.
    Section 3.3 relies on annotator-written GTFAs plus two reviewer layers; no independent solve-rate or second-answer validation is reported.
  • ad hoc to paper The 3-of-5 model-failure filter selects genuinely hard native reasoning items rather than ambiguous or broken prompts.
    Section 3.3 defines dataset inclusion by model failures; this assumption is load-bearing for claims that low accuracy reflects model weakness.
  • domain assumption English translations preserve the logic, difficulty, and cultural grounding of original cultural and math prompts.
    Section 3.2 states translators focused on logic over literal words, but no equivalence validation is provided; required for the English-vs-Original comparison.
  • domain assumption GPT-4.1 judge labels are valid across all four languages and categories.
    Section 4 reports over 95% agreement with human judgments and Scott's Pi 0.88, but the judge is an LLM that may share systematic biases; agreement was measured on a subset.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs." pith.science (2026). https://pith.science/paper/2TYANLBB

@misc{pith2026250717476,
  author       = {Pith},
  title        = {Pith review of: MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2TYANLBB}},
  note         = {Machine review of arXiv:2507.17476}
}
read the original abstract

Although recent Large Language Models (LLMs) have shown rapid improvement on reasoning benchmarks in English, the evaluation of such LLMs' multilingual reasoning capability across diverse languages and cultural contexts remains limited. Existing multilingual reasoning benchmarks are typically constructed by translating existing English reasoning benchmarks, biasing these benchmarks towards reasoning problems with context in English language/cultures. In this work, we introduce the Multilingual Native Reasoning Challenge (MultiNRC), a benchmark designed to assess LLMs on more than 1,000 native, linguistic and culturally grounded reasoning questions written by native speakers in French, Spanish, and Chinese. MultiNRC covers four core reasoning categories: language-specific linguistic reasoning, wordplay & riddles, cultural/tradition reasoning, and math reasoning with cultural relevance. For cultural/tradition reasoning and math reasoning with cultural relevance, we also provide English equivalent translations of the multilingual questions by manual translation from native speakers fluent in English. This set of English equivalents can provide a direct comparison of LLM reasoning capacity in other languages vs. English on the same reasoning questions. We systematically evaluate current 14 leading LLMs covering most LLM families on MultiNRC and its English equivalent set. The results show that (1) current LLMs are still not good at native multilingual reasoning, with none scoring above 50% on MultiNRC; (2) LLMs exhibit distinct strengths and weaknesses in handling linguistic, cultural, and logical reasoning tasks; (3) Most models perform substantially better in math reasoning in English compared to in original languages (+10%), indicating persistent challenges with culturally grounded knowledge.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Combining multiple LLMs' reasoning traces into weighted DAGs gives an auditable consensus graph that matches self-consistency and modestly improves on majority voting.

Reference graph

Works this paper leans on

37 extracted references · 16 canonical work pages · cited by 1 Pith paper

  1. [1]

    Llama 4: Advancing multimodal intelligence

    Meta AI. Llama 4: Advancing multimodal intelligence. https://ai.meta.com/blog/llama-4-multimodal-intelligence/, 2024. Accessed: 2025-06-21

  2. [2]

    Claude 3.7 sonnet and claude code

    Anthropic . Claude 3.7 sonnet and claude code. Anthropic News, 2025. URL https://www.anthropic.com/news/claude-3-7-sonnet. 5 min read

  3. [3]

    Arc prize 2024: Technical report

    Francois Chollet, Mike Knoop, Gregory Kamradt, and Bryan Landers. Arc prize 2024: Technical report. ArXiv preprint, abs/2412.04604, 2024. URL https://arxiv.org/abs/2412.04604

  4. [4]

    Think you have solved question answering? try arc, the ai2 reasoning challenge

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. ArXiv preprint, abs/1803.05457, 2018. URL https://arxiv.org/abs/1803.05457

  5. [5]

    Training verifiers to solve math word problems

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. ArXiv preprint, abs/2110.14168, 2021. URL https://arxiv.org/abs/2110.14168

  6. [6]

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities

    Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025

  7. [7]

    thinking

    Google DeepMind. Gemini model and “thinking” updates: March 2025. https://blog.google/technology/google-deepmind/gemini-model-thinking-updates-march-2025/, 2025. Accessed: 2025-06-21

  8. [9]

    Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies

    Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9: 0 346--361, 2021. doi:10.1162/tacl_a_00370. URL https://aclanthology.org/2021.tacl-1.21

Show all 37 references
  1. [10]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. ArXiv preprint, abs/2501.12948, 2025. URL https://arxiv.org/abs/2501.12948

  2. [11]

    EXAMS : A multi-subject high school examinations dataset for cross-lingual and multilingual question answering

    Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. EXAMS : A multi-subject high school examinations dataset for cross-lingual and multilingual question answering. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.), Pro...

  3. [12]

    Measuring mathematical problem solving with the math dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track ...

  4. [13]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview...

  5. [14]

    Benchmax: A comprehensive multilingual evaluation suite for large language models

    Xu Huang, Wenhao Zhu, Hanxu Hu, Conghui He, Lei Li, Shujian Huang, and Fei Yuan. Benchmax: A comprehensive multilingual evaluation suite for large language models. ArXiv preprint, abs/2502.07346, 2025. URL https://arxiv.org/abs/2502.07346

  6. [15]

    Multiloko: a multilingual local knowledge benchmark for llms spanning 31 languages

    Dieuwke Hupkes and Nikolay Bogoychev. Multiloko: a multilingual local knowledge benchmark for llms spanning 31 languages. ArXiv preprint, abs/2504.10356, 2025. URL https://arxiv.org/abs/2504.10356

  7. [16]

    Eliciting better multilingual structured reasoning from llms through code

    Bryan Li, Tamer Alkhouli, Daniele Bonadiman, Nikolaos Pappas, and Saab Mansour. Eliciting better multilingual structured reasoning from llms through code. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 5...

  8. [17]

    Logiqa: A challenge dataset for machine reading comprehension with logical reasoning

    Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. Logiqa: A challenge dataset for machine reading comprehension with logical reasoning. In Christian Bessiere (ed.), Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligen...

  9. [18]

    Aime 2024

    Mathematical Association of America . Aime 2024. https://www.maa.org/math-competitions/aime, 2024. Problem set and results

  10. [19]

    New models and developer products announced at spring update

    OpenAI . New models and developer products announced at spring update. https://openai.com/blog/new-models-and-developer-products-announced-at-spring-update, 2024. Accessed: 2024-06-13

  11. [20]

    Introducing gpt-4.1 in the api

    OpenAI. Introducing gpt-4.1 in the api. https://openai.com/index/gpt-4-1/, 2024. Accessed: 2024-06-21

  12. [21]

    Introducing o3 and o4 mini

    OpenAI. Introducing o3 and o4 mini. https://openai.com/index/introducing-o3-and-o4-mini/, 2025. Accessed: 2025-06-21

  13. [22]

    Arkil Patel, Satwik Bhattamishra, and Navin Goyal. Are NLP models really able to solve simple math word problems? In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (eds.),...

  14. [23]

    SQ u AD : 100,000+ questions for machine comprehension of text

    Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. SQ u AD : 100,000+ questions for machine comprehension of text. In Jian Su, Kevin Duh, and Xavier Carreras (eds.), Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp.\ 23...

  15. [24]

    Gpqa: A graduate-level google-proof q&a benchmark

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling

  16. [25]

    Winogrande: An adversarial winograd schema challenge at scale

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligenc...

  17. [26]

    William A. Scott. Reliability of content analysis: The case of nominal scale coding. Public Opinion Quarterly, 19 0 (3): 0 321--325, 1955

  18. [27]

    Language models are multilingual chain-of-thought reasoners

    Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Le...

  19. [28]

    M4u: Evaluating multilingual understanding and reasoning for large multimodal models

    Hongyu Wang, Jiayu Xu, Senwei Xie, Ruiping Wang, Jialin Li, Zhaojie Xie, Bin Zhang, Chuyan Xiong, and Xilin Chen. M4u: Evaluating multilingual understanding and reasoning for large multimodal models. ArXiv preprint, abs/2405.15638, 2024 a . URL https://arxiv.org/abs/2405.15638

  20. [29]

    Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

    Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. Mmlu-pro: A more robust and challenging multi-task language unders...

  21. [30]

    Mmlu-prox: A multilingual benchmark for advanced large language model evaluation, 2025

    Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Yun Xing, Junjue Wang, Huitao Li, Xin Li, Kunyu Yu, Nan Liu, Qingyu Chen, Douglas Teodoro, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, and Irene Li. Mmlu-prox: A multilingual benchmark for advanc...

  22. [31]

    Reclor: A reading comprehension dataset requiring logical reasoning

    Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. Reclor: A reading comprehension dataset requiring logical reasoning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net, 2020. URL https://open...

  23. [32]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. H ella S wag: Can a machine really finish your sentence? In Anna Korhonen, David Traum, and Llu \' s M \`a rquez (eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguist...

  24. [33]

    M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models

    Wenxuan Zhang, Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.), Advances i...

  25. [34]

    AGIE val: A human-centric benchmark for evaluating foundation models

    Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. AGIE val: A human-centric benchmark for evaluating foundation models. In Kevin Duh, Helena Gomez, and Steven Bethard (eds.), Findings of the Association for Comput...

  26. [35]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  27. [36]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  28. [37]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  29. [38]

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.