A native-authored French, Spanish, and Chinese reasoning benchmark shows current LLMs score below 50% and improve by about 10% on math when questions are in English.
Claude 3.7 sonnet and claude code
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
A native-authored French, Spanish, and Chinese reasoning benchmark shows current LLMs score below 50% and improve by about 10% on math when questions are in English.