{"id":"dd9b79ee-1225-421c-a088-9c57880f60aa","arxiv_id":"2506.07418","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new multilingual, image-based Kangaroo math benchmark shows Gemini 2.0 Flash, Qwen-VL 2.5 72B, and GPT-4o lead current multimodal LLMs, but all remain far below human accuracy on visual math reasoning.","lead":"Researchers built a multilingual benchmark from Kangaroo math contest questions that include diagrams and tested nine multimodal AI models in English, French, Spanish, and Catalan. They found that even the best models, Gemini 2.0 Flash, Qwen-VL 2.5 72B, and GPT-4o, stay far below human-level accuracy and often fail to use diagram information.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Image-vs-text gap is confounded by question difficulty: non-image questions come from a different and apparently easier set, so the 'underutilization of diagrams' conclusion is not established by the reported comparison.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the image-versus-no-image comparison lacks a control for question difficulty and content. I agree that this is the central threat to the paper's interpretive claim rather than to the benchmark itself. The model ranking on image-based tasks is computed on a fixed set of questions and therefore survives the confound, but the abstract's second finding and the discussion's underutilization conclusion depend on the uncontrolled comparison. The paper's own observation that higher-level questions contain fewer images makes the confound concrete: if text-only questions are drawn from easier levels, the accuracy advantage on those questions is expected regardless of whether models use diagrams. The suggested paired ablation or difficulty/topic-matched stratification would settle the question. Since the reader already assigned CONDITIONAL and the required fix is feasible, I recommend no change to the verdict.","tokens_in":11280,"tokens_out":5388,"duration_ms":70343,"concrete_test":"From the released GitHub dataset, build paired versions of every image-based KMC question: original figure plus text, and the same text stem/options with the figure removed (or replaced by an equivalent precise textual description). Evaluate all models on both versions. If accuracy does not drop substantially when the image is removed, or if the original image versus no-image gap reverses on matched pairs, the 'underutilization' conclusion is not supported; if it drops sharply, the conclusion holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the paper's second finding: the accuracy gap between image-based and non-image questions is interpreted as evidence that models 'underutilize diagrammatic information.' For that inference to hold, the two question sets must be comparable in difficulty and content. The paper does not provide this control: Table 1 and Table 4 compare different KMC questions with images against different KMC questions without images, and the Figure 2 caption explicitly notes that models are not evaluated on the same questions in each language. The confound is admitted indirectly in Section 4, where the authors explain the upward accuracy trend with difficulty as 'linked to the reduced presence of images in higher-level questions,' meaning image presence is correlated with difficulty. If text-only questions are easier, both the observed text-over-image gain and the smaller gain for weak models are explained without assuming underutilization. Topic mix is also unmatched: Figure 4 defines topic categories only for image-based questions, so the non-image set has no topic control either. The headline ranking of Gemini on image tasks is less affected by this confound, but the abstract's second finding and the discussion's 'underutilized diagrammatic and visual information' conclusion rest on the uncontrolled comparison. Because the benchmark itself is the main contribution, this is an addressable methodological gap rather than a fatal flaw.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a multilingual visual-mathematics benchmark derived from Kangaroo Mathematics Competition (KMC) tests in English, French, Spanish, and Catalan, and evaluates nine multimodal LLMs on accuracy for image-based and text-only questions. The authors report four findings: (1) no model excels across all mathematical topics; (2) most models perform better on questions without images, which is interpreted as underutilization of diagrammatic information; (3) substantial variation across languages and difficulty levels, with Gemini 2.0 Flash achieving the highest precision on image-based tasks; and (4) a qualitative analysis suggesting Gemini and GPT-4o engage in structured reasoning while Pixtral and Llama often fall back on heuristics or random choice. The dataset and code are released on GitHub.","tokens_in":11525,"tokens_out":2374,"duration_ms":32098,"significance":"If the comparisons were properly controlled, the paper would provide a useful multilingual, diagram-focused benchmark for MLLM mathematical reasoning, with a current capability ranking across nine models and a released dataset that can support further research. The authors' choice of KMC tests, the four-language setup, and the topic categorization for image-based questions are strengths, as is the availability of the data and code. However, the headline inference about underutilization of diagrams rests on an uncontrolled comparison between different question sets, so the significance of the main interpretative claim is currently limited.","major_comments":[{"comment":"The comparison between image-based and non-image accuracy is computed on different KMC question sets, not on matched items or on the same questions rendered with and without images. The abstract and Section 5 interpret the resulting accuracy gap as evidence that models 'underutilize diagrammatic and visual information,' but this conclusion does not follow from the reported design. The authors themselves note in Section 4 that the upward accuracy trend with difficulty level 'may be linked to the reduced presence of images in higher-level questions,' which means image presence is correlated with difficulty. Without a matched control for difficulty and topic mix, the observed text-over-image gain could simply reflect that the text-only questions are easier or have a different topic composition. I recommend either restricting the claim to a descriptive observation about accuracy on two different subsets, or adding a controlled comparison (for example, text-only renderings of the same image-based questions, or per-difficulty/per-topic stratification).","section":"Section 4, Table 4, Figure 1"},{"comment":"The Figure 2 caption states that models are 'not evaluated on the same questions in each language.' Consequently, cross-language accuracy differences and difficulty-level comparisons are confounded by the fact that each language version of the KMC contains different items. The third finding in the abstract, 'substantial variation exists across languages and difficulty levels,' is therefore only a descriptive statement about these particular test forms; it does not establish that the models' multilingual ability differs. To support the language-variation claim, the authors would need per-item translated equivalents or an explicit analysis of item difficulty across languages.","section":"Figure 2 and language-level comparisons"},{"comment":"The analysis aimed at distinguishing reasoning from recitation is reported qualitatively and without quantitative coding criteria or inter-rater reliability. The text states that Pixtral and Llama 'frequently returned No answer responses' and that Gemini and GPT-4o demonstrated 'more coherent and structured reasoning,' but no counts, examples, or error taxonomy are provided. As written, the fourth abstract finding is not supported by the reported evidence. I suggest adding a small quantitative coding scheme for reasoning quality, or explicitly labeling this part as anecdotal and removing it from the abstract's list of findings.","section":"Section 5, recitation analysis"}],"minor_comments":[{"comment":"The word 'Kangarooo' appears to be a typo and should read 'Kangaroo.'","section":"Section 4, first paragraph"},{"comment":"The phrase 'an escellent text to check for recitation' should be corrected to 'an excellent test to check for recitation.'","section":"Introduction, last paragraph before Section 2"},{"comment":"The 'Results (%)' column is described as showing image-based and non-image results on the left and right sides, but the caption could state this explicitly in the table header rather than only in the caption text.","section":"Table 1, caption"},{"comment":"The sentence 'The figure should be interpreted separately with non-normalized values' is unclear; please specify which values are non-normalized and why the heatmaps must be read separately.","section":"Figure 3, caption"}],"recommendation":"major_revision","confidential_remarks":"The paper's benchmark and dataset are potentially useful contributions, and the ranking of models on image-based questions is likely robust as a descriptive result. The main reasons for the major-revision recommendation are the uncontrolled image-versus-text comparison and the language-level confound, both of which affect headline abstract claims. These are addressable within the manuscript's scope by either adding controlled analyses or softening the causal language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper delivers a genuinely new resource — a multilingual Kangaroo benchmark with images, a topic taxonomy, and nine MLLMs evaluated across four languages. The accuracy tables and language/difficulty breakdowns are plausible and useful. But the second headline finding — that models underutilize diagrams because they do better on text-only questions — is not supported by the design, since image and text question sets are different and uncontrolled for difficulty or topic.\n\nWhat's new: the dataset itself. Prior Kangaroo work is text-only; MathVista and similar benchmarks don't cover these languages at this age range. The topic heatmap for image questions and the co-occurrence analysis are nice additions.\n\nSoft spots: the image-vs-text gap is confounded. Section 4 notes accuracy rises with difficulty while image presence falls, so the text-only advantage could simply reflect easier questions. The Figure 2 language comparison also uses different questions per language. The reasoning-vs-recitation analysis is qualitative; 'no answer' and random-guessing behaviors are described but not quantified. There are no error bars or significance tests, and the dataset is cited only as 'Kangaroo Language Test on GitHub' without a stable identifier.\n\nNone of this is fatal. The model ranking on image tasks is less affected, and the benchmark is a standalone contribution. But the abstract and discussion should be revised to present the underutilization claim as a hypothesis rather than an established result.\n\nAnyone working on multimodal math reasoning will get value from the benchmark itself. This paper belongs in a venue for evaluation resources. Send it to peer review; the authors should be asked to control the image/text comparison, add error bars or per-question matching, and release the dataset under a proper identifier.","headline":"Useful new multilingual visual math benchmark and model snapshot, but the underutilization-of-diagrams conclusion is confounded by uncontrolled question sets and should be reframed as a hypothesis.","tokens_in":12058,"tokens_out":2291,"would_cite":true,"duration_ms":26334,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Even the strongest multimodal model fails more than half of image-based Kangaroo math questions, while doing much better on text-only ones, evidence that diagrams are underused.","keywords":["multimodal large language models","visual mathematics","Kangaroo contest","diagram understanding","multilingual evaluation","mathematical reasoning","benchmark"],"falsifier":"Present the same Kangaroo problem in two forms, one with the original diagram and one with the diagram's essential information restated in text, and check whether models' accuracy changes; if models do not lose accuracy when the diagram is removed or made decorative, the claim that they underutilize diagrams is falsified, and if the text-only versions of the same problems are easier, then the benchmark's image/text gap reflects question difficulty rather than diagram use.","tokens_in":11127,"feed_emoji":"📐","tokens_out":7043,"duration_ms":73878,"temperature":0.7,"pith_summary":"The paper builds a multilingual benchmark from Kangaroo mathematics contest problems in English, French, Spanish, and Catalan, with and without accompanying images, and runs nine multimodal large language models on them. It argues that these models perform only moderately on visual mathematics and that most improve when images are removed, which the authors read as evidence that models underutilize diagrammatic information. The best model on image-based questions, Gemini 2.0 Flash, reaches 45.4% accuracy, with Qwen-VL 2.5 72B and GPT-4o close behind, far below human-level performance on the same contest questions. A secondary analysis attempts to separate genuine reasoning from memorization, finding Gemini and GPT-4o produce more structured reasoning while smaller models often fall back on guessing or refuse to answer.","feed_headline":"Best multimodal model misses over half of visual math problems","feed_subtitle":"Kangaroo math benchmark in four languages: models answer text-only questions far better than diagram questions","key_machinery":"The load-bearing object is the multilingual Kangaroo dataset: contest questions from 2014 to 2024 in four languages, each annotated with whether it contains an image, categorized by mathematical topic, and presented to models in a structured prompt that asks for reasoning before the final answer. The key comparison is accuracy on image-based versus non-image questions, which the authors treat as a probe for whether models use diagrammatic information. A co-occurrence heatmap analysis and an analysis of answers on fresh contest questions from the Comunitat Valenciana are used to test whether models reason rather than recite.","core_discovery":"The core discovery is a capability ranking and a behavioral pattern: on image-based Kangaroo questions, Gemini 2.0 Flash achieves the highest precision (45.4%), followed by Qwen-VL 2.5 72B (43.5%) and GPT-4o (40.2%), while most models shift to higher accuracy on text-only questions (Gemini 2.0 Flash 75.9%, Qwen-VL 2.5 72B 70.6%, GPT-4o 65.3%). The authors interpret the accuracy gap as indicating that models underutilize diagrammatic information rather than being unable to reason about the underlying mathematics. No model excels across all mathematical topics (geometry, visual algebra, logic, patterns, combinatorics), and precision varies by language and difficulty, with visual logic and reasoning being the hardest category for all models.","pith_inferences":["Editorial inference: The clean way to test the underutilization claim is to compare matched versions of the same problem with and without the diagram; the paper's current design compares different questions, so the language and difficulty mix can explain part of the gap.","Editorial inference: The benchmark could be extended to measure diagram comprehension directly, for instance by asking models to read off marked angles, lengths, or counts from figures before solving, separating perception errors from reasoning errors.","Editorial inference: The near-human performance gap and the topic-specific leaders suggest that progress may come less from scaling and more from training that forces models to ground each reasoning step in the image, possibly using intermediate structured representations of the diagram."],"forward_implications":["If the accuracy gap between image and text questions does reflect underutilization of diagrams, then current MLLMs are leaving a large fraction of visual math performance on the table, and training that teaches models to use diagrams should be prioritized over pure parameter scaling.","The benchmark gives future models a concrete target: reaching human-level performance on Kangaroo image questions requires improving on the current best image-based precision of 45.4% by a wide margin, with the hardest categories being visual logic and reasoning.","Because no model leads on every mathematical topic, in-domain specialization is not yet possible; the topic-level heatmap suggests each family has specific strengths, but the co-occurrence analysis shows that correlated failures will limit simple ensemble strategies.","The strong performance of Gemini 2.0 Flash and Qwen-VL 2.5 72B on text-only questions, combined with their bigger drops on image questions, indicates that multilingual textual math reasoning is already strong while visual grounding is the bottleneck."],"supporting_citations":[{"why":"Provides the English-language Kangaroo contest questions used to build the text-only and image-based evaluation set.","marker":"[41]"},{"why":"Supplies the Spanish-version contest questions in the multilingual benchmark.","marker":"[42]"},{"why":"Supplies the French-version contest questions in the multilingual benchmark.","marker":"[43]"},{"why":"Supplies the Catalan-version contest questions, including the fresh test used for the reasoning-versus-recitation analysis.","marker":"[44]"},{"why":"Prior evaluation of text-only Kangaroo problems with dedicated reasoners, which this study extends to multimodal models.","marker":"[14]"},{"why":"MathVista, the visual-math benchmark that frames the comparison for measuring visual reasoning in multimodal models.","marker":"[20]"},{"why":"Motivates the analysis distinguishing recitation from reasoning, which the paper applies to newly released contest questions.","marker":"[13]"}],"fun_headline_variants":["Gemini tops visual math, but still fails half","Multilingual math benchmark: diagrams stump AI","No AI model excels at visual math across languages","Text-only questions boost AI math scores","Visual math: AI underuses diagrams, study finds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central comparison assumes that questions with images and questions without images are equally hard, so that any accuracy gap must come from how models handle diagrams rather than from differences in the questions themselves.","fun_headline_variants_meta":{"raw":{"variants":["Gemini tops visual math, but still fails half","Multilingual math benchmark: diagrams stump AI","No AI model excels at visual math across languages","Text-only questions boost AI math scores","Visual math: AI underuses diagrams, study finds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000221,"raw_usage":{"total_tokens":1484,"prompt_tokens":1014,"completion_tokens":470,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":400}},"tokens_in":630,"tokens_out":470,"duration_ms":5768,"temperature":1.0,"reasoning_tokens":400,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:34:15.164868+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Present the same Kangaroo problem in two forms, one with the original diagram and one with the diagram's essential information restated in text, and check whether models' accuracy changes; if models do not lose accuracy when the diagram is removed or made decorative, the claim that they underutilize diagrams is falsified, and if the text-only versions of the same problems are easier, then the benchmark's image/text gap reflects question difficulty rather than diagram use.","supporting_citations":[{"cited_title":"Kangaroo mathematics competition, 2025","cited_arxiv_id":null,"evidence_quote":"Provides the English-language Kangaroo contest questions used to build the text-only and image-based evaluation set."},{"cited_title":"Concurso Canguro de Matem´ aticas, 2025.https://canguromat.es/Last retrieved, April 25th, 2025","cited_arxiv_id":null,"evidence_quote":"Supplies the Spanish-version contest questions in the multilingual benchmark."},{"cited_title":"Concours Kangourou de Math´ ematiques, 2025.https://www.aksf","cited_arxiv_id":null,"evidence_quote":"Supplies the French-version contest questions in the multilingual benchmark."},{"cited_title":"Concurs Cangur de Matem` atiques, 2025.https://scm.iec","cited_arxiv_id":null,"evidence_quote":"Supplies the Catalan-version contest questions, including the fresh test used for the reasoning-versus-recitation analysis."},{"cited_title":"Rhomrasi, Y","cited_arxiv_id":null,"evidence_quote":"Prior evaluation of text-only Kangaroo problems with dedicated reasoners, which this study extends to multimodal models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Motivates the analysis distinguishing recitation from reasoning, which the paper applies to newly released contest questions."}],"review_version":1}