{"id":"ecede837-cbb2-4559-ae02-2f7a8ccef587","arxiv_id":"2607.05992","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"PluraMath extends PolyMath with human-validated math problems in 18 mid-to-extreme low-resource languages and benchmarks 27 reasoning LLMs, finding a persistent high- vs low-resource performance gap.","lead":"PluraMath adds 18 underrepresented languages to math-reasoning LLM benchmarks via native-speaker-validated translations. It lets researchers measure whether strong reasoning models still fail when the problem is not in English or Chinese.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Measured high- vs underrepresented-language math gaps on PluraMath may partly reflect residual translation or instruction-template artifacts if native-speaker validation did not enforce solution-preserving equivalence.","rationale":"The reader’s weakest_assumption correctly isolates the load-bearing premise for the strongest claim. Open release of dataset, pipeline, and evaluation framework would make this a useful extension of PolyMath; the scientific interpretation of the multi-scale gap as a multilingual reasoning gap (rather than a translation/instruction confound) remains conditional on demonstrated solution-preserving equivalence. No internal inconsistency or circularity is visible from the available text; disagreement with consensus is not at issue. The concern is empirical and settleable by a re-solve audit plus inspection of IAA, QA metrics, language-selection criteria, and public artifacts. Dataset papers of this type are accept-shaped when those hold, so the CONDITIONAL verdict and LOW confidence are appropriate; no adjustment is warranted. The concrete test above is the single check that would most directly land or clear the concern.","tokens_in":2062,"tokens_out":605,"duration_ms":33952,"concrete_test":"Independently sample ≥50 PluraMath items per language (or the full set if smaller). Have a second native-speaker mathematician re-solve each item from the released text only (no original), recording (a) whether the intended answer is uniquely recoverable under the paper’s scoring rubric, (b) perceived difficulty vs. a matched PolyMath control, and (c) any instruction/format failures. If >10% of items fail unique-answer recovery or show resource-level-correlated difficulty/format shift, recompute the headline gap after filtering those items; a material shrink weakens the pure-reasoning attribution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that a persistent high- vs underrepresented-language gap in mathematical reasoning holds across 27 models at four scales, with stronger results largely tracking instruction-following—requires that PluraMath items preserve mathematical content, difficulty, and answer format relative to PolyMath. The pipeline uses pre-computed translations that native speakers “thoroughly validated,” but the abstract does not secure what validators checked (surface fluency vs. re-solving for unique correct answers under the same rubric), residual error rates, IAA, or whether instruction templates were language-adapted without format drift. If validation was primarily linguistic rather than solution-preserving, residual artifacts (ambiguous wording, difficulty shift, script/format issues, template mismatch) would systematically depress underrepresented-language scores and inflate the apparent reasoning gap. The secondary association with instruction-following is vulnerable to the same confound if that ability is measured on the same translated prompts. This does not assert the gap is false; it asserts that attribution to multilingual mathematical reasoning (vs. evaluation pipeline) is the least secure link in the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript introduces PluraMath, an extension of the PolyMath mathematical-reasoning benchmark from 18 high-resource languages to 18 additional underrepresented languages spanning six language families (mid- to extreme low-resource). Items are produced via a human-curated pipeline in which native speakers validate pre-computed translations. The authors then evaluate 27 reasoning LLMs across four scale bands (small, mid-size, large, and closed-source ensembles) and report a persistent performance gap between high-resource and underrepresented languages, with stronger scores largely associated with better instruction-following. The dataset, acquisition pipeline, and evaluation framework are fully open-sourced.","tokens_in":2234,"tokens_out":1188,"duration_ms":35127,"significance":"If the items preserve mathematical content, difficulty, and answer format across languages, PluraMath would be a concrete and timely contribution: it expands multilingual math evaluation into languages that existing suites systematically omit, and the multi-scale evaluation of 27 models supplies a useful snapshot of current capabilities. Full open-sourcing of the dataset, pipeline, and evaluation framework is a clear strength that can lower the barrier for community-driven extension to further underrepresented languages. The secondary link between performance and instruction-following is a potentially actionable finding for model development, provided it is not confounded by evaluation artifacts.","major_comments":[{"comment":"The central claim—that measured high- vs underrepresented-language gaps reflect multilingual mathematical reasoning rather than evaluation artifacts—depends on solution-preserving equivalence of PluraMath items to their PolyMath sources. The abstract and construction description state that native speakers “thoroughly validated” pre-computed translations, but do not report residual error rates, inter-annotator agreement, or an explicit validation protocol that requires re-solving for a unique correct answer under a fixed rubric (as opposed to surface fluency or grammaticality). Without those controls, residual wording ambiguity, difficulty shift, script/format drift, or answer-format mismatch can systematically depress underrepresented-language scores and inflate the apparent reasoning gap. This is load-bearing for attribution; the manuscript should quantify validation quality and, where ","section":null},{"comment":"The secondary finding that stronger results are “largely associated with better instruction-following ability” is vulnerable to the same confound if instruction-following is scored on the same translated prompts or templates. The manuscript should (i) define how instruction-following is measured (separate probe vs. same PluraMath items), (ii) state whether instruction templates were language-adapted and how format consistency was enforced, and (iii) show that the association survives controls for residual translation quality or template mismatch. Otherwise the association cannot be cleanly attributed to model capability rather than prompt surface form.","section":null},{"comment":"Cross-language and cross-scale gap claims need explicit statistical support and matched evaluation conditions. The fine-grained analysis should report per-language and per-family scores with uncertainty (e.g., bootstrap CIs or paired tests against the PolyMath high-resource baseline under identical decoding and answer-extraction settings), and should document any language-specific answer-normalization rules. Without this, “persistent gap” remains a qualitative summary rather than a secured empirical result, especially for extreme low-resource languages where small absolute score differences can be dominated by extraction failures.","section":null}],"minor_comments":[{"comment":"Clarify the exact relationship to PolyMath (Wang et al., 2025): which problem subsets were reused, whether difficulty tiers were preserved, and whether any items were rewritten rather than translated.","section":null},{"comment":"List the 18 underrepresented languages and their language-family groupings early (table or figure) so readers can map resource level to results without hunting the appendix.","section":null},{"comment":"Define the four model-scale bands (parameter ranges or explicit model lists) in the main text so the “27 models across four scales” claim is auditable without supplementary material.","section":null},{"comment":"Report decoding hyperparameters, answer-extraction method, and any language-specific post-processing in a single reproducible evaluation subsection; open-sourcing the framework is valuable but the paper should still document the protocol.","section":null},{"comment":"In the abstract and introduction, avoid implying that PolyMath covers “only high-resource languages” without a brief resource-level criterion (e.g., Common Crawl share or labeled-data availability) so the high- vs underrepresented contrast is operationally clear.","section":null},{"comment":"If figures plot high-resource vs underrepresented aggregates, add per-language strip or family-level breakdowns so outliers in extreme low-resource settings are visible.","section":null}],"recommendation":"major_revision","confidential_remarks":"The contribution is timely and the open-sourcing commitment is a genuine plus for cs.CL. My major_revision recommendation is driven almost entirely by the missing quantitative validation of solution-preserving translation equivalence and the potential confound with instruction-following measurement—both fixable within the manuscript’s scope via audits, IAA/error rates, and clearer metric definitions. I do not see an internal inconsistency or a reason to reject on novelty grounds; the stress-test concern about residual translation artifacts is the main correctness-risk for the central gap claim. If the full paper already contains residual-error rates, re-solve audits, and a separated instruction-following probe that I could not verify from the materials available here, the recommendation could move to minor_revision after a quick check of those sections."},"author_rebuttal":{"model":"grok-4.5","summary":"We thank the referee for a careful and constructive review. The three major comments correctly identify load-bearing reporting gaps around (i) translation/validation quality controls, (ii) the operational definition of instruction-following and possible confounds with prompt surface form, and (iii) statistical support for cross-language and cross-scale gap claims. We agree that these points must be addressed before the central attribution—that measured gaps reflect multilingual mathematical reasoning rather than evaluation artifacts—can be secured. We will revise the manuscript accordingly: expand the validation protocol with residual-error and agreement metrics, clarify how instruction-following is measured and controlled, and add per-language/family scores with uncertainty under matched evaluation conditions. Full open-sourcing of dataset, pipeline, and evaluation code remains unchanged and will incorporate the additional diagnostics. We believe these revisions directly answer the referee’s concerns and strengthen the contribution.","responses":[{"response":"We agree that solution-preserving equivalence is load-bearing for attribution, and that the current manuscript under-specifies validation quality. Our pipeline already required native-speaker validators to check that each item retained the same mathematical content, difficulty intent, and unique gold answer as the PolyMath source (not merely fluency), with rejection and re-translation on failure. However, we did not report residual error rates, inter-annotator agreement, or the full rubric in the text. In revision we will: (1) spell out the explicit validation protocol and acceptance criteria (including re-solving / answer uniqueness under a fixed rubric); (2) report residual error rates from a held-out double-annotation sample, stratified by language and resource band; (3) report inter-annotator agreement on accept/reject and on answer equivalence; and (4) document any residual issues (script/format drift, answer-format mismatch) and how they were handled. Where residual risk remains for extreme low-resource languages, we will state it explicitly rather than over-claim equivalence. These additions will appear in the dataset-construction section and appendix, with the open-sourced pipeline updated to match.","revision_made":"yes","referee_comment":"The central claim—that high- vs underrepresented-language gaps reflect multilingual mathematical reasoning rather than evaluation artifacts—depends on solution-preserving equivalence of PluraMath items to PolyMath sources. The manuscript states that native speakers “thoroughly validated” pre-computed translations, but does not report residual error rates, inter-annotator agreement, or an explicit validation protocol that requires re-solving for a unique correct answer under a fixed rubric (vs. surface fluency). Without those controls, residual ambiguity, difficulty shift, script/format drift, or answer-format mismatch can systematically depress underrepresented-language scores and inflate the apparent reasoning gap. The manuscript should quantify validation quality."},{"response":"We agree the association is currently under-specified and could be confounded by prompt surface form. In the present draft, “instruction-following” was operationalized primarily as format compliance and successful answer extraction on the same PluraMath items (e.g., producing a parseable final answer in the required form), not via an independent probe. Instruction templates were language-adapted by native speakers as part of the same validation pipeline, with a shared answer-extraction schema, but we did not fully document adaptation rules or enforce/report format consistency metrics. In revision we will: (i) define instruction-following explicitly (format compliance / extractability rates, and any separate diagnostic if retained); (ii) document language adaptation of templates and the shared extraction/normalization rules; (iii) report the association after controlling for residual translation-quality indicators and template-mismatch flags (e.g., partial correlations or stratified analyses by validation residual band); and (iv) qualify the claim where it does not survive those controls. We will not overstate a causal capability story if the association is partly surface-form driven. The evaluation framework release will expose the exact templates and extraction code used.","revision_made":"yes","referee_comment":"The secondary finding that stronger results are “largely associated with better instruction-following ability” is vulnerable to the same confound if instruction-following is scored on the same translated prompts or templates. The manuscript should (i) define how instruction-following is measured (separate probe vs. same PluraMath items), (ii) state whether instruction templates were language-adapted and how format consistency was enforced, and (iii) show that the association survives controls for residual translation quality or template mismatch. Otherwise the association cannot be cleanly attributed to model capability rather than prompt surface form."},{"response":"We agree that “persistent gap” must be backed by uncertainty estimates and matched conditions, not only qualitative summary. All models were already run under a shared decoding and answer-extraction setup, but the manuscript did not report per-language/family uncertainty or formal comparisons to the PolyMath high-resource baseline, and language-specific normalization rules were under-documented. In revision we will: (1) report per-language and per-family scores with bootstrap confidence intervals; (2) add paired or stratified tests against the high-resource PolyMath baseline under identical decoding and extraction settings; (3) document all answer-normalization and extraction rules, including any language-specific exceptions (numeral systems, script variants, answer delimiters); and (4) separate extraction-failure rates from content-correct rates so that extreme low-resource gaps are not driven solely by unparseable outputs. We will revise figures/tables and the analysis section accordingly, and release the evaluation scripts that reproduce the CIs and tests. This converts the gap claim into a secured empirical result with explicit caveats where sample size or extraction noise dominates.","revision_made":"yes","referee_comment":"Cross-language and cross-scale gap claims need explicit statistical support and matched evaluation conditions. The fine-grained analysis should report per-language and per-family scores with uncertainty (e.g., bootstrap CIs or paired tests against the PolyMath high-resource baseline under identical decoding and answer-extraction settings), and should document any language-specific answer-normalization rules. Without this, “persistent gap” remains a qualitative summary rather than a secured empirical result, especially for extreme low-resource languages where small absolute score differences can be dominated by extraction failures."}],"tokens_in":1891,"tokens_out":1346,"duration_ms":28083,"standing_objections":[]},"desk_editor":{"model":"grok-4.5","letter":"Colleague —\n\nPunchline: PluraMath is a practical, open extension of PolyMath to 18 mid-to-extreme low-resource languages across six families. The real contribution is the released dataset, pipeline, and multi-model numbers. The high- vs underrepresented-language math gap across 27 models is real but unsurprising.\n\nWhat they did well: They used a human-curated pipeline with native speakers validating pre-computed translations, covered a genuine range of resource levels, and open-sourced the data, acquisition pipeline, and evaluation framework. That is exactly the kind of barrier-lowering work the abstract promises. Benchmarking small, mid, large, and closed-source models gives a usable snapshot for the community. No circularity problem — scores come from independent model outputs on held-out items.\n\nSoft spots, in proportion: The load-bearing assumption is that native-speaker validation preserved mathematical content, difficulty, and answer format, not just surface fluency. The abstract says “thoroughly validated” but does not report IAA, residual error rates, or whether validators re-solved for unique correct answers under a fixed rubric. If validation was mostly linguistic, residual artifacts (wording ambiguity, format drift, template mismatch) could systematically depress underrepresented-language scores and inflate the apparent reasoning gap. The secondary link to instruction-following inherits the same risk if measured on the same prompts. These are reporting and design questions for referees, not reasons to dismiss the work. Novelty is incremental — an explicit extension of PolyMath — which is fine for a dataset paper of this type.\n\nWho this is for: Multilingual LLM evaluation and low-resource NLP people. A reading group that cares about coverage and eval hygiene will get value from the language selection and the model table. It deserves a serious referee; desk-rejecting careful open benchmarks like this is a mistake.\n\nRecommendation: Send to peer review. Ask referees to pressure-test the translation QA metrics, solution-preserving checks, and the instruction-following analysis. If the artifacts match the open-source claim and the validation holds up, this is a useful addition.\n\n— me","headline":"Useful open extension of PolyMath to 18 underrepresented languages; the resource is the contribution, the gap finding is expected.","tokens_in":3002,"tokens_out":532,"would_cite":true,"duration_ms":28847,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"PluraMath extends math-reasoning tests to 18 underrepresented languages and finds a persistent high-resource gap across 27 models","keywords":["multilingual mathematical reasoning","underrepresented languages","LLM evaluation","PolyMath extension","instruction following","low-resource languages","benchmark construction","reasoning LLMs"],"falsifier":"Re-translate a stratified sample of PluraMath items with an independent human pipeline (or back-translate and re-solve) and re-evaluate the same 27 models; if the high-versus-low-resource gap shrinks or vanishes while English scores stay stable, the original gap was an artifact of residual translation inequivalence.","tokens_in":2941,"feed_emoji":"🔢","tokens_out":772,"duration_ms":11789,"temperature":0.7,"pith_summary":"Mathematical reasoning is a standard way to test and tune large language models, but almost every public benchmark is still dominated by English, Chinese, and a few other high-resource languages. This paper introduces PluraMath, a carefully validated extension of the existing PolyMath suite that adds 18 more languages spanning six families, from mid-resource to extremely low-resource settings. Native speakers checked pre-computed translations so that problem content, difficulty, and answer format stay equivalent. The authors then run 27 reasoning models at four scales on the new set and show that scores remain systematically lower on the underrepresented languages. Better multilingual instruction-following, not raw scale alone, is the main factor that closes the gap. By open-sourcing the data, the translation pipeline, and the evaluation code, the work aims to make it easier for underrepresented language communities to build their own math-reasoning benchmarks.","feed_headline":"Math-reasoning gap stays wide for 18 underrepresented languages","feed_subtitle":"27 models show stronger scores track instruction-following more than scale on the new PluraMath suite","key_machinery":"A human-curated translation-and-validation pipeline that takes PolyMath problems, produces candidate translations, and has native speakers verify mathematical equivalence, difficulty, and answer format across 18 additional languages spanning six families.","core_discovery":"Across 27 reasoning LLMs spanning small, mid-size, large, and closed-source ensembles, mathematical reasoning accuracy on PluraMath remains markedly lower for the 18 newly added underrepresented languages than for high-resource ones; stronger results track mainly with a model’s ability to follow instructions in those languages rather than with parameter count alone.","pith_inferences":["If instruction-following is the dominant bottleneck, targeted multilingual instruction tuning on math templates may close more of the gap than additional pre-training on raw low-resource text.","The same validation pipeline could be reused for other structured reasoning domains (code, formal logic, scientific word problems) where surface form must not alter the underlying answer.","Persistent gaps on extreme low-resource languages may eventually force model developers to treat those languages as first-class evaluation axes rather than optional extras."],"forward_implications":["Benchmarking suites that ignore mid- and low-resource languages will systematically overstate the multilingual reasoning ability of current LLMs.","Improvements in multilingual instruction-following should raise math-reasoning scores on underrepresented languages more reliably than simply scaling model size.","Open release of the validated items and pipeline lowers the cost for other language communities to build parallel math-reasoning tests.","Closed-source ensembles that already show stronger instruction following are likely to retain their relative advantage on the new languages unless open models close the instruction gap."],"fun_headline_variants":["PluraMath shows math-reasoning gap remains for 18 underrepresented languages","Instruction-following beats model scale for math on low-resource languages","27 LLMs lag in math accuracy beyond high-resource languages on PluraMath","New PluraMath suite confirms persistent multilingual math-reasoning shortfalls","Model size alone fails to close math gaps in underrepresented languages"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Native-speaker checks of pre-computed translations are assumed to keep problem meaning, difficulty, and answer format equivalent enough that score gaps can be blamed on model reasoning rather than leftover translation or template artifacts.","fun_headline_variants_meta":{"raw":{"variants":["PluraMath shows math-reasoning gap remains for 18 underrepresented languages","Instruction-following beats model scale for math on low-resource languages","27 LLMs lag in math accuracy beyond high-resource languages on PluraMath","New PluraMath suite confirms persistent multilingual math-reasoning shortfalls","Model size alone fails to close math gaps in underrepresented languages"]},"model":"grok-4.5","cost_usd":0.009714,"raw_usage":{"total_tokens":2171,"prompt_tokens":780,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":97140000,"prompt_tokens_details":{"text_tokens":780,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1313,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":780,"tokens_out":78,"duration_ms":15064,"temperature":1.0,"reasoning_tokens":1313,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T19:05:19.599476+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-translate a stratified sample of PluraMath items with an independent human pipeline (or back-translate and re-solve) and re-evaluate the same 27 models; if the high-versus-low-resource gap shrinks or vanishes while English scores stay stable, the original gap was an artifact of residual translation inequivalence.","supporting_citations":[],"review_version":1}