{"id":"eab326e0-52d4-40cc-8889-cefaefb44273","arxiv_id":"2608.04160","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The measured native-versus-translate reasoning gap on MGSM depends strongly on the output-token budget, nearly vanishing at saturation and reversing direction under tight caps.","lead":"Multilingual AI evaluations usually report accuracy at one output length, and this paper shows that choice is a hidden variable. On three MGSM languages, the native-versus-translate gap swings by up to 57 points depending on the output-token budget, and at tight caps can even reverse which strategy scores better.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Load-bearing concern: the Eq. (1) artifact is defined through FLORES-200 premiums, but the paper's Appendix D shows the Qwen Swahili 5-point peak disappears under the behavioral trace-length ratio; the promised same-content trace-premium validation is still outstanding.","rationale":"The paper is unusually careful: internal freezes, independent decodes, audits, and explicit limitations. I looked for a more fundamental flaw and did not find one. The independent replication at pre-specified peak budgets is strong evidence that the replay-frame peaks are not a shared-trajectory artifact. The residual weakest point is the normalizer: Eq. (1) makes r the definition of the artifact, and the paper's own sensitivity analysis isolates a cell where the conclusion flips under a plausible alternative. Because the authors themselves flag the same-content trace-premium validation as outstanding, this is a recognized open assumption rather than a hidden error. Keeping the reader's CONDITIONAL verdict is appropriate; releasing the artifacts and completing that validation would resolve the concern.","tokens_in":18172,"tokens_out":6443,"duration_ms":69071,"concrete_test":"Run the outstanding same-content trace-premium validation: for each MGSM item, take the stored Qwen NATIVE reasoning trace and create an English rendering of the same reasoning content (e.g., translate the trace with the same model or a strong MT system), tokenize both sides with Qwen's tokenizer, and estimate r_trace as the L/English token ratio. Recompute Δ_L(B) for B=128/192/256 with r_trace in Eq. (1). If Qwen Swahili's peak drops below 5 points while German and Thai remain above 5, the confirmatory peaks are normalizer-dependent and the artifact claim must be scoped to FLORES-based normalization; if r_trace matches FLORES and all three peaks exceed 5, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central estimand is Δ_L(B)=acc_N(⌊r_{m,L}B⌋)−acc_N(B) (Eq. 1), so every reported artifact magnitude inherits the FLORES-200 token premium r_{m,L} as a length normalizer for MGSM reasoning traces. That normalizer is load-bearing precisely where the quantitative claim is strongest: the three confirmatory Qwen peaks. Appendix D shows that the Qwen Swahili 5-point artifact requires a premium of at least 1.254, while the behavioral NATIVE-to-TRANSLATE trace-length ratio is 1.179; substituting the behavioral ratio would leave Swahili below the 5-point threshold, and Llama Thai never reaches 5 points even at 1.5× FLORES. The authors acknowledge in the Limitations that the 'same-content trace-premium validation' remains outstanding, so the choice between FLORES and a reasoning-trace-specific premium is unresolved. If the correct same-content normalizer is closer to the behavioral ratio for Swahili, the headline peak of 15.0 points (independent 13.7) is not a budget artifact of the reported size. The qualitative budget-regime dependence is supported independently by the raw gap curves and the announcement experiment, so this does not overturn the paper; it does mean the central quantitative conclusion is conditional on an untested normalizer.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes treating the output-token cap as an independent variable in multilingual evaluation. It defines a length-normalized estimand Δ_L(B)=accN(⌊r_{m,L}B⌋)−accN(B), where r is the FLORES-200 token premium, so the native-vs-translate gap becomes a finite increment of the native accuracy curve. On MGSM with Qwen3-8B and Llama-3.1-8B-Instruct, it reports that the gap is near zero at the frozen B*=1024 (Qwen de 0.00, th 0.15, sw 0.05) while peaking at tight budgets (de 34.2, th 38.9, sw 15.0), with peaks replicated on 540,000 independent hard-capped decodes. It also reports that announcing a budget changes behavior at a fixed enforced cap (one significant cell, Thai NATIVE +5.1), that vocabulary extension closes the gap only where truncation still binds, and that a correct-emission sub-CDF identity from one long-cap run tracks the peaks. The paper is explicitly methodology-focused, with audits on language ID, translation quality, parser robustness, decoder parity, and normalizer sensitivity.","tokens_in":18395,"tokens_out":3839,"duration_ms":31733,"significance":"If correct, the paper's central claim is important: single-budget accuracy measurements do not provide a stable estimate of the multilingual reasoning gap, and the cap should be swept and length-normalized. The paper has genuine methodological strengths: the estimand is an identity, the confirmatory families were pre-specified and frozen, the independent hard-capped replication is substantial (540k decodes), the six-test Holm family is reported with tail-conservatism calibration, and the paper is unusually transparent about its own circularity and limitations (e.g., the sub-CDF agreement is explicitly flagged as agreeing by construction with replay and is checked against independent decodes). The announcement experiment is a clean design that isolates disclosure from truncation. The adaptation triage is honestly scoped as a heuristic. These strengths make the paper publishable if the normalizer concern is resolved or adequately bounded.","major_comments":[{"comment":"The central quantitative claim is conditional on a load-bearing normalizer. Equation (1) defines the budget artifact entirely through the FLORES-200 premium r_{m,L}, but Appendix D shows the Qwen Swahili 5-point artifact disappears if the observed behavioral trace-length ratio (1.179) is used instead of the FLORES premium (threshold 1.254), and Llama Thai never reaches 5 points even at 1.5×FLORES. The Limitations explicitly state that the same-content trace-premium validation is outstanding. Please either supply that validation (e.g., measuring token premiums on parallel MGSM item content in both languages for the exact NATIVE/TRANSLATE-ACT traces) or relegate the Swahili 5-point claim and any magnitude claims that inherit the FLORES premium to exploratory status; the qualitative conclusion that budget regimes change the measured gap is not affected by this issue.","section":"§4.2/Appendix D"},{"comment":"The announcement experiment's headline cell (Thai NATIVE +5.1 points) is one significant cell out of four, and the direction is described as 'announcing a tighter budget raised accuracy', which the paper reports honestly. However, the conclusion in the abstract that 'accuracy is not a function of the enforced cap alone' is stated as a general consequence of one significant cell. I would ask for a more careful wording: the data support an effect in one cell, not a general statement, though the monotone dose-response in that cell (63.20/59.90/58.10) is suggestive. This is a presentation issue but it matters for how the abstract is read.","section":"Limitations/§4.4"},{"comment":"The adaptation triage conclusion that the vocabulary extension closes the gap only where truncation binds is supported by the data, but the scope caveats in Appendix G weaken the generality more than the main text acknowledges: the extension is trained on NATIVE traces only, the English control is aggregate (not per-string), the fold split is fixed, and G(4096) is the largest stored prefix, not a non-binding regime. Since the section is explicitly scoped as a heuristic, this is not a blocking issue, but I would ask that the first two caveats be moved into the main text of §6, because Table 3 is easily over-read as a general cost-benefit result.","section":"§6/Appendix G"},{"comment":"The tail-conservatism factor is described as 'a conservative safeguard rather than a verified family-wise calibration', and the corrected type-I rate (0.00917) exceeds the target (0.00833). This is an honest and unusual disclosure, but the reader should be told whether the 1.3 factor was pre-specified before seeing the data or chosen after observing the anti-conservatism. The text says 'pre-specified 1.3× tail-conservatism factor' but also says calibration 'found mild anti-conservatism', which is ambiguous. Please clarify the chronology; if the factor was chosen after the fact, the confirmatory status of the H1 and H3 tests is weaker than presented.","section":"§4.1/Appendix E"}],"minor_comments":[{"comment":"The phrase 'length normalization moves it by up to 38.9 points where the cap binds' is a replay-ledger quantity; the abstract later mentions 540,000 independent decodes, so consider clarifying that the 38.9 is from the stored-prefix sweep, not the independent sample.","section":"Abstract"},{"comment":"The phrase 'This procedureally matched secondary analysis' has a typo ('procedureally' for 'procedurally') and the sentence 'also rejects nothing' is oddly phrased given that the test was not designed to reject anything; consider 'also yields no rejection'.","section":"§4.1"},{"comment":"The Swahili row of Table 1 cites NATIVE accuracy 8.70 versus TRANSLATE-ACT 0.60 at B=128, a degenerate floor cell. The main text cautions about this, but the table note should repeat it because the table alone is misleading without it.","section":"Table 1 / §4.2"},{"comment":"The sentence 'Qwen Swahili is the exception: its threshold of 1.254 sits above the behavioral ratio of 1.179' is load-bearing and should be echoed in §4.2, not only in the appendix, for the reasons in my major comment 1.","section":"Appendix D"},{"comment":"The sentence 'We cannot price the prompting rung, because adopting TRANSLATE-ACT closes G by definition' is confusing: adopting TRANSLATE-ACT replaces the NATIVE arm rather than being an incremental rung. Please rephrase to say that the comparator is the strategy itself, so it cannot be priced as a separate intervention.","section":"§6"},{"comment":"'capped_eos false' and 'capped_eos true' are used without definition; define them at first use.","section":"Appendix H"},{"comment":"Several 2026 preprints appear in the reference list without in-text citation numbers (e.g., Lasbordes et al., Liu et al., Su et al., Silvestri and Cetin, Zhang et al.). Please reconcile citations and list entries.","section":"References"},{"comment":"The relationship between the exploratory blind LLM adjudication and the outstanding frozen human GlotLID validation should be stated once in one place; currently the reader must piece it together from §5 and the Limitations.","section":"§5/Limitations"}],"recommendation":"major_revision","confidential_remarks":"This paper is well above the quality bar in its internal discipline: the estimand identity, pre-specified families, independent decodes, and explicit circularity flagging are a model for the field. The block is narrow: the FLORES-premium normalizer is load-bearing for the peak magnitudes, and the paper itself admits the same-content validation is outstanding. That admission is commendable but it does not discharge the obligation for the main claim. If the authors can either provide the same-content premium validation or restrict the confirmatory claim to the qualitative regime-dependence and the frozen equivalence at B*=1024 (which does not depend on the normalizer at all), I would support acceptance. I would also suggest the editorial team check whether the 2026-dated references comply with the journal's preprint policy, though that is not a scientific issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you do multilingual evaluation. The paper's central claim—that the measured native-vs-translate gap on MGSM depends on the output-token budget—is well supported. The identity Δ_L(B) = acc_N(⌊rB⌋) − acc_N(B) is a simple rearrangement, but the paper turns it into a useful diagnostic and then backs it with unusually careful empirics. The strongest parts are the prospective freezing of confirmatory families (git tags before data), the independent hard-capped decodes (540k records) that reproduce the three Qwen peaks, and the transparent handling of the B*=1024 null, which the paper treats as a saturation boundary condition rather than evidence of absence. The announcement effect at a fixed enforced cap (Thai native accuracy moves 5.1 points under an announced budget change) is a genuinely nice result. The correct-emission sub-CDF consistency check is also handled honestly: the authors explicitly say replay agreement is by construction, so they check it against the independent decodes instead.\n\nThe soft spots are real but not fatal. The load-bearing normalizer is the FLORES-200 token premium r_{m,L}, computed on parallel devtest sentences, used to length-normalize MGSM reasoning traces. Equation (1) defines the budget artifact through that premium, so every reported magnitude inherits it. The paper's own Appendix D shows the Qwen Swahili 5-point artifact requires a premium above the behavioral trace-length ratio; under the behavioral ratio that cell drops below 5 points. Llama Thai never reaches 5 points even at 1.5× FLORES. The authors flag the same-content trace-premium validation as outstanding, so the choice between FLORES and a reasoning-trace-specific premium is unresolved. That means the quantitative headline numbers are conditional, not settled. The qualitative regime-dependence, however, is supported by the raw gap curves and the announcement experiment, so the central message stands.\n\nSecond, data and code are not actually released yet; the paper says it intends to release them. That blocks exact reproduction now. Third, the vocabulary-extension triage in §6 is exploratory and properly scoped, so it shouldn't be over-read.\n\nOverall, this is a sincere, well-audited paper that deserves a serious referee. Send it to review, and make the normalizer sensitivity and the release of artifacts the two conditions for acceptance. I would bring it to reading group and would cite it if I worked on multilingual evaluation.","headline":"A careful, mostly convincing demonstration of budget-regime dependence in the MGSM gap, with the exact magnitudes conditional on an unvalidated length normalizer.","tokens_in":18966,"tokens_out":2735,"would_cite":true,"duration_ms":27079,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The native-versus-translate gap on MGSM is a token-budget artifact, not a stable measure of reasoning ability.","keywords":["multilingual reasoning gap","token budget","output cap","MGSM","FLORES-200","chain-of-thought","length normalization","budget sweep"],"falsifier":"Retake the Qwen Swahili peak (discovered at $B = 128$) but scale the native budget by the measured behavioral trace-length ratio (1.179) instead of the FLORES premium (1.94); the paper's analysis predicts the artifact falls below 5 points. If independently hard-capped decodes at the trace-length-scaled cap still show a gap above 5 points, the premium-based budget-artifact explanation for that cell would be wrong.","tokens_in":17897,"feed_emoji":"📏","tokens_out":9108,"duration_ms":75070,"temperature":0.7,"pith_summary":"The paper claims that the native-versus-translate gap on the MGSM reasoning benchmark is substantially a token-budget artifact: languages need different numbers of tokens to express the same content, so a fixed output cap penalizes the native language. The gap is shown to equal the gain in native accuracy obtained by scaling the cap by the language's token premium; at a large frozen budget the gap is near zero (0.00–0.15 points for Qwen), while at tight caps it peaks at 15.0–38.9 points, and those peaks were prospectively confirmed on 540,000 independently hard-capped decodes. If the paper is right, single-budget accuracy measurements do not give a stable estimate of the multilingual reasoning gap, and evaluators should report accuracy across the budget regime rather than at one cap.","feed_headline":"Move the token cap and the multilingual reasoning gap moves with it","feed_subtitle":"Measured at a large budget the native-vs-translate gap drops to near zero on MGSM; single-budget scores are misleading.","key_machinery":"The load-bearing object is the estimand identity $\\Delta_L(B) = \\mathrm{acc}_N(\\lfloor r_{m,L} B \\rfloor) - \\mathrm{acc}_N(B)$, with $r_{m,L}$ the FLORES-200 token premium: the ratio of tokens the model's own tokenizer needs for parallel devtest content in language $L$ versus English. This reduces the gap to a finite increment of the native accuracy curve, so the gap's peak location follows the native answer-emission distribution and its height depends on the premium-scaled interval. A companion identity, $\\Delta_L(B) = G(\\lfloor r B \\rfloor) - G(B)$ with $G(t) = P(C = 1, E \\le t)$ the correct-emission sub-CDF, predicts the peak from one long-cap run to within 0.65 points on the three pre-specified MGSM cells. The empirical machinery is the ledger of stored 4096-token generations scored at every prefix budget, paired with an independent sample in which every budget is decoded under its own hard cap.","core_discovery":"The central claim is that the measured native-vs-translate gap on MGSM is not a stable quantity but a function of the output-token budget, formalized by the identity $\\Delta_L(B) = \\mathrm{acc}_N(\\lfloor r_{m,L} B \\rfloor) - \\mathrm{acc}_N(B)$, where $r_{m,L}$ is the FLORES-200 token premium of language $L$ over English for model $m$. The gap is therefore just how many more native answers become correct when the native budget is scaled by the premium; the translate arm cancels out. At the frozen budget $B^* = 1024$ the gap is near zero for all three Qwen languages and all six confirmatory tests fail to reject, whereas at tight caps the gap reaches 34.2 (German), 38.9 (Thai), and 15.0 points (Swahili), and these three peaks survive on an independent sample of 540,000 hard-capped decodes. A third family of tests varies only the announced budget at a fixed enforced cap and finds that announcing 128 rather than 2048 tokens moves Thai native accuracy by 5.1 points, showing that accuracy is not a function of the enforced cap alone. The residual difference above saturation is interpreted as a strategy-performance gap, not an identified reasoning deficit.","pith_inferences":["If the budget-artifact account generalizes, prior single-budget estimates of multilingual reasoning gaps on other benchmarks should be re-read with suspicion; a minimal re-analysis would report native accuracy at $B$ and at the premium-scaled cap, which is enough to bound the artifact.","The announcement effect suggests that budget disclosure is a behavioral variable in deployed systems: the same enforced cap can yield different accuracy depending on what the prompt says, so benchmark comparisons that include the cap in the prompt are comparing more than the model's reasoning.","The correct-emission sub-CDF identity offers a cheap diagnostic for choosing evaluation budgets: one long-cap run can locate where the budget artifact peaks, so future work could pre-register peak budgets from a pilot run instead of sweeping many budgets.","The vocabulary-extension result implies that token-count interventions are only as valuable as the answer-emission tail they rescue; a production team deciding between a bigger cap and a language-specific tokenizer should first measure how much accuracy longer prefixes actually recover on its own traces."],"forward_implications":["Single-budget accuracy estimates of the multilingual reasoning gap are unstable: the measured gap swings by up to 57 points across budgets and can even reverse which strategy scores higher at tight caps.","Length-normalizing the native budget with the FLORES-200 premium moves the gap by up to 38.9 points where the cap binds, and the effect vanishes once native accuracy saturates.","Vocabulary extension closes the gap only where the cap still truncates: 0.0 points at the frozen budget, up to 4.9 points where 19% of traces still truncate.","Accuracy is not a function of the enforced cap alone: announcing a budget changes behavior at a fixed cap, so a deployment that discloses its budget must evaluate under the disclosure it ships.","Evaluations should treat the output cap as an independent variable and report accuracy across a budget sweep, marking the region where answer emission binds."],"supporting_citations":[{"why":"Introduces MGSM and the native-vs-translate evaluation setup that the paper re-examines under output budgets.","marker":"Shi et al., 2023"},{"why":"Introduces the FLORES-200 benchmark from which the token premiums are computed.","marker":"NLLB Team et al., 2022"},{"why":"Establishes the FLORES parallel-text lineage used as the length proxy.","marker":"Goyal et al., 2022"},{"why":"Documents that tokenization cost differs across languages, motivating the premium-scaled budgets.","marker":"Ahia et al., 2023"},{"why":"Shows tokenizers introduce unfairness between languages, supporting the claim that a fixed cap is a hidden variable.","marker":"Petrov et al., 2023"},{"why":"Supplies the budget-forcing and test-time-scaling context that the paper's budget sweeps build on.","marker":"Muennighoff et al., 2025"}],"fun_headline_variants":["Token budget swings the multilingual reasoning gap by 57 points","Multilingual reasoning gap is a token-budget artifact","At tight caps the gap is 39 points; at large caps it's zero","Accuracy at a single token cap is misleading"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the FLORES-200 token premium, computed on parallel devtest sentences, is the right length normalizer for MGSM reasoning traces; if the true trace-length ratio differs, the paper's own sensitivity analysis shows the artifact's magnitude and even its existence change in specific cells (Qwen Swahili depends on the premium exceeding the behavioral ratio of 1.254 versus 1.179).","fun_headline_variants_meta":{"raw":{"variants":["Token budget swings the multilingual reasoning gap by 57 points","Multilingual reasoning gap is a token-budget artifact","At tight caps the gap is 39 points; at large caps it's zero","Accuracy at a single token cap is misleading"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000783,"raw_usage":{"total_tokens":3586,"prompt_tokens":1206,"completion_tokens":2380,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":822,"completion_tokens_details":{"reasoning_tokens":2312}},"tokens_in":822,"tokens_out":2380,"duration_ms":17297,"temperature":1.0,"reasoning_tokens":2312,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T00:22:01.073596+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retake the Qwen Swahili peak (discovered at $B = 128$) but scale the native budget by the measured behavioral trace-length ratio (1.179) instead of the FLORES premium (1.94); the paper's analysis predicts the artifact falls below 5 points. If independently hard-capped decodes at the trace-length-scaled cap still show a gap above 5 points, the premium-based budget-artifact explanation for that cell would be wrong.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the FLORES parallel-text lineage used as the length proxy."},{"cited_title":"Mortensen, Noah A","cited_arxiv_id":null,"evidence_quote":"Documents that tokenization cost differs across languages, motivating the premium-scaled budgets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows tokenizers introduce unfairness between languages, supporting the claim that a fixed cap is a hidden variable."}],"review_version":1}