{"id":"cd47352f-a8a4-4fc4-be75-ebbadf41a9d5","arxiv_id":"2608.11110","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"With five confounds corrected, four frontier tool-using models keep 71-73% of their action-policy consistency when the language changes, and the apparent small-model ordering is largely a chance-floor artifact.","lead":"Researchers measured whether AI agents take the same action steps for the same task in 41 languages, running 2.38 million agent rollouts across 8 models. They found that large frontier models retain about 71-73% of their action policy across languages, and that much apparent multilingual failure is caused by measurement artifacts such as trace-extraction regexes.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Model-independence claim rests on an invalid variance decomposition: 24 cells nested in 4 models are treated as independent, so the 5.7% model-share has no valid uncertainty and cannot distinguish a true constant from a 4-model coincidence.","rationale":"The paper is a large-scale, carefully controlled empirical study with an unusually honest treatment of confounds: it measures a chance floor by permutation, establishes trace-length causality by intervention, and reports failures (GPT-OSS, Aya-Expanse) as measurement results rather than discarding them. Those strengths are real and should be credited. The load-bearing weakness I identify is not the existence of the 71–73% spread but the statistical argument that upgrades four models to a 'model-independent quantity.' The paper's own Appendix S states that the per-cell decomposition (n=24, η²=5.7%) 'is the stronger form of the evidence and is what §4 relies on.' That decomposition is a one-way ANOVA over cells that share model identity, so the effective sample size for the model effect is four. Point estimates of variance components from four clusters are notoriously unstable, and no uncertainty is reported around 5.7%. A cluster bootstrap or REML analysis would likely produce a wide interval, potentially covering model shares that would undermine the 'model-independent' phrasing. This is an internal statistical issue, not an external-validity guess: it can be settled with the released per-cell data. The concrete test is therefore a re-analysis, not new experiments. If the CI is narrow and small, the claim survives; if not, the conclusion should be softened to a four-model observation. The reader's verdict of CONDITIONAL already accommodates this level of uncertainty, and the paper's other contributions — the protocol, the confound pricing, the causal pivot analysis, the measurement pathology — remain intact regardless. I therefore see no reason to change the verdict; CONDITIONAL is the right call, with the variance-decomposition analysis as a natural condition for acceptance.","tokens_in":39346,"tokens_out":10195,"duration_ms":87982,"concrete_test":"Using the released per-cell tilde-I values (Table 10 / Appendix R), fit a linear mixed model tilde-I ~ (1|model) + (1|benchmark) by REML and compute a 95% profile-likelihood CI for the model variance component and for the model share of total variance. Also run a cluster bootstrap that resamples whole models with replacement (e.g., 10,000 draws) and report the resulting CI for η²_model. If the CI upper bound exceeds roughly 30% — or if the bootstrap distribution is too wide to exclude large model effects — the 'model identity explains only 5.7%' statement is not statistically supported, and the conclusion should be downgraded from 'model-independent quantity' to 'four selected models agree under this protocol.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central upgrade — from 'four frontier models agree' to 'policy retention is close to a model-independent quantity' — rests on the variance decomposition in §4/Table 3: across 24 (model×benchmark) cells, model identity explains η²=5.7% of the variance in tilde-I. This statistic is computed as a one-way ANOVA over cells, but the 24 cells are not independent: they are nested within four models, and cells from the same model share that model's training recipe, scaffold behavior, and any model-specific confound. With only four clusters, the between-model variance component is estimated with enormous uncertainty; no confidence interval, significance test, or cluster-robust procedure is reported. The paper explicitly leans on this decomposition to answer the n=4 objection (§4: 'the data already answers it'; Appendix S calls it 'the stronger form of the evidence'). That is not valid: the effective sample size for model-level generalization is 4, not 24. A cluster bootstrap resampling whole models, or a REML mixed model with random model intercepts, would likely produce a CI for the model share spanning a wide range, possibly including values an order of magnitude larger. The large residual (67.4%), which in this design is mostly model×benchmark interaction, further shows that per-cell variation is not small. Absent a defensible uncertainty estimate, the headline 'explains only 5.7%' cannot support the model-independence claim; the data support only the weaker statement that the four selected models have similar raw tilde-I under this metric.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a measurement protocol for cross-lingual action-policy retention in tool-using agents. It defines same-language and cross-language trace agreement from paired replicates, corrects five confounds, and normalizes by self-consistency to obtain tilde-I = I_cross / I_within. Across 8 models, 6 parallel benchmarks, 41 languages and 2.38M rollouts, the study reports that under greedy decoding four frontier models retain 71-73% of their action policy across languages, with model identity explaining 5.7% of the per-cell variance. It further locates a breakdown below roughly 10B parameters, identifies an English pivot as the causal mechanism, and documents a trace-extraction regex artifact that can manufacture multilingual failure.","tokens_in":39574,"tokens_out":8771,"duration_ms":79740,"significance":"This is an unusually careful empirical study. The matched-replicate design, length matching in both directions, measured permutation chance floor, design checks, and release of code and traces are all strengths. The GPT-OSS regex case is a particularly convincing control and the causal length manipulation is valuable. If the model-independence claim survives statistical scrutiny, the 71-73% convergence is an important finding for multilingual agent evaluation. The main weakness is that the headline inference from 24 cells nested in four models to a 'model-independent quantity' is not supported by the current variance decomposition, and the chance-inclusive retention level is presented without sufficient emphasis on the chance-corrected 15-18% figure. These issues are fixable with reanalysis and more careful framing.","major_comments":[{"comment":"The claim that 'the data already answer' the n=4 objection is not supported by the reported variance decomposition. The 24 cells in Table 3 are nested within four models, so a one-way eta-squared computed on cells treats non-independent observations as independent; the effective sample size for any statement about generalization across models is four, not twenty-four. The task-level bootstrap intervals in Table 10 and Appendix I quantify within-model precision only, not model-level uncertainty, and Appendix S in fact concedes that the intervals are task-level. Please redo the decomposition with a cluster bootstrap resampling whole models, a mixed-effects model with random intercepts for model and benchmark, or a leave-one-model-out sensitivity analysis, and report the resulting uncertainty or range for the 5.7% figure and for the 71-73% band. The residual (67.4%) is not small and includes benchmark and model-by-benchmark variation, so it does not support the reading that per-cell variation is negligible. Without such an estimate, the abstract's 'model-independent quantity' is an overstatement; the data support 'four frontier models agree closely under the chosen metric and greedy decoding.'","section":"§4 and Table 3; Appendix S"},{"comment":"The paper states that all five confounds are removed, but the headline 71-73% retention is not chance-corrected. Appendix Q shows that the chance-corrected retention is 15-18%, roughly one-fifth of the headline value, because unrelated traces already agree at c about 0.56. Since the chance floor biases the ratio toward 1 and differentially across models, the chance-inclusive 71-73% should not be presented without this qualifier. Please either report the chance-corrected level as the primary retention figure while noting that the band is preserved, or explicitly state in the abstract and Section 4 that 71-73% is chance-inclusive and that the above-chance retention is 15-18%.","section":"Abstract and Appendix Q"},{"comment":"The model-independence claim is also conditional on the choice of the trace-similarity metric S. Appendix S acknowledges that no alternative metric was tested, but S is not a neutral measurement: C2 shows the metric is strongly length-sensitive, and a set-based or edit-distance family could spread the four models' ratios beyond the 2.6-point band. Because the abstract promotes 'close to a model-independent quantity,' a sensitivity analysis with at least one alternative S, or a clear restriction of the claim to matching-block similarity, is needed before the stronger form of the claim is made.","section":"§4 and Appendix S"}],"minor_comments":[{"comment":"Section 1 says 'Every correction we apply makes it larger, never smaller,' but Section 6 reports that self-consistency voting and the scale extension remove results and lower tilde-I; please align the wording, for example by saying 'every confound correction' instead of 'every correction.'","section":"§1 and §6"},{"comment":"The term 'pre-registered' is used to mean predictions written into the analysis script before compute was spent; this is not an external preregistration. Please say so explicitly and describe how the script version and timestamps were fixed, so readers do not infer a formal registry.","section":"Appendix N and §5"},{"comment":"The statement that the per-cell decomposition is 'the stronger form of the evidence' should be reconciled with the acknowledged n=4 limitation; the main text should not resolve the n=4 objection by invoking the same decomposition without reporting its model-level uncertainty.","section":"Appendix S"}],"recommendation":"major_revision","confidential_remarks":"I would not reject this paper: the empirical core is strong, the design checks are exemplary, and the GPT-OSS regex case is a valuable cautionary result. My main concern is that the abstract and Section 4 overclaim model independence on the basis of a nested variance decomposition with only four models, and that the chance-inclusive 71-73% headline is presented without sufficiently foregrounding the 15-18% chance-corrected level. Both are fixable: a cluster-robust or mixed-model reanalysis, a leave-one-model-out sensitivity check, and reframed wording would make the paper's claims match the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Most of this paper is as good as empirical NLP gets. The matched-replicate estimand is the right move; the five confounds are priced and two are established causally; the permutation null for the chance floor is real; and the GPT-OSS regex case is a genuinely useful warning. Releasing 2.38M traces, prompts, and code makes it checkable. This deserves a serious referee.\n\nThe soft spot is the leap from 'four frontier models land at 71-73%' to 'policy retention is close to a model-independent quantity.' The four-model agreement is a fact: per-model means under greedy decoding are tightly clustered and the task-level intervals are narrow. But the paper then tries to answer the n=4 objection with a variance decomposition over 24 cells, claiming model identity explains only 5.7% of the variance. That statistic treats the cells as independent, and they are not: each model contributes six cells that share its training, scaffold behavior, and confounds. With four clusters, the between-model variance component has enormous uncertainty. A cluster bootstrap over models or a REML mixed model would likely give a wide interval on that 5.7%. So the 'data already answers it' claim in §4 is not valid. The data support 'these four specific models agree'; they do not support a universal constant. The paper half-acknowledges this in Appendix S, but the abstract and §4 state the stronger version.\n\nOther weaknesses are minor. The single similarity metric is acknowledged. The removal direction of the pivot ablation does not generalize cleanly, but the pre-registered mandate ordering is good. Voting is at one temperature. None threaten the main measurement.\n\nBottom line: a careful, reproducible study that genuinely reframes how to compare multilingual agents. The central empirical result — normalized cross-lingual retention around 0.71-0.73 for four current frontier models — is well supported. The model-independence claim needs to be pulled back, and the η² over nested cells should be replaced or dropped. That is a revision, not a rejection.","headline":"Strong empirical study, but the 'model-independent' claim rests on an invalid variance decomposition; the four-model agreement stands, the universal constant does not.","tokens_in":40180,"tokens_out":3194,"would_cite":true,"duration_ms":27751,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that after removing five measurement confounds, frontier tool-using agents retain 71–73% of their action policy when the task language changes, and model identity explains only 5.7% of the variance across cells.","keywords":["cross-lingual policy retention","tool-using agents","action traces","trace similarity","evaluation confounds","English pivot","model reproducibility","multilingual evaluation"],"falsifier":"Run the same matched-replicate protocol at greedy decoding on a new frontier model that is not one of the four (for instance, a model from a fourth vendor that emits parseable traces), and compute $\\tilde{I}$ pooled over the six benchmarks; if the value falls outside the [0.708, 0.733] band, the model-independence claim is refuted. Alternatively, replace the matching-block similarity with an edit distance under a tool-cost matrix and check whether the four models still land within 2.6 points of each other; the paper's own Appendix S notes this is untested.","tokens_in":39094,"feed_emoji":"🌐","tokens_out":5979,"duration_ms":40680,"temperature":0.7,"pith_summary":"This paper argues that the conventional way of evaluating multilingual tool-using agents—comparing final answers and throwing away the intermediate action traces—misses the product that matters: the sequence of tool calls that determines cost, failure modes, and auditability. It rebuilds the measurement so that the action policy itself is compared across languages, removing five confounds: a missing same-language baseline, trace-length bias, empty-trace inflation, a reproducibility ceiling, and a chance floor measured by permutation. After correcting all five, the paper finds that under greedy decoding four very different frontier models each retain 71–73% of their own policy consistency across languages, with model identity explaining only 5.7% of the variance across 24 model-benchmark cells. The same measurement shows that the apparent ordering among smaller models is largely an artifact of the chance floor, and that agents route non-English tasks through an English pivot that is causally load-bearing and cannot be prompted away. A single trace-extraction regex is shown to have manufactured an apparent multilingual failure, raising one model's measured accuracy twenty-sixfold.","feed_headline":"Frontier agents keep 71-73% of their policy across languages","feed_subtitle":"Five confounds corrected, model identity explains just 5.7% of the variance; smaller models break the pattern.","key_machinery":"The load-bearing object is the ceiling-corrected estimand $\\tilde{I}=I_{\\mathrm{cross}}/I_{\\mathrm{within}}$, computed from matched replicates: every (model, benchmark, language) cell is generated twice under identical decoding and token budget, varying only the serving seed, so that same-language agreement $I_{\\mathrm{within}}$ and cross-language agreement $I_{\\mathrm{cross}}$ are both cross-seed and differ only in language. The protocol removes the five confounds by strict empty-trace exclusion, two-directional length matching, a permutation-measured chance floor, and the division by $I_{\\mathrm{within}}$ itself, with task-level bootstrap intervals. This machinery is what converts a raw gap into a quantity that can be compared across models.","core_discovery":"The paper's central discovery is that cross-lingual policy retention—defined as $\\tilde{I}=I_{\\mathrm{cross}}/I_{\\mathrm{within}}$, the share of a model's own same-language reproducibility that survives a change of language—is nearly model-independent at the frontier. Under greedy decoding, four models spanning dense and mixture-of-experts architectures, 24B–235B parameters, and different training recipes each land between 71% and 73% retention, with the four-model band only 2.6 percentage points wide and model identity explaining 5.7% of the variance across 24 cells. The paper further shows that this result emerges only after correcting five confounds, each of which otherwise biases or reverses conclusions: a missing same-language baseline, trace-length sensitivity, empty-trace inflation, the reproducibility ceiling, and a chance floor measured by permutation at $c\\approx 0.56$. It locates the boundary of the regularity below roughly 10B parameters, where the raw ordering of models is largely a chance-floor artifact, and it identifies a causal mechanism: agents pivot non-English tasks through English, a behaviour that is load-bearing for cross-lingual agreement and that resists direct instruction to abandon it.","pith_inferences":["If the 71–73% band is a genuine frontier property, a natural testable extension is to run the same protocol on newly released frontier models; the prediction is that they land inside the band, and a miss would bound the regime.","The chance-floor correction implies that the above-chance retention is roughly 15–18% rather than 71–73%, so downstream citations of the headline figure should state which scale they mean; this is a direct arithmetic consequence of the paper's own correction.","The causal length dose-response suggests benchmark designers can move retention by changing trace length, so cross-benchmark comparisons are not meaningful without length control; this is our editorial inference, not stated as a recommendation by the paper.","The English-pivot result connects to latent-English interpretability work, but extends it: prompting is unlikely to remove the pivot, so deployment in non-English settings should budget for translation-related costs and failure modes."],"forward_implications":["Final-answer accuracy is not behavioural agreement: two languages can agree on every answer while the action route differs in cost, failure mode, and auditability.","Uncorrected cross-lingual invariance gaps rank models by determinism rather than by language robustness; the observed rank correlation between the sampling-inclusive and greedy gaps is $\\rho=-0.80$.","The English pivot is causally load-bearing: removing it lowers length-matched cross-lingual agreement in proportion to usage, and instructed abandonment fails at over 99% non-compliance.","Self-consistency voting is a variance reducer, not a retention improver: on the ceiling-corrected estimand it costs 1.6–1.9 points with disjoint intervals.","Trace-extraction parsing must be reported alongside every headline number, since a single regex changed one model's measured accuracy twenty-sixfold."],"supporting_citations":[{"why":"Supplies the ReAct-style scaffold whose Thought/Action trace format this study adopts and makes the measured object.","marker":"Yao et al., 2023"},{"why":"The closest prior work asking whether models retrieve the same facts across languages; this paper extends the question from facts to procedures and adds the matched baseline.","marker":"Qi et al., 2023"},{"why":"Source of FLORES-200, one of the six parallel benchmarks; provides the verified sentence-level alignment keys used for cross-lingual task identity.","marker":"NLLB Team et al., 2022"},{"why":"Source of XQuAD, a parallel extractive QA benchmark used in the suite.","marker":"Artetxe et al., 2020"},{"why":"Source of XNLI, a parallel natural-language-inference benchmark used in the suite.","marker":"Conneau et al., 2018"},{"why":"Source of Belebele, a 122-language parallel reading-comprehension benchmark; required re-keying before index joins were valid.","marker":"Bandarkar et al., 2024"},{"why":"Source of XCOPA, a parallel causal-commonsense benchmark; contributes the only suite with no English arm.","marker":"Ponti et al., 2020"},{"why":"Representational evidence that multilingual transformers route through English-aligned latent states; provides the hypothesis the paper tests behaviourally with the pivot ablation.","marker":"Wendler et al., 2024"},{"why":"Self-consistency voting method; the paper tests whether voting improves cross-lingual retention and finds it does not on the ceiling-corrected estimand.","marker":"Wang et al., 2023"}],"fun_headline_variants":["Frontier agents keep 71-73% of policy across languages","Cross-lingual policy retention: frontier models at 71-73%","Agents retain policies across languages: 71-73% at frontier","Policy retention across 41 languages: frontier models converge","Tool-using agents keep 71-73% of actions across languages"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model-independence claim rests on treating the 24 (model, benchmark) cells as independent in a variance decomposition even though they are nested within four models and use a single fixed trace-similarity metric; if those models are not representative, or a different similarity metric would spread the ratios, the 71–73% convergence could be an artifact of selection and metric rather than a property of frontier policies.","fun_headline_variants_meta":{"raw":{"variants":["Frontier agents keep 71-73% of policy across languages","Cross-lingual policy retention: frontier models at 71-73%","Agents retain policies across languages: 71-73% at frontier","Policy retention across 41 languages: frontier models converge","Tool-using agents keep 71-73% of actions across languages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1493,"prompt_tokens":1146,"completion_tokens":347,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":762,"completion_tokens_details":{"reasoning_tokens":255}},"tokens_in":762,"tokens_out":347,"duration_ms":3104,"temperature":1.0,"reasoning_tokens":255,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:59:47.624897+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same matched-replicate protocol at greedy decoding on a new frontier model that is not one of the four (for instance, a model from a fourth vendor that emits parseable traces), and compute $\\tilde{I}$ pooled over the six benchmarks; if the value falls outside the [0.708, 0.733] band, the model-independence claim is refuted. Alternatively, replace the matching-block similarity with an edit distance under a tool-cost matrix and check whether the four models still land within 2.6 points of each other; the paper's own Appendix S notes this is untested.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of XNLI, a parallel natural-language-inference benchmark used in the suite."},{"cited_title":"XCOPA: A multilingual dataset for causal commonsense reasoning","cited_arxiv_id":null,"evidence_quote":"Source of XCOPA, a parallel causal-commonsense benchmark; contributes the only suite with no English arm."}],"review_version":1}