{"id":"33809d88-416e-47d0-bbe7-f4ea91e0c29b","arxiv_id":"2603.17945","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"ShapleyLaw treats languages as players in a cooperative game and uses Shapley values to predict test loss and optimize multilingual pretraining mixture ratios.","lead":"The paper models multilingual pretraining as a cooperative game and uses Shapley values to measure each language’s contribution, including cross-lingual transfer. It claims this yields better loss predictions and better language-mixture ratios than prior multilingual scaling laws.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The cooperative-game framing assumes Shapley values of loss-reduction coalitions correctly quantify cross-lingual transfer for continuous mixture optimization inside a scaling law, but this mapping is unmotivated and unverifiable from the abstract alone.","rationale":"The reader's weakest_assumption already isolates the precise load-bearing step—the unmotivated identification of Shapley values of loss-reduction coalitions with the quantities needed for mixture optimization. Because only the abstract is present, no further technical flaw (or confirmation) can be located; the same unverifiability that produced the UNVERDICTED / LOW-confidence judgment therefore remains. No adjustment to the verdict is warranted, and the concrete test simply operationalizes the missing formal check.","tokens_in":1934,"tokens_out":468,"duration_ms":16346,"concrete_test":"If the full paper or code becomes available, extract the exact definition of v(S) and the formula that inserts the Shapley vector into the scaling law; recompute exact Shapley values for any 3-language subset whose per-coalition losses are reported, re-optimize the mixture ratios under that law, and verify whether the predicted test-loss improvement over the strongest baseline exceeds the experimental standard error on a held-out language set.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that multilingual pretraining is a cooperative game with languages as players, characteristic function v(S) equal to test-loss reduction on coalitions S, and Shapley values (average marginal contributions) the right scalars to insert into a predictive scaling law that then yields superior mixture ratios. This rests on three unstated conditions: (1) v(S) is well-defined and independent of the complementary languages and of training dynamics outside S; (2) the resulting values satisfy the transferable-utility axioms needed for continuous ratio optimization rather than discrete set selection; (3) the functional form that embeds those values recovers ordinary monolingual scaling when transfer vanishes and improves prediction when it does not. The abstract supplies none of the formal derivation, approximation scheme for the 2^n evaluations, or consistency checks, so the claimed outperformance in loss prediction and mixture optimization has no demonstrated grounding.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript proposes ShapleyLaw, a multilingual scaling law obtained by casting multilingual pretraining as a cooperative game whose players are languages and whose characteristic function is the reduction in test loss realized by coalitions of languages. Shapley values of this game are taken to quantify each language’s contribution, including cross-lingual transfer, and are then inserted into a scaling law that predicts test loss under arbitrary language-mixture ratios and that is used to optimize those ratios. The abstract reports that ShapleyLaw outperforms existing multilingual scaling-law baselines on both loss prediction and mixture optimization.","tokens_in":2181,"tokens_out":878,"duration_ms":19474,"significance":"If the formal construction is sound and the empirical gains hold, the work would supply a principled, game-theoretic account of cross-lingual transfer inside multilingual scaling laws—an important and currently under-modeled factor in data-mixture design. Explicit transfer quantification via Shapley values, together with claimed improvements in both prediction and optimization, would be of practical interest for multilingual pretraining and of conceptual interest at the intersection of cooperative game theory and scaling laws. Because only the abstract is available, however, neither the derivation nor the experimental evidence can be assessed, so the significance remains conditional.","major_comments":[{"comment":"Abstract: the central claim rests on treating multilingual pretraining as a cooperative game with characteristic function v(S) equal to test-loss reduction for coalition S. For the resulting Shapley values to isolate cross-lingual transfer usable in continuous mixture optimization, v(S) must be well-defined and essentially independent of languages outside S and of training dynamics beyond the coalition. The abstract supplies neither a formal definition of v nor any consistency check of this independence; without them the mapping from discrete coalitions to continuous ratios is unmotivated and the claimed superiority over baselines cannot be evaluated.","section":"Abstract"},{"comment":"Abstract: embedding discrete Shapley values into a predictive scaling law for continuous mixture ratios requires a functional form that recovers ordinary monolingual scaling when transfer vanishes and that improves prediction when it does not. No such form, limiting-case argument, or recovery check is stated. This is load-bearing for both the prediction and the optimization claims.","section":"Abstract"},{"comment":"Abstract: exact Shapley values require 2^n characteristic-function evaluations. For any realistic number of languages this is intractable, so an approximation scheme (and its error analysis) is indispensable. The abstract does not mention any approximation, sampling method, or complexity bound; without one the method is not practically usable and the reported experiments cannot be interpreted.","section":"Abstract"},{"comment":"Abstract: the experimental claims of outperformance on “model performance prediction and language mixture optimization” are unsupported by any baseline definitions, metrics, error bars, data splits, ablations of the game-theoretic terms, or statistical tests. With only the abstract available these claims are not checkable and therefore cannot underwrite acceptance.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract is dense and introduces several technical notions (cooperative game, characteristic function, Shapley contribution, mixture-ratio scaling law) without even brief parenthetical definitions; a short clarifying sentence for each would improve accessibility.","section":"Abstract"},{"comment":"The term “ShapleyLaw” is introduced without indicating whether it denotes a specific closed-form expression, a family of laws, or an estimation procedure; consistent terminology would help.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"Only the abstract was supplied for this review; the full manuscript (formal definitions, approximation algorithms, experimental protocol, tables/figures) was unavailable. A proper technical assessment is therefore impossible. I recommend that the editor either obtain the complete paper and re-assign the review, or treat the present report as a provisional abstract-level screening only. The conceptual risks flagged above are real but may well be addressed in the unseen full text; they should not be read as a final rejection recommendation."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"We only have the abstract for 2603.17945, so this is a thin read. The punchline is simple: they cast multilingual pretraining as a cooperative game with languages as players, take Shapley values of test-loss reduction as the measure of each language’s contribution (including transfer), and plug those into a scaling law they call ShapleyLaw. They claim better loss prediction and better mixture ratios than prior multilingual scaling laws that ignore transfer.\n\nWhat is actually new is the framing. Treating mixture optimization as a cooperative game and using Shapley values for cross-lingual contribution is a legitimate move relative to the usual power-law fits that treat languages more independently. If the full paper delivers a clean characteristic function, a workable approximation for the 2^n coalitions, and a scaling form that reduces to the monolingual case when transfer is zero, that would be a useful tool for people who actually set language ratios in pretraining. Credit where due: the problem is real, and the abstract states a clear, falsifiable claim (outperformance on prediction and optimization).\n\nThe soft spots are exactly what you expect from abstract-only. We have no equations, no definition of v(S), no approximation scheme, no baselines, no error bars, no ablations. The stress-test concern is fair: it is not obvious that average marginal contributions over discrete coalitions are the right continuous scalars for mixture ratios, or that v(S) is independent of the complement and of training dynamics. That mapping needs formal justification and consistency checks; without them the outperformance claim is just an assertion. Circularity risk is moderate rather than fatal—Shapley values will depend on how you estimate coalition losses—but we cannot score it from the abstract.\n\nThis is for people who care about multilingual data mixture and scaling laws. A serious referee should see the full paper if the methods section is real; I would not desk-reject on the abstract alone. I would not cite it yet and I would not bring it to reading group until we have the math and the numbers. Send it to review if the full text exists and is coherent; otherwise it stays a title and a promise.","headline":"Abstract-only: plausible game-theoretic framing for multilingual mixture ratios, but no equations or evidence to check the claim.","tokens_in":2791,"tokens_out":526,"would_cite":false,"duration_ms":5096,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Multilingual pretraining is a cooperative game among languages, and Shapley values of each language's contribution yield a scaling law that better predicts test loss and optimizes mixture ratios.","keywords":["multilingual pretraining","scaling laws","language mixture ratios","cross-lingual transfer","Shapley values","cooperative game theory","test-loss prediction"],"falsifier":"Train models on a fixed set of languages under several mixture ratios, measure actual test losses, then check whether the losses predicted by ShapleyLaw (using Shapley values estimated from a smaller set of coalitions) match the observed losses more closely than baseline multilingual scaling laws; a clear failure of that match would falsify the claim.","tokens_in":2823,"feed_emoji":"⚖️","tokens_out":604,"duration_ms":8218,"temperature":0.7,"pith_summary":"The paper argues that existing multilingual scaling laws fail because they ignore cross-lingual transfer: how training data in one language improves performance on others. By treating languages as players in a cooperative game whose payoff is the reduction in test loss from any coalition of languages, the authors compute each language's Shapley value—the average marginal contribution of that language across all possible mixtures. Those values become the coefficients of a new scaling law, ShapleyLaw, that predicts test loss under any mixture and therefore can be optimized to choose better training ratios. Experiments reported in the abstract show that this game-theoretic law outperforms prior multilingual scaling laws both at predicting held-out loss and at finding mixture ratios that improve final model quality.","feed_headline":"Shapley values fix multilingual scaling by capturing language transfer","feed_subtitle":"Treating languages as players in a cooperative game yields better test-loss predictions and mixture ratios.","key_machinery":"Shapley values of languages in a cooperative game whose payoff is test-loss reduction: each language's average marginal contribution across all coalitions becomes the weight that captures both its direct utility and its transfer to other languages, and those weights enter the scaling formula used for prediction and mixture optimization.","core_discovery":"Multilingual pretraining can be cast as a cooperative game in which languages are players and the characteristic function is the reduction in test loss achieved by any coalition of languages; the resulting Shapley values quantify each language's true contribution, including cross-lingual transfer, and plug directly into a scaling law (ShapleyLaw) that predicts test loss more accurately and yields better language-mixture ratios than previous methods.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Shapley values cast languages as players to fix multilingual scaling","Cooperative game theory quantifies true language transfer for better laws","ShapleyLaw predicts test loss via each language's coalition contribution","Languages as game players yield superior multilingual mixture ratios","Cross-lingual transfer measured by Shapley values upgrades scaling laws"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That multilingual pretraining is well-modeled as a cooperative game whose characteristic function is simply the reduction in test loss over coalitions of languages, so that the resulting Shapley values are the right quantities to insert into a scaling law.","fun_headline_variants_meta":{"raw":{"variants":["Shapley values cast languages as players to fix multilingual scaling","Cooperative game theory quantifies true language transfer for better laws","ShapleyLaw predicts test loss via each language's coalition contribution","Languages as game players yield superior multilingual mixture ratios","Cross-lingual transfer measured by Shapley values upgrades scaling laws"]},"model":"grok-4.5","effort":"low","cost_usd":0.005302,"raw_usage":{"total_tokens":1429,"prompt_tokens":722,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":53020000,"prompt_tokens_details":{"text_tokens":722,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":640,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":722,"tokens_out":67,"duration_ms":5852,"temperature":1.0,"reasoning_tokens":640,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T22:48:50.797714+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train models on a fixed set of languages under several mixture ratios, measure actual test losses, then check whether the losses predicted by ShapleyLaw (using Shapley values estimated from a smaller set of coalitions) match the observed losses more closely than baseline multilingual scaling laws; a clear failure of that match would falsify the claim.","supporting_citations":[],"review_version":1}