{"id":"7067ed21-a2d3-4234-a733-b4e657da6a11","arxiv_id":"2501.05629","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Across 204 languages, model scale has little effect on zero-shot multilingual performance, improves two-shot classification linearly, and helps translation mainly for instruction-tuned models.","lead":"This paper measures how larger multilingual language models perform on classification and translation across 204 languages, separating languages seen during training from unseen ones. It finds that scale barely helps in zero-shot use, but helps clearly in two-shot classification, and that only an instruction-tuned model scales well for translation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"BLOOMZ translation scaling is confounded because part of the FLORES-200 evaluation data were used in instruction-tuning, yet the abstract credits only this model with clear scaling benefits.","rationale":"The reader’s weakest assumption was the seen/unseen language mapping. That is a legitimate concern, but the more load-bearing issue is the acknowledged overlap between BLOOMZ’s instruction-tuning data and the FLORES-200 evaluation set. The abstract’s statement about translation scaling depends entirely on BLOOMZ, and the paper flags the overlap but fails to control for it. This is an explicit, self-acknowledged confound, not a speculative mapping error. Even if the seen/unseen mapping is perfect, the translation scaling claim for BLOOMZ remains suspect. I still recommend CONDITIONAL acceptance because the classification findings and the resource-level correlation analyses are not directly tainted and the overlap might turn out to be with the dev set rather than the test subset. But the condition must require the authors to demonstrate disjointness or re-run the translation evaluation. Therefore the verdict stays CONDITIONAL, matching the reader’s verdict, even though the specific concern I identify differs from the reader’s weakest assumption.","tokens_in":13594,"tokens_out":8712,"duration_ms":79204,"concrete_test":"Check the exact 204-sentence test subset per language used in the FLORES-200 evaluation against the xP3 instruction-tuning dataset (including all task templates and language pairs). If any test sentence appears in xP3 (or a near-duplicate), recompute BLOOMZ’s translation scaling on a verifiably disjoint held-out subset and compare slopes. If the scaling slope vanishes or drops below the other models, the abstract’s claim about the instruction-tuned model is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract—“For translation tasks, however, only the instruction-tuned model showed clear benefits from scaling”—rests on the BLOOMZ translation results. The paper itself states in the Text generation section: “We do not show the results of bloomz in this plot and instead plot them separately later because part of the Flores-200 dataset was used to instruction-tune bloomz.” This is a direct acknowledgment that BLOOMZ’s instruction-tuning data overlap with the FLORES-200 dataset used for evaluation. If any of the 204 test sentences per language used for scoring BLOOMZ appear in the xP3 instruction-tuning mix, the observed “clear” scaling in translation is confounded by memorization: larger models memorize more examples, producing artificially steeper scaling. The paper never quantifies the overlap or demonstrates that the test subset is disjoint from instruction-tuning data. Without that guarantee, the translation pillar of the central claim is unsupported. The seen/unseen mapping issue identified by the reader is real but secondary; this contamination directly invalidates one of the three headline findings.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies scaling behavior of three multilingual model families (XGLM, BLOOM, BLOOMZ) across 204 languages on two tasks: topic classification with SIB-200 and machine translation with FLORES-200. Models are evaluated at multiple sizes under zero-shot and two-shot prompting, and languages are categorized as seen or unseen based on the models' reported pretraining language lists. The paper reports scaling slopes, resource-level correlations, and claims that scaling effects depend strongly on task and setting: zero-shot performance is mostly flat, two-shot classification improves linearly with size, and among translation models only the instruction-tuned BLOOMZ shows clear scaling benefits. The resource-level analysis claims that general resource levels predict performance better than language-specific pretraining data proportions.","tokens_in":13823,"tokens_out":3395,"duration_ms":36848,"significance":"If the findings are robust, the paper would provide a useful large-scale empirical map of how model size, pretraining visibility, and prompting interact across more than 200 languages, extending prior scaling studies that cover far fewer languages. The use of publicly listed pretraining language distributions, open model families, and two standardized multilingual benchmarks is a strength, as is the breadth of the evaluation (over two million scored instances). However, the central claims currently rest on several load-bearing assumptions and statistical procedures that are not yet supported: a potential train/evaluation overlap for BLOOMZ on FLORES-200, a fragile seen/unseen mapping, and slope and correlation estimates computed from very few model sizes without uncertainty quantification. These issues materially affect the headline conclusions rather than being presentation concerns.","major_comments":[{"comment":"The abstract's translation claim ('only the instruction-tuned model showed clear benefits from scaling') rests on the BLOOMZ results, but the paper itself states in the Text generation section that part of the FLORES-200 dataset was used to instruction-tune BLOOMZ. Because the translation evaluation uses FLORES-200, the observed scaling slope could be inflated by memorization of instruction-tuning examples if any of the 204 test sentences per language overlap with xP3 training data. The manuscript never quantifies this overlap nor reports results on a verified disjoint subset. Please either demonstrate that the evaluation instances are disjoint from the instruction-tuning data, or re-estimate the BLOOMZ translation scaling on a clean subset and state whether the 'clear benefits' conclusion survives.","section":"Seen and Unseen Languages"},{"comment":"The seen/unseen dichotomy is load-bearing for most comparisons, but it depends entirely on the accuracy of the model-card pretraining language lists and on an ISO 639-3 to ISO 639-1 mapping that resolves multiple scripts using 'the most common script type.' The paper itself qualifies these languages as 'potentially unseen.' Any miscategorization would change every seen-versus-unseen comparison, including the main scaling disparities. Please provide a per-language mapping table, justify the script-resolution decisions, and report a sensitivity analysis that excludes ambiguous languages or uses stricter criteria for 'unseen' status.","section":"Table 3; Figure 1"},{"comment":"The central scaling statements are based on linear fits to only four or five model sizes per family and setting, yet Table 3 reports slopes without confidence intervals, standard errors, or significance tests. In particular, the claims that zero-shot performance is 'mostly flat' and that two-shot classification shows 'clear linear improvements' require interval estimates to distinguish genuine trends from noise around a small number of points. Please add uncertainty quantification (for example, bootstrap intervals over languages or model sizes) and, where possible, permutation or correlation tests to support the qualitative claims.","section":"Appendix; Table 4"},{"comment":"The abstract's claim that 'overall resource levels, not just the proportions of pretraining languages, are better predictors of model performance' is supported only by Pearson correlations in Table 4, with no confidence intervals, p-values, or comparisons between dependent correlations. The table also aggregates over a small number of seen languages (30 for XGLM, roughly 45 for BLOOM/BLOOMZ), and the manuscript notes that for unseen languages the correlations are poor. Please report significance tests and interval estimates for the PD, RL, and RL* correlations, and frame the 'better predictor' claim according to what those tests support.","section":"Prompts, Zero-shot, and Few-shot In-context Learning"},{"comment":"Each language contributes only 204 test sentences, and few-shot demonstrations are drawn from a single random sample. Per-language F1 and especially per-language SacreBLEU scores computed from 204 sentences are noisy, yet the analysis uses them as fixed points when fitting slopes and resource-level correlations. Please assess stability with respect to the demonstration sample (for example, by repeating with multiple random seeds) or with per-language bootstrap intervals, and state how the conclusions change if noisier low-resource languages are downweighted.","section":null}],"minor_comments":[{"comment":"The sentence 'In Section we saw that few-shot ICL hurts smaller models' has an incomplete section reference; please supply the actual section number or name.","section":"Conclusion"},{"comment":"The caption renders '& - seen; & - unseen' instead of a readable legend convention; please define the line styles explicitly in the caption text.","section":"Figure 4 caption"},{"comment":"The text says 'Similar correlation were obtained for 2-shot' but no corresponding table or numbers are provided; either include the 2-shot correlations or remove the unsupported statement.","section":"Appendix; Table 4"},{"comment":"Table 3 omits BLOOMZ rows for text generation because BLOOMZ is plotted separately, but the abstract's translation claim depends on that excluded model; consider adding the BLOOMZ generation slopes to the table with an explicit contamination caveat.","section":"Text generation using bloomz; Table 3"},{"comment":"The prompt templates used for classification and translation are referenced only by prior work; including the exact prompt strings in an appendix would improve reproducibility.","section":"Prompts, Zero-shot, and Few-shot In-context Learning"}],"recommendation":"major_revision","confidential_remarks":"The BLOOMZ/FLORES-200 overlap is acknowledged in the manuscript itself but never resolved; this is the main risk to the abstract's translation claim. If the authors can establish disjointness or redo the analysis on a clean subset, the paper is potentially publishable. The seen/unseen mapping and lack of uncertainty quantification are also substantial, but they appear fixable within the current scope. Given the journal context, I would not recommend reject unless the contamination cannot be quantified and the authors insist on the current abstract wording."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: this is a genuinely broad empirical map of how three multilingual model families scale across 204 languages, but the translation leg of the central claim is partly compromised by training/evaluation overlap, and the abstract overstates the resource-level finding.\n\nWhat is actually new and useful: it is the first evaluation I know of that spans 200+ languages, splits languages into seen/unseen per model, and includes both classification and generation in zero- and two-shot settings. That is a lot of compute and careful dataset assembly. The classification results are interesting: zero-shot scaling is flat, two-shot scaling appears linear for BLOOM/BLOOMZ, and smaller models sometimes do worse with few-shot than zero-shot. The family-level analysis (Indo-European vs Niger-Congo) is a nice touch. The slope table (Table 3) is honest about small slopes in generation.\n\nNow the soft spots. The biggest is BLOOMZ in translation. The paper itself notes that part of FLORES-200 was used to instruction-tune BLOOMZ, yet the abstract credits only BLOOMZ with clear translation scaling. The authors should quantify the overlap or show that the 204 test sentences per language are disjoint from the instruction-tuning data. Without that, the BLOOMZ scaling slope in translation is confounded by a memorization effect that grows with model size. This weakens one of the three headline findings.\n\nSecond, the slopes are fitted from four or five model sizes per family with no confidence intervals or significance tests. With that few points, calling something 'clear linear improvement' is a stretch. The use of a single random demonstration sample per language and only 204 test sentences per language also means high variance. I would like to see standard errors or at least a stability check across different demonstration draws.\n\nThird, the abstract says 'overall resource levels ... are better predictors of model performance,' but the appendix shows that for unseen languages the correlation with resource level is near zero (Table 4). The paper's own conclusion is more careful: resource level helps for seen languages, not unseen. The abstract should match that.\n\nThe seen/unseen mapping is a real dependency, as you note, but the paper is transparent about the 'potentially unseen' qualifier, and the main classification patterns are robust enough that I do not see it as fatal.\n\nWho is this for: anyone working on multilingual model evaluation or low-resource NLP. The paper gives a useful map and some cautionary notes for few-shot prompting. With the contamination issue fixed (or quantified) and statistical grounding improved, it would be a solid contribution. It deserves a serious referee, but I would not accept the current abstract at face value.\n\nMy recommendation: send it to peer review with a request for major revision. In the current form, the translation finding is unsupported; the classification part is interesting.","headline":"Broad 204-language scaling study with a useful classification map, but the translation finding for BLOOMZ is confounded by training/eval overlap and the abstract overstates the resource-level result.","tokens_in":14307,"tokens_out":3091,"would_cite":true,"duration_ms":29328,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multilingual scaling laws are not universal: across 204 languages, model size has little effect on zero-shot text classification, produces linear gains in two-shot classification, and helps translation only for an instruction-tuned model.","keywords":["multilingual language models","model scaling","zero-shot evaluation","few-shot in-context learning","low-resource languages","text classification","machine translation","seen and unseen languages"],"falsifier":"Obtain the actual pretraining corpora for XGLM, BLOOM, and BLOOMZ and check whether the languages labeled 'potentially unseen' actually appear in them; if a substantial fraction do, the flat zero-shot scaling and the seen-unseen gaps are artifacts of the labeling. Recomputing the scaling slopes using only languages with unambiguous ISO and script mappings would give a direct test.","tokens_in":13449,"feed_emoji":"🌐","tokens_out":7731,"duration_ms":66873,"temperature":0.7,"pith_summary":"This paper sets out to answer a practical question: when a multilingual language model is made larger, do languages see a performance improvement, and does it matter whether the language appeared in pretraining? The authors evaluate three model families (XGLM, BLOOM, BLOOMZ) at parameter counts from 560 million to 7.5 billion across 204 languages on topic classification and machine translation, in zero-shot and two-shot settings. They find that scaling does not follow one trend: zero-shot classification stays mostly flat as models grow, two-shot classification shows clear linear gains with scale, and translation improves with scale only for the instruction-tuned BLOOMZ. Seen languages outperform unseen languages throughout, and a language's general resource level predicts performance better than its share of pretraining data. If this is right, developers cannot assume a universal scaling law for multilingual systems; the task, the prompting setup, and pretraining visibility determine whether scale helps.","feed_headline":"Scale barely moves zero-shot multilingual scores","feed_subtitle":"Across 204 languages, gains from model size depend on task, prompting, and whether the language was seen in training.","key_machinery":"The argument is carried by the empirical scaling curve: each model family comes in several sizes trained on the same corpus, so task performance can be plotted against parameter count and summarized by the slope of a linear fit. The comparisons are organized by a seen-versus-unseen split built from pretraining language lists and by six resource levels, applied to both tasks under zero-shot and two-shot prompting. The slopes in Table 3 — near zero for zero-shot classification, positive for two-shot classification, and near zero for most translation settings — are the evidence for the central claim.","core_discovery":"The paper's central claim is that multilingual scaling behavior is task- and setting-dependent rather than universal. In text classification on SIB-200, macro-F1 is largely flat from 560M to 7.5B parameters under zero-shot prompting; under two-shot prompting, larger models show clear linear improvement for both seen and unseen languages, with XGLM's unseen-language performance as the one slight exception. In xx-to-English translation on FLORES-200, XGLM and BLOOM show little or no scaling, whereas the instruction-tuned BLOOMZ shows clear improvements from scale under both settings, although two-shot demonstrations hurt it. Seen languages consistently beat potentially unseen languages, including within the same language family, and larger models narrow the gap for seen languages but not for unseen ones. The paper further claims that overall resource level, rather than the language's proportion in pretraining data, is the stronger predictor of performance, and that this is true for seen languages but not for unseen ones.","pith_inferences":["Inference: if overall resource level is the stronger predictor, then adding pretraining data for a single low-resource language may not move its score much; improving related-language coverage or general data diversity could matter more.","Inference: the k=2 ceiling may itself shape the result; testing larger k on models with longer context windows could reveal whether translation scaling emerges later.","Inference: the contrast between flat zero-shot and linear two-shot classification suggests the scaling signal lies in in-context learning capacity rather than static multilingual knowledge; comparing English-prompt versus target-language-prompt results would test this.","Inference: because BLOOMZ translation degrades when given two examples, instruction-tuned models may be sensitive to prompt format; varying example ordering or templates could restore the benefit."],"forward_implications":["Zero-shot multilingual classification benchmarks will show little separation across model sizes from 560M to 7.5B; users should expect flat scores without demonstrations.","Two-shot demonstrations are what unlock scaling benefits in classification, so few-shot evaluation is necessary to observe improvements from larger models.","For translation, increasing model size alone will not reliably improve low-resource quality; instruction tuning is the observed exception.","Resource-level groupings, not pretraining corpus proportions, should guide expectations for a language's performance."],"supporting_citations":[{"why":"Supplies the SIB-200 topic-classification benchmark covering 204 languages and dialects.","marker":"Adelani et al. 2024"},{"why":"Supplies the FLORES-200 parallel dataset used for xx-to-English machine translation evaluation.","marker":"NLLB Team et al. 2022"},{"why":"Describes the XGLM model family and the 30 languages in its pretraining corpus.","marker":"Lin et al. 2022"},{"why":"Describes BLOOM and the ROOTS pretraining corpus with 46 natural languages.","marker":"Scao et al. 2023"},{"why":"Describes the BLOOMZ instruction-tuned variant and the xP3 tuning data.","marker":"Muennighoff et al. 2023"},{"why":"Provides the six-level language resource categorization that structures the resource-level analyses.","marker":"Joshi et al. 2020"},{"why":"Defines SacreBLEU, the metric used to score translation outputs.","marker":"Post 2018"},{"why":"Establishes the few-shot in-context learning paradigm behind the two-shot setting.","marker":"Brown et al. 2020a"}],"fun_headline_variants":["Zero-shot multilingual: scale is no help, few-shot is","Model scale flattens zero-shot, lifts two-shot language scores","Scaling helps two-shot text classification, not zero-shot","Translation scaling only for instruction-tuned models","Seen languages outperform unseen; scale narrows gap only for seen"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the pretraining language lists used to label each language as seen or unseen are accurate, and that the ISO-code and script mapping that connects dataset languages to those lists is correct; if a language is miscategorized, every seen-versus-unseen comparison and the scaling conclusions built on it become unreliable.","fun_headline_variants_meta":{"raw":{"variants":["Zero-shot multilingual: scale is no help, few-shot is","Model scale flattens zero-shot, lifts two-shot language scores","Scaling helps two-shot text classification, not zero-shot","Translation scaling only for instruction-tuned models","Seen languages outperform unseen; scale narrows gap only for seen"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001063,"raw_usage":{"total_tokens":4443,"prompt_tokens":915,"completion_tokens":3528,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":3445}},"tokens_in":531,"tokens_out":3528,"duration_ms":22877,"temperature":1.0,"reasoning_tokens":3445,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:12:28.767831+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Obtain the actual pretraining corpora for XGLM, BLOOM, and BLOOMZ and check whether the languages labeled 'potentially unseen' actually appear in them; if a substantial fraction do, the flat zero-shot scaling and the seen-unseen gaps are artifacts of the labeling. Recomputing the scaling slopes using only languages with unambiguous ISO and script mappings would give a direct test.","supporting_citations":[],"review_version":1}