{"id":"1466d483-26f4-49a4-af98-db1562f57f40","arxiv_id":"2502.08489","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.","lead":"Researchers at Barcelona Supercomputing Center trained and released Salamandra, a family of open-source language models from 2 to 40 billion parameters covering 35 European languages and code. The models and training and evaluation code are public under an Apache 2.0 license, offering a transparent multilingual alternative to closed or English-centric models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prometheus-2 judge scores for Catalan, Basque, and Galician lack human validation; instructed-model competitiveness is not yet supported.","rationale":"The paper is a transparent technical report accompanying released models. It discloses training data sources, hyperparameters, evaluation scripts, and model weights, which are strong independent contributions that can be tested by the community. The central claim, however, is that Salamandra achieves competitive performance compared to similarly sized open-source models. That claim is necessarily comparative and therefore depends critically on the validity of the evaluation instruments. I agree with the reader's general concern that the evaluation design may not be a valid proxy for real multilingual ability. My stress-test narrows the concern to the LLM-as-a-Judge setup, which is the most load-bearing component because it has no ground truth at all: Prometheus-2 was fine-tuned on English evaluation data only, and the paper's own remarks in Section 5.1 reveal that its behavior on non-English inputs required manual prompt iteration. For low-resource Iberian languages such as Basque and Galician, a mis-calibrated judge could reverse relative rankings, making the comparative claim unsupported. The base-model results in Table 9 use established datasets with gold labels, which is less concerning, though IberoBench's provenance still deserves scrutiny. The reader's CONDITIONAL verdict is the right disposition: the engineering and openness contributions are real, but the competitive performance claim requires independent human or at least judge-calibration validation. My concrete test would settle whether the judge is reliable for the languages where Salamandra's advantage is claimed. Since the reader already asked for exactly this kind of validation, no change to the verdict is needed.","tokens_in":64715,"tokens_out":5160,"duration_ms":54450,"concrete_test":"Sample 200 responses per language (Spanish, Catalan, Basque, Galician) from Salamandra 7B Instructed and two top baselines (e.g., Gemma-2 9B and Mistral-7B-v0.3) from Tables 11-14. Have three native speakers per language score each response using the same rubrics as the judge. Compute the Spearman correlation between human mean scores and Prometheus-2 scores for each language, and compare the model rankings under human scores versus judge scores. If any language shows Spearman rho < 0.7 or a change in which model is ranked first, the LLM-as-a-Judge tables do not support the instructed-model competitiveness claim and the report should be revised to withhold comparative conclusions pending human evaluation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim of competitive performance for Salamandra, especially for Iberian languages, rests on evaluation results in Tables 9-18. The instructed-model comparisons in Tables 11-18 rely on Prometheus-2 8x7B as an LLM judge (Section 5.2.2). Prometheus-2 is a Mixtral-8x7B model fine-tuned on English-only human-annotated evaluation data; the report uses an English system prompt and English rubrics while the queries and responses are in Spanish, Catalan, Basque, Galician, French, German, and Italian. The paper itself acknowledges in Section 5.1 that Prometheus-2 'reacted' to non-English languages and required prompt iteration, but no human validation of the judge's scores is reported. Section 5.1 also states that a comprehensive human evaluation setup is 'on our current roadmap,' which is an explicit admission that the comparative quality claims for instructed models currently lack ground-truth support. If Prometheus-2 is biased toward Mixtral-like outputs or systematically misjudges fluency and correctness in low-resource Iberian languages, the model rankings in Tables 11-18 could change, undermining the central competitive claim. This is not a claim of dishonesty; it is a missing validation step that is load-bearing because the models are positioned as open alternatives for European languages, particularly Catalan, Basque, and Galician.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This technical report introduces Salamandra, a family of from-scratch decoder-only large language models of 2B, 7B, and 40B parameters, trained on a corpus covering 35 European languages and code, with instruction-tuned checkpoints, a vision proof-of-concept, and open releases under the Apache 2.0 license. The paper documents architecture, tokenizer design, data curation and processing pipelines, pre-training and continued-training recipes, post-training choices, and a large evaluation campaign: base and instructed model comparisons in Tables 9–18 using the LM Evaluation Harness, IberoBench and other benchmarks, an LLM-as-a-Judge setup with Prometheus-2 8x7B, and bias, cognitive-bias, and red-teaming analyses. The central claim is that Salamandra achieves competitive performance against similarly sized open-source models, especially in Iberian languages. The report is transparent about limitations, including an unfinished 40B checkpoint, the absence of preference alignment, and the lack of a comprehensive human evaluation setup.","tokens_in":64928,"tokens_out":6021,"duration_ms":63851,"significance":"If the evaluation results hold, this is a useful contribution to the European-language open-model landscape: the suite provides open weights at three sizes, public training and evaluation scripts, detailed data curation documentation, a model card, a datasheet, and an unusually honest discussion of weaknesses. The tokenizer fertility analysis and the open red-teaming pipeline are also valuable artifacts. However, the headline competitiveness claim is not yet fully established. The instructed-model comparisons in Tables 11–18 rely on an LLM judge that is not validated against human judgments for the non-English languages, and several headline Iberian-language benchmarks are developed or extended by the same team. With added human validation, uncertainty quantification, and clearer separation of preliminary 40B results, this could become a solid reference for multilingual open-weight models.","major_comments":[{"comment":"The instructed-model evaluation uses Prometheus-2 8x7B as judge with an English system prompt and English rubrics while the queries and responses are in Spanish, Catalan, Basque, Galician, French, German, and Italian. The paper reports in §5.1 that Prometheus-2 'reacted' to non-English languages and required prompt iteration, and that a comprehensive human evaluation setup is 'on our current roadmap.' No human agreement or calibration check for the judge is reported. Because the instructed-model competitiveness claim rests on these scores, please add a human validation sample (e.g., 100–200 instances per language and task) with judge–human agreement metrics, or remove or substantially qualify the comparative instructed-model claims.","section":"§5.1, §5.2.2, Tables 11–18"},{"comment":"The base-model claim of 'strong capabilities' and 'competitive performance' is based on single point estimates without confidence intervals, standard errors, or significance tests. Several comparisons are within a few points (e.g., xstorycloze en: Salamandra 7B 79.09 vs. EuroLLM 9B 80.41), and §5.2.1 notes that library versions, tensor parallelism, and the use of vLLM can shift scores by 1–2% and in some cases more. The authors should either provide bootstrap intervals or other uncertainty estimates, state a minimum meaningful difference, and refrain from ranking models on gaps smaller than the reported variability.","section":"§5.1, Table 9"},{"comment":"IberoBench and EsBBQ are developed or extended by the same research group and are used for several headline Catalan, Basque, and Galician results. This is not invalid by itself, but it creates a correctness risk for the central comparative claim because these instruments have not been independently audited. Please (i) identify which headline conclusions depend exclusively on these in-house benchmarks, (ii) report contamination checks between these benchmarks and the pre-training or instruction-tuning data (the §3.6.1 Aya filtering indicates that evaluation-benchmark overlap is already a considered concern), and (iii) add at least one external or third-party benchmark per headline language.","section":"§5.2.1, §6.2.1"},{"comment":"The 40B results in Table 9 come from an intermediate checkpoint whose training is still ongoing and that has not undergone the annealing phase, as stated in §3.4. Presenting these results among the final released checkpoints, and referring in the abstract to 'three different sizes,' can mislead readers about what is actually shipped. Please move the 40B results to a clearly separated preliminary section or table, and state in the abstract that the 40B release is an incomplete, non-final checkpoint.","section":"§3.4, §5.2.3, Table 9"}],"minor_comments":[{"comment":"The row labelled 'Hungarian hr Balto-Slavic' should be labelled 'Croatian hr Balto-Slavic'; Hungarian is 'hu' and is listed separately in the same table and in Figure 4.","section":"Table 2"},{"comment":"The sentence 'A script for re-setting reserved tokens is provided is provided' contains a duplicated phrase and should be corrected.","section":"§2.3.2"},{"comment":"The text 'LM Evalaution Harness' should read 'LM Evaluation Harness.'","section":"§5.2.2"},{"comment":"There are unexplained 'nan' entries in the table (e.g., xquad_en for Occiglot-eu5 7B and Teuken 7B); please state how missing values were handled and whether they affect the comparisons.","section":"Table 9"},{"comment":"There are small typos: 'Aya 23 8B us generally more resistant' should read 'is generally more resistant,' and 'One significant issues' should be 'One significant issue.'","section":"§6.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an honest and unusually transparent technical report, and the main missing validation step is fixable within the scope of a revision. I do not see evidence of misconduct, but the abstract and the instructed-model tables overstate the support for the competitive claim until human validation of the judge is provided. If this is intended for a peer-reviewed venue rather than a report-style repository, the authors should consider reframing the contribution around the open release, the data/recipe disclosure, and the evaluation methodology, rather than the headline ranking."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look before you invest time: this is a technical report for Salamandra, a from-scratch 2B/7B/40B decoder-only family trained on 35 European languages plus code, released under Apache-2.0 with training and evaluation scripts public. That alone is a real contribution. The base-model benchmark tables are extensive and the tokenizer fertility study is genuinely useful; the data curation section is among the most detailed you will see in this kind of report. They also disclose what they did not do: no preference alignment, an unfinished 40B checkpoint, and math performance that is frankly weak. The soft spot is exactly where the stress-test note lands. The instructed-model comparisons in Tables 11-18 rely on Prometheus-2 8x7B as a judge, a model fine-tuned on English-only human evaluation data, applied to Catalan, Basque, Galician, and other languages with English prompts and rubrics. The paper admits the judge 'reacted' to non-English languages and that human evaluation is 'on our current roadmap.' That is a load-bearing gap for the claim of competitive instructed-model quality, not a minor footnote. Without human validation or an independent audit, the rankings in those tables could shift. The base-model claims rest on more solid ground because they use external datasets like MGSM, Belebele, and FLORES, though even there the report gives point estimates without confidence intervals. I also note that IberoBench and EsBBQ are developed or extended by the same group, so the headline Iberian-language results are partly self-referential. That is not fatal, but it is worth keeping in mind when reading 'strong capabilities.' The paper is unusually honest about its limitations, which makes me trust the parts that look good. Bottom line: this deserves a serious referee. It is a solid engineering and data documentation contribution, and the open release lets anyone reproduce the core results independently. But the instructed-model comparative claims need either human evaluation or a clearly validated multilingual judge before they should be taken as established. I would accept it for peer review with a request for revision on that point, and I would not want the benchmark tables published without error bars or a discussion of the judge's language bias. If you work on multilingual models or evaluation, this is worth your time; cite it for the tokenizer study and the data pipeline, not for the instructed-model rankings.","headline":"A genuinely open multilingual model family with a strong engineering report, but the headline competitive claim for instructed models hangs on an unvalidated LLM judge and needs referee attention.","tokens_in":744,"tokens_out":684,"would_cite":true,"duration_ms":23974,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This report claims Salamandra, a from-scratch family of 2B, 7B, and 40B open models trained on 35 European languages plus code, reaches competitive performance against similarly sized open-source models.","keywords":["open-source language models","multilingual European languages","Spanish","Catalan","Basque","Galician","instruction tuning","bias and safety evaluation"],"falsifier":"Give independent human annotators the same prompts used in the judge-based evaluation (story completion, math, paraphrase, reading comprehension, summarization, and translation) for Salamandra and the comparison models in Catalan, Basque, Galician, and Spanish, with model identity hidden, and see whether human preference rankings match the reported score ordering; a clear reversal would falsify the claim.","tokens_in":64482,"feed_emoji":"🦎","tokens_out":6129,"duration_ms":64312,"temperature":0.7,"pith_summary":"This paper introduces Salamandra, a suite of decoder-only language models in three sizes (2B, 7B, and 40B), trained from scratch on open-access text from 35 European languages plus code. The central claim is that these models achieve competitive performance against similarly sized open-source models on multilingual benchmarks, especially in Spanish, Catalan, Basque, and Galician. If the claim holds, researchers and companies get a fully open, Apache-2.0 alternative for European languages where strong open models are scarce. The paper additionally makes its training recipes, data curation methods, and evaluation scripts publicly available, so the suite is reproducible in spirit.","feed_headline":"Open-model family hits competitive scores in 35 European languages","feed_subtitle":"Apache-2.0 2B, 7B, and 40B checkpoints target Spanish, Catalan, Basque, and Galician; recipes and scripts are public.","key_machinery":"The load-bearing machinery is the training and evaluation stack: a 256,000-token byte-pair-encoding tokenizer trained on a per-language balanced sample; multilayered data curation with language detection, deduplication, and quality scoring; factor and epoch sampling that emphasizes Spanish, Catalan, Galician, and Basque while undersampling English and code; a two-stage pre-training schedule with a continued high-quality-data phase; supervised instruction tuning on multilingual chat data with added embedding noise; and an evaluation pipeline built on an Iberian-language benchmark plus a separate judge model for open-ended tasks. The evaluation pipeline carries the competitive-performance claim, while the data pipelines are what make the models genuinely multilingual.","core_discovery":"The authors claim that Salamandra models, trained entirely on open-access data, reach performance competitive with comparable open baselines across a broad multilingual evaluation. On the reported benchmarks, the 7B and 40B variants often lead or tie in Catalan, Basque, and translation tasks, while mathematical reasoning and natural language inference remain relatively weak. The instruction-tuned checkpoints improve instruction following and several downstream categories, but they hurt translation and generative truthfulness. The 40B model is described as an intermediate checkpoint whose training has not finished and has not undergone the annealing phase. No released checkpoint has received preference-based alignment, and the paper states such alignment is future work.","pith_inferences":["Because several evaluation instruments were built or translated by the same team and the judge-based scores were not checked against human raters, an independent multilingual human evaluation is the natural next test before betting on the competitive claim.","If the balanced tokenizer training transfers as the fertility numbers suggest, similar uniform-language tokenizer recipes could help other low-resource languages, though the report does not directly measure downstream transfer.","The vision experiments are explicitly a proof of concept; treating them as production multimodal capability would go beyond what the paper claims.","Given the absence of preference alignment, safety-sensitive applications should treat the released checkpoints as base material for further tuning, not as final products."],"forward_implications":["If the reported capabilities hold, the Apache-2.0 release gives researchers and companies deployable open models for Spanish, Catalan, Basque, Galician, and many other European languages without licensing barriers.","The published training recipes, data-selection logic, and evaluation scripts make the suite reproducible in spirit, so others can audit or extend the approach.","The instructed checkpoints should be used knowing they improve instruction following but lose ground on translation and truthfulness; task choice matters.","The 40B checkpoint, though unfinished, already leads the family and several comparisons, so completing its training and adding an annealing phase should improve it further.","The safety analysis shows that larger and instructed models lean more on social stereotypes and that red-teaming resistance varies by language, so downstream deployments need their own alignment and safety testing."],"supporting_citations":[{"why":"Supplies the Iberian-language benchmark tasks that anchor the main multilingual performance claim.","marker":"[19]"},{"why":"Provides the evaluation harness used to run the gold-standard automatic benchmarks.","marker":"[62]"},{"why":"Serves as the judge model for LLM-as-a-Judge scoring of open-ended tasks.","marker":"[88]"},{"why":"Defines a major open-source baseline family and the staged/annealing training recipe the report follows.","marker":"[53]"},{"why":"Provides a widely used 7B open baseline against which Salamandra 7B is compared.","marker":"[78]"},{"why":"Provides the Gemma-2 family as a strong comparable-size baseline throughout the evaluation.","marker":"[190]"},{"why":"Provides EuroLLM, a directly comparable multilingual European model used as a baseline.","marker":"[123]"},{"why":"Supplies FLOR, a smaller language-adaptation baseline, and informs the language-adaptation discussion.","marker":"[45]"},{"why":"Supplies the FineWeb-Edu high-quality subset used in the continued pre-training phase.","marker":"[148]"},{"why":"Supplies StarCoder training data that contributes code and is linked to improved reasoning.","marker":"[105]"}],"fun_headline_variants":["Salamandra LLMs: 2B to 40B, open and multilingual","Open-source LLM family hits competitive multilingual scores","35-language LLMs trained from scratch, fully open","Apache-2.0 LLMs: competitive multilingual performance","Open LLMs: public data, training code, and recipes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion stands only if the automatic benchmarks and the judge model actually measure real capability in Catalan, Basque, Galician, and Spanish; if those instruments are biased, the competitive result is not established.","fun_headline_variants_meta":{"raw":{"variants":["Salamandra LLMs: 2B to 40B, open and multilingual","Open-source LLM family hits competitive multilingual scores","35-language LLMs trained from scratch, fully open","Apache-2.0 LLMs: competitive multilingual performance","Open LLMs: public data, training code, and recipes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001965,"raw_usage":{"total_tokens":7661,"prompt_tokens":910,"completion_tokens":6751,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":6666}},"tokens_in":526,"tokens_out":6751,"duration_ms":53459,"temperature":1.0,"reasoning_tokens":6666,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T04:53:27.955613+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give independent human annotators the same prompts used in the judge-based evaluation (story completion, math, paraphrase, reading comprehension, summarization, and translation) for Salamandra and the comparison models in Catalan, Basque, Galician, and Spanish, with model identity hidden, and see whether human preference rankings match the reported score ordering; a clear reversal would falsify the claim.","supporting_citations":[],"review_version":1}