{"id":"b3538c9f-c278-4c74-9159-f897d7498cf3","arxiv_id":"2607.09424","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Soofi S 30B-A3B, a hybrid Mamba-MoE model pretrained on ~27T tokens with deliberately up-weighted German, reports the highest English and German aggregate scores among fully open base models in its comparison while matching dense 14-27B models at ~3B active parameters.","lead":"A consortium of German research institutes trained Soofi S 30B-A3B, an open 30-billion-parameter model that activates only ~3B parameters per token and reports top English and German benchmark scores among fully open base models after roughly 27 trillion tokens of training. The report documents the full per-source data mixture, publishes training and evaluation code, and claims an 8-9x long-context decode-throughput advantage over dense 14-27B models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Contamination audit in §4.3 covers only the QA-base slice and only the mislabeled-split failure mode; the web and MT-German sources most likely to hide benchmark items were never screened against the eval suite, so the benchmark-based central claim rests on an unverified premise.","rationale":"I agree with the reader that the completeness of the Section 4.3 contamination audit is the load-bearing premise, and I sharpen the mechanism. The audit has two narrowings: scope (only the QA-base constituents, ~0.05% of the Phase 2 pool) and mechanism (only the mislabeled-split failure mode). The principal documented leakage pathway for web-scale pretraining — benchmark items present in crawl data — is not addressed for this run: the corpus includes 11.6T effective web tokens up-sampled to 3 epochs and a 571B German MT of ClimbMix, while the -DE eval benchmarks are German renderings of English items, so MT-German web text is a direct near-duplicate channel to the German eval sets. The n-gram screening that would detect this is disclosed (Section 4.3, 'Remediation') as a forward-looking practice, and no screening result for the actual run is reported. The paper's own Figure 15 concedes that paraphrased contamination is indistinguishable from genuine gains at the trajectory level, so the ex-post discussion cannot rescue the premise. I considered two alternatives and set them aside. The throughput claim (4.82k TPS/GPU, 8-9x over dense baselines at 40K) is structurally supported by the hybrid Mamba-MoE architecture and the disclosed latency-subtraction protocol; even a factor-of-two measurement error would leave the qualitative claim intact. The aggregate-reproducibility gap (German aggregate 85.3 vs ~76 simple mean of the 11 visible German tasks) is real, but the paper promises release of per-task results, so it is a verification issue rather than evidence of incorrectness, and it is secondary to the contamination premise. The paper's disclosure record — the GPQA incident, the discarded annealing stage, the checkpoint-merge ablation, the Minerva protocol correction — supports reading the report as honest, which is why this is a conditional-accept situation, not a rejection. My concern confirms rather than moves the reader's verdict: release weights/code/data, run an independent contamination screen including the MT-German sources, and recompute the affected aggregates. Hence UNCHANGED.","tokens_in":52458,"tokens_out":18443,"duration_ms":192409,"concrete_test":"Obtain the released training-mixture artifacts (corrected QA-base, the German KletterMix translation of ClimbMix, MultiSynt/MT, and the Nemotron-CC tiers actually used) together with the full item sets of every reported English and German benchmark, including -DE variants. Run a paraphrase-tolerant overlap screen — e.g., ≥8-gram exact matches plus embedding cosine >0.9 on non-boilerplate text — between every eval item and the training sources, with particular focus on the MT-German sources and the up-sampled web tiers. Decision rule: if any reported benchmark has evaluation items with training near-duplicates above background, recompute that benchmark and the 77.3/85.3 aggregates with those items removed; the concern lands if the +5.9 German margin over Apertus or the +1.5 English margin over Olmo 3 disappears or reverses. A clean screen retires the objection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central capability claim is entirely benchmark-based, and its load-bearing premise is Section 4.3's assertion that only four constituents leaked evaluation material into training. That assertion rests on an audit with a narrow scope and a single hypothesized mechanism. Scope: the re-audit covers 'all QA-base constituents' — roughly 3.3B tokens, 0.05% of the Phase 2 pool (Section 3.3) — not the other ~99.95% of the 26.68T-token corpus. Mechanism: the audit looks only for evaluation-only benchmarks whose sole published split is mislabeled 'train'. The dominant leakage pathway for web-scale corpora is different: benchmark items are endemic in Common Crawl-derived text, and Soofi's corpus deliberately up-samples web tiers (Nemotron-CC at up to 3 epochs, 11.6T effective tokens) and includes a 571B German machine translation of ClimbMix (Section 3.5.3). The -DE evaluation benchmarks are themselves German renderings of English items, so MT-German web text containing the underlying English items is a direct near-duplicate channel to the German eval sets — exactly the benchmarks carrying the flagship +5.9 German margin over Apertus. The remedy that would close this pathway, screening the final mixture against the full evaluation suite via n-gram overlap, is listed under 'Remediation' as a forward-looking practice; no screening result for this run is reported. The paper's own Figure 15 concedes paraphrased contamination 'is indistinguishable, at the trajectory level, from genuine capability gains', so the ex-post trajectory analysis cannot carry the premise either. Independent verification is required before the 77.3/85.3 aggregates can be read as capability measurements rather than possible leakage artifacts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Soofi S 30B-A3B, a 31.6B-parameter mixture-of-experts hybrid Mamba-Transformer base model with ~3.2B active parameters, pretrained on roughly 26.68T tokens with deliberately up-weighted German. The authors document a three-phase curriculum (20T diverse pretraining, ~6.6T high-quality annealing, ~0.1T long-context extension), release full per-source token accounting, training hyperparameters, intermediate checkpoints, and evaluation code, and evaluate the model against 15 (elsewhere 16 or 17) open and open-weight baselines. The central claims are that Soofi S is the strongest fully open model in the evaluation on English and German aggregates, matches dense 14–27B international models on aggregate performance at a fraction of the active parameter cost, achieves best-in-comparison code aggregates, and sustains 8–9× the aggregate decode throughput of dense baselines at 40K context. The paper also discloses a benchmark-contamination incident involving GPQA in the QA-base pretraining constituent and describes a remediation that removes GPQA from reported aggregates and adds forward-looking screening practices.","tokens_in":52745,"tokens_out":5126,"duration_ms":59644,"significance":"If the capability claims survive scrutiny, this is a significant contribution: a fully documented, sovereign, German-English pretraining run with unusually complete data accounting, an architecture-identical baseline that cleanly isolates the data recipe, and a credible serving-efficiency advantage from the hybrid Mamba-MoE design. The release of weights, selected checkpoints, exact per-source token counts, hyperparameters, training/evaluation code, and even a discarded final-annealing stage is exemplary for reproducibility. The long-context weakness on common-word extraction is disclosed honestly. However, the headline capability claims are entirely benchmark-based, and the paper's own contamination disclosure establishes that benchmark material entered the training mixture through at least one pathway. The completeness of the contamination audit is therefore load-bearing for the central claims.","major_comments":[{"comment":"The audit's scope is a load-bearing limitation. The re-audit explicitly covers 'all QA-base constituents' — roughly 3.3B tokens, about 0.05% of the Phase 2 pool (Section 3.3) — and only the failure mode of evaluation-only benchmarks whose sole published split is mislabeled 'train'. The remaining ~99.95% of the ~26.68T-token corpus, including the deliberately up-sampled English web tiers (Nemotron-CC, 11.6T effective tokens in Phase 1) and the 571B-token German machine translation of ClimbMix, is not screened against the evaluation suite. This matters doubly for the German benchmarks: the -DE evaluation sets are German renderings of English items, so MT-German web text containing the underlying English items is a direct near-duplicate channel to the German eval sets that carry the flagship +5.9 German aggregate margin. The paper's own Figure 15 concedes that paraphrased contamination is i","section":"Section 4.3 (GPQA Contamination Disclosure)"},{"comment":"The manuscript states that QA-base contains 'paraphrased training splits of 25 standard NLP benchmarks in English and German', and that the model trained on QA-base. The evaluation suite includes benchmarks from the same families (code, math, QA, knowledge). Training on paraphrased train splits of a benchmark family can inflate downstream scores on that family even when no evaluation item is duplicated. The contamination disclosure in §4.3 addresses only evaluation-set leakage of four specific datasets, not this broader train-split exposure. The paper does not list the 25 benchmarks, nor does it analyze which of the reported English or German eval tasks have train-split overlap with QA-base. Because several of the largest reported margins are on code and math tasks (HumanEval +10.8, MBPP-DE +13.4, Minerva +24.2), this is potentially a direct confound for the 'strongest fully open model'","section":"Section 3.3 and Tables 4–5"},{"comment":"The n-gram screening of final training mixtures against the evaluation suite is described as a forward-looking practice ('final training mixtures are screened ... before training'), not as a result for this run. Removing GPQA from the reported aggregates is a necessary correction but does not repair the possibility that other evaluation material entered through unscreened web or MT-German data, nor does it address the train-split exposure identified above. The claims in the Contributions section and Conclusion — 'strongest fully open model', 'matches dense 14–27B models', 'first European sovereign model to sit on the same capability-per-active-parameter frontier' — are therefore stronger than the evidence currently supports. A revision should either supply screening results for this run or substantially qualify these claims, e.g., by stating that they hold 'barring undetected contaminati","section":"Section 4.3, 'Remediation'"}],"minor_comments":[{"comment":"The number of comparison models is inconsistent: the Abstract says 'among 17 open base models', Section 4 says 'against 15 open-source and open-weight base models', and the Conclusion says 'unified evaluation of 16 open base models'. Please reconcile these counts.","section":"Abstract and Conclusion"},{"comment":"Equation (1) subtracts t(1) from t(1024), which removes prefill cost only if the per-token decode time is approximately linear in output length. The paper should state this linearity assumption explicitly, since the 'TTFT-like' t(1) values are reported separately and the aggregation of prefill and decode into a single TPS figure may be sensitive to the chosen output-length range.","section":"Section 4.4, Eq. (1)"},{"comment":"The long-context comparison with Nemotron 3 Nano is a strength, but the RULER CWE collapse beyond 32K is a substantial capability gap that is only visible in Appendix E. Consider foregrounding this limitation in the main text rather than only in an appendix.","section":"Section 2.2 and Appendix E"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's open-data and open-code commitments are genuinely strong, and the architecture-matched baseline with Nemotron 3 Nano is a good design. My main concern is that the central benchmark-based claims rest on a contamination audit that is self-performed, narrow in scope (QA-base only), and demonstrably incomplete with respect to the web and MT-German portions of the corpus. This is not a reason to reject outright — the issue is addressable by additional screening or by explicitly scoping the claims — but it is load-bearing. I would want to see either concrete evidence that the final mixture of this run is free of eval-set overlap beyond QA-base, or a revised set of claims that do not depend on that premise."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know up front. The good news: this is the most candid large-scale pretraining report I've read in a long time. Full per-source accounting for all three phases, proxy ablations, discarded annealing stages, checkpoint-merge results, a Minerva protocol correction, and a contamination incident owned in detail. The catch: the headline capability claims — 'strongest fully open model,' +5.9 German aggregate over Apertus — are benchmark-based, and the load-bearing premise, that the eval suites were clean, is established by an audit that covers about 0.05% of the Phase 2 pool and exactly one failure mechanism.\n\nThe §4.3 re-audit covers the QA-base constituents only, roughly 3.3B tokens, and looks for benchmarks whose sole published split is mislabeled 'train.' It found GPQA, TruthfulQA, BLiMP, and Inverse Scaling. But the dominant contamination pathway for web-scale corpora is different: eval items are endemic in Common Crawl text, and this corpus deliberately up-samples those tiers, including 571B tokens of machine-translated German ClimbMix. The -DE eval benchmarks are German renderings of English items; the German training data is heavily MT of English web. The near-duplicate channel is right there, and it was never screened for this run. The paper's own Figure 15 concedes paraphrased contamination is indistinguishable from genuine gains at the trajectory level — so the training logs can't carry the premise either. N-gram screening of the final mixture is listed under 'Remediation' as a practice for future runs.\n\nNone of this breaks the paper. The throughput section is the strongest part: a clean latency-subtraction protocol on a single B200, TP=1, batch 32, showing flat decode throughput from 4K to 256K and an 8-9x advantage over dense 14-27B models. Contamination doesn't inflate TPS, and the mechanism — six GQA layers in a Mamba-2 backbone — is sound. The architecture-identical comparison with Nemotron 3 Nano is the right experimental design for isolating the data recipe. And the paper reports its own RULER regression (CWE collapse to 3% at 256K–1M where the reference keeps 60-64%); that is hard to square with systematic spin, and it's a real point in the authors' favor.\n\nMinor soft spots: the withheld benchmark group is never named, and headline margins like +0.6 English aggregate over Nemotron sit at noise level with no error bars. The self-performed audit would be fine in principle; the scope is the problem, not the authorship.\n\nWho this is for: anyone building European or German-capable models, and anyone writing about eval contamination. It deserves a serious referee. The claims are checkable — weights, code, and data accounting are promised — so the audit question is an empirical one, not a matter of faith. The right outcome is major revision: extend the audit to the web and MT-German sources against the full eval suite, name the withdrawn benchmarks, add variance estimates. If an independent audit comes back clean, this is an important result.","headline":"An unusually honest, well-disclosed German-English pretraining report whose headline benchmark claims are conditional on a contamination audit that covers only the QA-base slice and never screens the web or MT-German channels most likely to hide eval items.","tokens_in":53530,"tokens_out":8371,"would_cite":true,"duration_ms":87597,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Soofi S, a fully open German–English model activating 3.2 of its 31.6B parameters per token, claims the top aggregates among fully open models in its comparison and matches dense 14–27B rivals while serving long contexts 8–9x faster.","keywords":["mixture of experts","hybrid Mamba-Transformer","German-English model","sovereign AI","open-source LLM","pretraining data transparency","long-context inference","benchmark contamination"],"falsifier":"Have an independent team screen the released per-source data accounting and corrected QA-base dataset against the evaluation items of every benchmark in the reported suites — English and German, including the machine-translated -DE variants — using n-gram and paraphrase-resistant overlap. A single remaining training-data copy of any reported benchmark's test item would falsify the aggregates as capability measurements. Separately replicate the serving protocol (batch 32, TP=1, single B200, latency-subtraction formula, 4K–256K contexts): failure to reproduce the ~4.8k aggregate decode TPS/GPU a","tokens_in":52228,"feed_emoji":"🇩🇪","tokens_out":13587,"duration_ms":119070,"temperature":0.7,"pith_summary":"This paper claims that a sovereign, genuinely open foundation model can sit on the same capability-per-active-parameter frontier as the strongest international releases. Soofi S 30B-A3B activates only ~3.2 of its ~31.6 billion parameters per token, and because 23 of its 52 layers are recurrent Mamba-2 layers rather than attention, its per-sequence cache stays near-constant as context grows — which the authors measure as an 8–9x decode-throughput advantage over dense 14–27B models at 40K context. Trained on ~27 trillion tokens with German deliberately up-weighted (7.2% in the diverse phase, 15.3% in the annealing phase), it posts the highest English and German aggregate scores among fully open models in the comparison and matches or beats every European sovereign baseline on its German suite. The paper also discloses that GPQA evaluation items leaked into the training data, removes that benchmark from all aggregates for every model, and documents the corrective safeguards. If the results hold, other language communities gain a fully documented, rebuildable template for capable and efficient models beyond English.","feed_headline":"Open German model matches 14–27B rivals at 3B active parameters","feed_subtitle":"The fully documented build tops every fully open rival on its suite and serves long contexts 8-9x faster at 40K tokens.","key_machinery":"The carrying mechanism is the hybrid Mamba–Transformer MoE stack: 52 layers interleaving 23 Mamba-2 sequence-mixing layers (fixed-size recurrent state), 23 sparse MoE layers with 128 routed and 2 shared experts (6 active per token), and 6 Grouped-Query-Attention layers, which are the only layers maintaining a key–value cache. This yields ~3.2B active parameters per token and an incremental cache footprint of ~6 KB per token per sequence — 11–53x smaller than dense comparators — which is what keeps decode throughput flat as context grows. Around this sits a three-phase Warmup–Stable–Decay curriculum: ~20T tokens of diverse quality-tiered pretraining, ~6.6T of high-quality annealing in which G","core_discovery":"The paper claims that Soofi S 30B-A3B — a fully open, sovereign German–English base model trained end-to-end on a German HPC cloud — is the strongest fully open model in its comparison on both English and German benchmarks, matches dense 14–27B international models on aggregate performance (English 77.3, German 85.3, both excluding the leaked GPQA benchmark) while activating only 3.2 of its 31.6 billion parameters per token, and matches or outperforms every European sovereign baseline in the comparison on every German benchmark in its suite. The near-constant inference cache is the mechanism: 23 Mamba-2 layers carry most sequence mixing with a fixed-size state, so only 6 attention layers acc","pith_inferences":["The architecture-identical comparison isolates the data recipe as the transferable asset; another language community could plausibly apply the same three-phase, native-language up-weighting curriculum to its own language pair and reproduce gains of the same shape without new architecture work.","Because the audit was reactive — completed only after external discovery — and the n-gram screening safeguard applies to future runs, the reported aggregates are best read as provisional upper bounds on capability until an independent overlap check clears every reported suite, including the machine-translated -DE items.","The near-constant-cache result reframes how models should be reported: two models with equal benchmark scores can differ by an order of magnitude in serving cost, so aggregate scores alone understate the deployment value of hybrid architectures at long context.","The acknowledged long-context weakness — collapse on common-word extraction beyond 32K, diagnosed as a data-mixture gap rather than a backbone limit — makes a concrete next step available: adding retrieval- and aggregation-style synthetic data in the 32K–1M window should close the gap while leaving the rest of the long-context profile intact."],"forward_implications":["German capability can be bought with data allocation: raising German to 15.3% of the annealing mixture lifts the German aggregate by 4.6 points over the architecture-identical reference while the English aggregate rises 0.6 points, showing bilingual depth need not trade away English.","Fully open releases — weights, per-source data accounting, hyperparameters, training and evaluation code — can reach the capability-per-active-parameter frontier of weight-only international releases, giving other communities a rebuildable template rather than a checkpoint.","At high concurrency and long context, serving cost tracks cache size and memory bandwidth more than parameter count: the design sustains ~4.8k decode tokens/second/GPU at 40K context, with throughput essentially flat from 4K to 256K and a window extended to 1M tokens.","A model can be built end-to-end on sovereign European infrastructure (~253,000 GPU-hours for the ~27T-token run) without relinquishing benchmark competitiveness, addressing deployment under local data-protection rules.","Contamination is a measurable, correctable failure: the incident report shows name-based split selection can leak benchmark items into training, trajectory monitoring does not reveal such leaks, and removing the affected benchmark for all models symmetrically preserves the relative rankings."],"fun_headline_variants":["Open German model beats all fully open rivals with 3B active params","German MoE model matches 14–27B rivals, tops fully open field","Sovereign German model outranks fully open baselines at 3B active","Fully open German model leads benchmarks, uses 3B per token","3B-active German MoE surpasses open rivals, matches dense giants"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The capability claims rest on the completeness of the contamination audit in Section 4.3: if any evaluation item from a reported benchmark — especially a machine-translated -DE variant — remains in the ~27T-token mixture, the headline aggregates measure memorization rather than capability; the audit was conducted by the training team only after outsiders discovered the GPQA leak, and the n-gram screening safeguard was added for future runs, not applied to this one.","fun_headline_variants_meta":{"raw":{"variants":["Open German model beats all fully open rivals with 3B active params","German MoE model matches 14–27B rivals, tops fully open field","Sovereign German model outranks fully open baselines at 3B active","Fully open German model leads benchmarks, uses 3B per token","3B-active German MoE surpasses open rivals, matches dense giants"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000171,"raw_usage":{"total_tokens":1138,"prompt_tokens":803,"completion_tokens":335,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":234}},"tokens_in":547,"tokens_out":335,"duration_ms":3968,"temperature":1.0,"reasoning_tokens":234,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T07:36:23.480411+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an independent team screen the released per-source data accounting and corrected QA-base dataset against the evaluation items of every benchmark in the reported suites — English and German, including the machine-translated -DE variants — using n-gram and paraphrase-resistant overlap. A single remaining training-data copy of any reported benchmark's test item would falsify the aggregates as capability measurements. Separately replicate the serving protocol (batch 32, TP=1, single B200, latency-subtraction formula, 4K–256K contexts): failure to reproduce the ~4.8k aggregate decode TPS/GPU a","supporting_citations":[],"review_version":3}