REVIEW 3 major objections 3 minor 103 references
Soofi S, a fully open German–English model activating 3.2 of its 31.6B parameters per token, claims the top aggregates among fully open models in its comparison and matches dense 14–27B rivals while serving long contexts 8–9x faster.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 07:36 UTC pith:UHY7DV6U
load-bearing objection An unusually honest, well-disclosed German-English pretraining report whose headline benchmark claims are conditional on a contamination audit that covers only the QA-base slice and never screens the web or MT-German channels most likely to hide eval items. the 3 major comments →
A Sovereign, Open-Source Foundation Model for German and English
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that Soofi S 30B-A3B — a fully open, sovereign German–English base model trained end-to-end on a German HPC cloud — is the strongest fully open model in its comparison on both English and German benchmarks, matches dense 14–27B international models on aggregate performance (English 77.3, German 85.3, both excluding the leaked GPQA benchmark) while activating only 3.2 of its 31.6 billion parameters per token, and matches or outperforms every European sovereign baseline in the comparison on every German benchmark in its suite. The near-constant inference cache is the mechanism: 23 Mamba-2 layers carry most sequence mixing with a fixed-size state, so only 6 attention layers acc
What carries the argument
The carrying mechanism is the hybrid Mamba–Transformer MoE stack: 52 layers interleaving 23 Mamba-2 sequence-mixing layers (fixed-size recurrent state), 23 sparse MoE layers with 128 routed and 2 shared experts (6 active per token), and 6 Grouped-Query-Attention layers, which are the only layers maintaining a key–value cache. This yields ~3.2B active parameters per token and an incremental cache footprint of ~6 KB per token per sequence — 11–53x smaller than dense comparators — which is what keeps decode throughput flat as context grows. Around this sits a three-phase Warmup–Stable–Decay curriculum: ~20T tokens of diverse quality-tiered pretraining, ~6.6T of high-quality annealing in which G
Load-bearing premise
The capability claims rest on the completeness of the contamination audit in Section 4.3: if any evaluation item from a reported benchmark — especially a machine-translated -DE variant — remains in the ~27T-token mixture, the headline aggregates measure memorization rather than capability; the audit was conducted by the training team only after outsiders discovered the GPQA leak, and the n-gram screening safeguard was added for future runs, not applied to this one.
What would settle it
Have an independent team screen the released per-source data accounting and corrected QA-base dataset against the evaluation items of every benchmark in the reported suites — English and German, including the machine-translated -DE variants — using n-gram and paraphrase-resistant overlap. A single remaining training-data copy of any reported benchmark's test item would falsify the aggregates as capability measurements. Separately replicate the serving protocol (batch 32, TP=1, single B200, latency-subtraction formula, 4K–256K contexts): failure to reproduce the ~4.8k aggregate decode TPS/GPU a
If this is right
- German capability can be bought with data allocation: raising German to 15.3% of the annealing mixture lifts the German aggregate by 4.6 points over the architecture-identical reference while the English aggregate rises 0.6 points, showing bilingual depth need not trade away English.
- Fully open releases — weights, per-source data accounting, hyperparameters, training and evaluation code — can reach the capability-per-active-parameter frontier of weight-only international releases, giving other communities a rebuildable template rather than a checkpoint.
- At high concurrency and long context, serving cost tracks cache size and memory bandwidth more than parameter count: the design sustains ~4.8k decode tokens/second/GPU at 40K context, with throughput essentially flat from 4K to 256K and a window extended to 1M tokens.
- A model can be built end-to-end on sovereign European infrastructure (~253,000 GPU-hours for the ~27T-token run) without relinquishing benchmark competitiveness, addressing deployment under local data-protection rules.
- Contamination is a measurable, correctable failure: the incident report shows name-based split selection can leak benchmark items into training, trajectory monitoring does not reveal such leaks, and removing the affected benchmark for all models symmetrically preserves the relative rankings.
Where Pith is reading between the lines
- The architecture-identical comparison isolates the data recipe as the transferable asset; another language community could plausibly apply the same three-phase, native-language up-weighting curriculum to its own language pair and reproduce gains of the same shape without new architecture work.
- Because the audit was reactive — completed only after external discovery — and the n-gram screening safeguard applies to future runs, the reported aggregates are best read as provisional upper bounds on capability until an independent overlap check clears every reported suite, including the machine-translated -DE items.
- The near-constant-cache result reframes how models should be reported: two models with equal benchmark scores can differ by an order of magnitude in serving cost, so aggregate scores alone understate the deployment value of hybrid architectures at long context.
- The acknowledged long-context weakness — collapse on common-word extraction beyond 32K, diagnosed as a data-mixture gap rather than a backbone limit — makes a concrete next step available: adding retrieval- and aggregation-style synthetic data in the 32K–1M window should close the gap while leaving the rest of the long-context profile intact.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Soofi S 30B-A3B, a 31.6B-parameter mixture-of-experts hybrid Mamba-Transformer base model with ~3.2B active parameters, pretrained on roughly 26.68T tokens with deliberately up-weighted German. The authors document a three-phase curriculum (20T diverse pretraining, ~6.6T high-quality annealing, ~0.1T long-context extension), release full per-source token accounting, training hyperparameters, intermediate checkpoints, and evaluation code, and evaluate the model against 15 (elsewhere 16 or 17) open and open-weight baselines. The central claims are that Soofi S is the strongest fully open model in the evaluation on English and German aggregates, matches dense 14–27B international models on aggregate performance at a fraction of the active parameter cost, achieves best-in-comparison code aggregates, and sustains 8–9× the aggregate decode throughput of dense baselines at 40K context. The paper also discloses a benchmark-contamination incident involving GPQA in the QA-base pretraining constituent and describes a remediation that removes GPQA from reported aggregates and adds forward-looking screening practices.
Significance. If the capability claims survive scrutiny, this is a significant contribution: a fully documented, sovereign, German-English pretraining run with unusually complete data accounting, an architecture-identical baseline that cleanly isolates the data recipe, and a credible serving-efficiency advantage from the hybrid Mamba-MoE design. The release of weights, selected checkpoints, exact per-source token counts, hyperparameters, training/evaluation code, and even a discarded final-annealing stage is exemplary for reproducibility. The long-context weakness on common-word extraction is disclosed honestly. However, the headline capability claims are entirely benchmark-based, and the paper's own contamination disclosure establishes that benchmark material entered the training mixture through at least one pathway. The completeness of the contamination audit is therefore load-bearing for the central claims.
major comments (3)
- [Section 4.3 (GPQA Contamination Disclosure)] The audit's scope is a load-bearing limitation. The re-audit explicitly covers 'all QA-base constituents' — roughly 3.3B tokens, about 0.05% of the Phase 2 pool (Section 3.3) — and only the failure mode of evaluation-only benchmarks whose sole published split is mislabeled 'train'. The remaining ~99.95% of the ~26.68T-token corpus, including the deliberately up-sampled English web tiers (Nemotron-CC, 11.6T effective tokens in Phase 1) and the 571B-token German machine translation of ClimbMix, is not screened against the evaluation suite. This matters doubly for the German benchmarks: the -DE evaluation sets are German renderings of English items, so MT-German web text containing the underlying English items is a direct near-duplicate channel to the German eval sets that carry the flagship +5.9 German aggregate margin. The paper's own Figure 15 concedes that paraphrased contamination is i
- [Section 3.3 and Tables 4–5] The manuscript states that QA-base contains 'paraphrased training splits of 25 standard NLP benchmarks in English and German', and that the model trained on QA-base. The evaluation suite includes benchmarks from the same families (code, math, QA, knowledge). Training on paraphrased train splits of a benchmark family can inflate downstream scores on that family even when no evaluation item is duplicated. The contamination disclosure in §4.3 addresses only evaluation-set leakage of four specific datasets, not this broader train-split exposure. The paper does not list the 25 benchmarks, nor does it analyze which of the reported English or German eval tasks have train-split overlap with QA-base. Because several of the largest reported margins are on code and math tasks (HumanEval +10.8, MBPP-DE +13.4, Minerva +24.2), this is potentially a direct confound for the 'strongest fully open model'
- [Section 4.3, 'Remediation'] The n-gram screening of final training mixtures against the evaluation suite is described as a forward-looking practice ('final training mixtures are screened ... before training'), not as a result for this run. Removing GPQA from the reported aggregates is a necessary correction but does not repair the possibility that other evaluation material entered through unscreened web or MT-German data, nor does it address the train-split exposure identified above. The claims in the Contributions section and Conclusion — 'strongest fully open model', 'matches dense 14–27B models', 'first European sovereign model to sit on the same capability-per-active-parameter frontier' — are therefore stronger than the evidence currently supports. A revision should either supply screening results for this run or substantially qualify these claims, e.g., by stating that they hold 'barring undetected contaminati
minor comments (3)
- [Abstract and Conclusion] The number of comparison models is inconsistent: the Abstract says 'among 17 open base models', Section 4 says 'against 15 open-source and open-weight base models', and the Conclusion says 'unified evaluation of 16 open base models'. Please reconcile these counts.
- [Section 4.4, Eq. (1)] Equation (1) subtracts t(1) from t(1024), which removes prefill cost only if the per-token decode time is approximately linear in output length. The paper should state this linearity assumption explicitly, since the 'TTFT-like' t(1) values are reported separately and the aggregation of prefill and decode into a single TPS figure may be sensitive to the chosen output-length range.
- [Section 2.2 and Appendix E] The long-context comparison with Nemotron 3 Nano is a strength, but the RULER CWE collapse beyond 32K is a substantial capability gap that is only visible in Appendix E. Consider foregrounding this limitation in the main text rather than only in an appendix.
Circularity Check
No circular steps: the paper's central claims are external measurements and controlled comparisons, not derivations that reduce to their own inputs; the contamination-audit scope is an evaluation-validity risk, not a construction-level circularity.
full rationale
The paper's central claims are empirical measurements on a released model, not predictions derived from fitted parameters. The architecture is adopted without modification from an external reference (Nemotron 3 Nano), and the serving-efficiency numbers are measured under the described latency-subtraction protocol (Eq. 1) against external baselines. The German-data proxy ablation in Appendix C is a design input that informed the final mixture; the final checkpoint's benchmark scores are reported as outcomes, not as predictions of the proxy model, and the proxy and final evaluation are not conflated into a single fitted quantity. The GPQA contamination section is a disclosure and re-audit rather than a derivation: the paper removes the contaminated benchmarks (GPQA-Diamond and GPQA-Diamond-DE) from all reported aggregates, so the headline comparisons are recomputed symmetrically without them. The claim that only four QA-base constituents share the split-name failure mode is a factual audit statement about data provenance, not an equation that reduces to its inputs; it could be wrong, but its wrongness would be an empirical error, not circularity. The audit's narrow scope — covering QA-base but not screening the web and MT-German corpus against the full evaluation suite for this run — is a genuine and load-bearing evaluation-validity risk for the benchmark-based claims, and should be weighed as a correctness concern, but it does not make the benchmark results equivalent to the training inputs by construction. Self-citations such as JQL, KletterMix, Propella, and Modalities are data-source or implementation references; they are not used to justify the headline capability comparisons. Therefore no circular step is established under the required evidentiary standard.
Axiom & Free-Parameter Ledger
free parameters (4)
- German share of pretraining mixture =
7.2% (Phase 1), 15.32% (Phase 2)
- Per-source epoch multipliers =
0-10 epochs per source (web HQ x3, Specialized x5, QA-base x10)
- Phase token budgets =
20T stable / 5T decay / 1.58T constant / 0.30T discarded / 0.10T long-context
- Evaluation-suite membership for headline aggregates =
77.3 EN / 85.3 DE aggregates; LBPP excluded from code aggregates; GPQA + held-out group withdrawn
axioms (4)
- domain assumption Benchmark scores measured with lm-evaluation-harness under identical prompts and few-shot settings are directly comparable across base models with different training corpora.
- ad hoc to paper The post-hoc contamination audit is complete: only GPQA, TruthfulQA, BLiMP, and Inverse Scaling leaked evaluation material into training.
- domain assumption The machine-translated German benchmark variants (-DE) are valid, leakage-free measures of German capability.
- domain assumption Equation 1's latency-subtraction protocol adequately isolates aggregate decode throughput from prefill cost on a single B200 at batch 32.
read the original abstract
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.
Figures
Reference graph
Works this paper leans on
-
[1]
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
J.Ainslie, J.Lee-Thorp, M.DeJong, Y.Zemlyanskiy, F.Lebrón, andS.Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895–4901, 2023
2023
-
[2]
M. Ali, M. Brack, M. Lübbering, E. Wendt, A. G. Khan, R. Rutmann, A. Jude, M. Kraus, A. A. Weber, F. Stollenwerk, D. Kaczér, F. Mai, L. Flek, R. Sifa, N. Flores-Herr, J. Koehler, P. Schramowski, M. Fromm, and K. Kersting. Judging quality across languages: A multilin- gual approach to pretraining data filtering with language models. In C. Christodoulopoulo...
2025
-
[3]
M. Ali, M. Fromm, K. Thellmann, J. Ebert, A. A. Weber, R. Rutmann, C. Jain, M. Lübbering, D. Steinigen, J. Leveling, K. Klug, J. S. Buschhoff, L. Jurkschat, H. Abdelwahab, B. J. Stein, K.-H. Sylla, P. Denisov, N. Brandizzi, Q. Saleem, A. Bhowmick, L. Helmer, C. John, P. O. Suarez, M. Ostendorff, A. Jude, L. Manjunath, S. Weinbach, C. Penke, O. Filatov, F....
Pith/arXiv arXiv 2025
-
[4]
E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, M. Debbah, É. Goffinet, D. Hesslow, J. Launay, Q. Malartic, D. Mazzotta, B. Noune, B. Pannier, and G. Penedo. The falcon series of open language models.arXiv preprint arXiv:2311.16867, 2023. 36
Pith/arXiv arXiv 2023
-
[5]
P. Apertus, A. Hernández-Cano, A. Hägele, A. H. Huang, A. Romanou, A.-J. Solergibert, B. Pasztor, B. Messmer, D. Garbaya, E. F. Ďurech, I. Hakimi, J. G. Giraldo, M. Ismayilzada, N. Foroutan, S. Moalla, T. Chen, V. Sabolčec, Y. Xu, M. Aerni, B. AlKhamissi, I. A. Mar- iñas, M. H. Amani, M. Ansaripour, I. Badanin, H. Benoit, E. Boros, N. Browning, F. Bösch, ...
arXiv 2025
-
[6]
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al. Program synthesis with large language models.arXiv preprint arXiv:2108.07732, 2021
Pith/arXiv arXiv 2021
-
[7]
Biderman, H
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, A. Skowron, L. Sutawika, and O. Van Der Wal. Pythia: A suite for analyzing large language models across training and scaling. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors,Proceedings of...
2023
-
[8]
BLOOM: A 176b-parameter open-access multilingual language model
BigScience Workshop. BLOOM: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022
Pith/arXiv arXiv 2022
-
[9]
Black, S
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. Mc- Donell, J. Phang, M. Pieler, U. S. Prashanth, S. Purohit, L. Reynolds, J. Tow, B. Wang, and S. Weinbach. GPT-NeoX-20B: An open-source autoregressive language model. In A. Fan, S. Ilic, T. Wolf, and M. Gallé, editors,Proceedings of BigScience Episode #5 – Workshop o...
2022
-
[10]
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S.Gray, B.Chess, J.Clark, C.Berner, S.McCandlish, A.Radford, I.Sutskever, andD.Amodei. Langu...
1901
-
[11]
Burchell, O
L. Burchell, O. de Gibert, N. Arefyev, M. Aulamo, M. Bañón, P. Chen, M. Fedorova, L. Guillou, B. Haddow, J. Hajič, J. Helcl, E. Henriksson, M. Klimaszewski, V. Komulainen, A. Kutuzov, J. Kytöniemi, V. Laippala, P. Mæhlum, B. Malik, F. Mehryary, V. Mikhailov, N. Moghe, A. Myntti, D. O’Brien, S. Oepen, P. Pal, J. Piha, S. Pyysalo, G. Ramírez-Sánchez, D. Sam...
2025
-
[12]
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021
Pith/arXiv arXiv 2021
-
[13]
Clark, I
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
2018
-
[14]
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021
Pith/arXiv arXiv 2021
-
[15]
D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu, et al. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1280–1297, 2024
2024
-
[16]
T. Dao and A. Gu. Transformers are ssms: Generalized models and efficient algorithms through structured state space duality.arXiv preprint arXiv:2405.21060, 2024
Pith/arXiv arXiv 2024
-
[17]
Deutsche Telekom AG. Germany’s first AI factory for industry officially goes into operation in Munich.https://www.telekom.com/en/media/media-information/archive/germany-s-f irst-ai-factory-for-industry-1101670, Feb. 2026. Industrial AI Cloud, Munich; accessed 2026-07-08
2026
-
[18]
S. Diao, Y. Yang, Y. Fu, X. Dong, D. Su, M. Kliegl, Z. Chen, P. Belcak, Y. Suhara, H. Yin, et al. Nemotron-climb: Clustering-based iterative data mixture bootstrapping for language model pre-training.Advances in Neural Information Processing Systems, 38, 2026
2026
-
[19]
Fedus, B
W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022
2022
-
[20]
K. Fujii, Y. Tajima, S. Mizuki, M. Kawamura, H. Shimada, T. Shiotani, K. Saito, M. Oi, T. Nakamura, T. Okamoto, S. Ishida, K. Hattori, Y. Ma, H. Takamura, R. Yokota, J. Sakuma, and N. Okazaki. Rewriting pre-training data boosts llm performance in math and code.arXiv preprint arXiv:2505.02881, 2026
arXiv 2026
-
[21]
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy. The pile: An 800gb dataset of diverse text for language modeling.arXiv preprint arXiv:2101.00027, 2020
Pith/arXiv arXiv 2020
-
[22]
L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou. The language model evaluation harness, 07 2024
2024
-
[23]
GENIOS – german business and press database.https://www.genios.de/browse/Alle, 2025
GBI-Genios Deutsche Wirtschaftsdatenbank GmbH. GENIOS – german business and press database.https://www.genios.de/browse/Alle, 2025. Commercially licensed corpus of 38 German newspaper and trade-press archives; obtained under a license that does not permit redistribution
2025
-
[24]
L. Gienapp, C. Schröder, S. Schweter, C. Akiki, F. Schlatt, A. Zimmermann, P. Genêt, and M. Potthast. The german commons - 154 billion tokens of openly licensed text for german language models.arXiv preprint arXiv:2510.13996, 2025
arXiv 2025
-
[25]
Reasoning-v1-20m, 2025
GlaiveAI. Reasoning-v1-20m, 2025. A synthetic reasoning dataset containing 22mil+ general reasoning questions and responses generated using deepseek-ai/DeepSeek-R1-Distill-Llama- 70B
2025
-
[26]
A. Gonzalez-Agirre, M. Pàmies, J. Llop, I. Baucells, S. D. Dalt, D. Tamayo, J. J. Saiz, F. Es- puña, J. Prats, J. Aula-Blasco, M. Mina, I. Pikabea, A. Rubio, A. Shvets, A. Sallés, I. La- cunza, J. Palomar, J. Falcão, L. Tormo, L. Vasquez-Reina, M. Marimon, O. Pareras, V. Ruiz- Fernández, and M. Villegas. Salamandra technical report.arXiv preprint arXiv:25...
Pith/arXiv arXiv 2025
-
[27]
Groeneveld, I
D. Groeneveld, I. Beltagy, E. Walsh, A. Bhagia, R. Kinney, O. Tafjord, A. Jha, H. Ivison, I. Magnusson, Y. Wang, S. Arora, D. Atkinson, R. Authur, K. Chandu, A. Cohan, J. Dumas, Y. Elazar, Y. Gu, J. Hessel, T. Khot, W. Merrill, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. Peters, V. Pyatkin, A. Ravichander, D. Schwenk, S. Shah, W. Smith, E. Strubell, ...
2024
-
[28]
Gurgurov and T
D. Gurgurov and T. Röhr. ReasonXL: A multilingual cross-domain reasoning corpus.https: //huggingface.co/datasets/toroe/Soofi-Think-SFT-10B-multilingual, 2026
2026
-
[29]
Gurgurov, T
D. Gurgurov, T. Röhr, and S. Ostermann. Nemotron-multilingual-reasoning: A multilingual science reasoning dataset.https://huggingface.co/datasets/DGurgurov/Nemotron-Multi lingual-Reasoning, 2025. Derived from nvidia/Llama-Nemotron-Post-Training-Dataset with machine-translated reasoning traces
2025
-
[30]
Hägele, E
A. Hägele, E. Bakouch, A. Kosson, L. B. Allal, L. Von Werra, and M. Jaggi. Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024
2024
-
[31]
Hendrycks, C
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding, 2021
2021
-
[32]
Hendrycks, C
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Stein- hardt. Measuring mathematical problem solving with the math dataset. In J. Vanschoren and S. Yeung, editors,Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021
2021
-
[33]
Hoffmann, S
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. Rae, and L. Sifre. An empirical analysis of compute-optimal large language model training. InAdvanc...
2022
-
[34]
Hsieh, S
C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, Y. Zhang, and B. Ginsburg. Ruler: What’s the real context size of your long-context language models?, 2024
2024
-
[35]
S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, et al. Minicpm: Unveiling the potential of small language models with scalable training strategies. arXiv preprint arXiv:2404.06395, 2024
Pith/arXiv arXiv 2024
-
[36]
Idahl, B
M. Idahl, B. Droste, B. Plüster, and J. P. Harries. propella-1: Multi-property document annotation for llm data curation at scale, 2026
2026
-
[37]
Idahl, J
M. Idahl, J. Tiedemann, S. Pyysalo, D. Salinas, T. Galica, S. Qian, T. N. Mateiu, Z. Li, A. Lokrantz, F. Vitiugin, A. F. T. Martins, J. Kanerva, F. Ginter, M. Lindemann, T. Isbis- ter, B. Moell, J. Lindh, J. Hajič, J. Jitsev, A. Kutuzov, S. Oepen, and G. Ramírez-Sánchez. Multisynt/mt: Trillion-token multi-parallel pre-training data translated across 36 la...
2026
-
[38]
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bres- sand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. Le Scao, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed. Mistral 7b.arXiv preprint arXiv:2310.06825, 2023
Pith/arXiv arXiv 2023
-
[39]
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. Bou Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. Le Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed. Mixtral of experts.arXiv prepri...
Pith/arXiv arXiv 2024
-
[40]
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Rad- ford, J. Wu, and D. Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020
Pith/arXiv arXiv 2001
-
[41]
J. Krajewski, J. Ludziejewski, K. Adamczewski, M. Pióro, M. Krutul, S. Antoniak, K. Ciebiera, K. Król, T. Odrzygóźdź, P. Sankowski, et al. Scaling laws for fine-grained mixture of experts. arXiv preprint arXiv:2402.07871, 2024
Pith/arXiv arXiv 2024
-
[42]
M. Kraus, R. Härle, S. Sztwiertnia, A. G. Khan, M. Ali, M. Fromm, and K. Kersting. Kletter- mix: Climbing toward high-quality german pretraining data.arXiv preprint arXiv:2606.03773, 2026
Pith/arXiv arXiv 2026
-
[43]
Kwiatkowski, J
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov. Natural questions: A benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:452–466, 2019
2019
-
[44]
Kydlíček, G
H. Kydlíček, G. Penedo, and L. von Werra. Finepdfs.https://huggingface.co/datasets/ HuggingFaceFW/finepdfs, 2025
2025
-
[45]
Lewkowycz, A
A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra. Solving quantitative reasoning problems with language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors,Advances in Neural Information Processing Syste...
2022
-
[46]
Liesenfeld, M
A. Liesenfeld, M. Dingemanse, D. Blankvoort, N. Kalra, and A. R. Golkhandan. European open source ai definitions.https://osai-index.eu/osai-definitions, 2026. Accessed: 2026-07-06
2026
-
[47]
A. H. Liu, K. Khandelwal, S. Subramanian, V. Jouault, A. Rastogi, A. Sadé, A. Jeffares, A. Jiang, A. Cahill, A. Gavaudan, A. Sablayrolles, A. Héliou, A. You, A. Ehrenberg, A. Lo, A. Eliseev, A. Calvi, A. Sooriyarachchi, B. Bout, B. Rozière, B. D. Monicault, C. Lanfranchi, C. Barreau, C. Courtot, D. Grattarola, D. Dabert, D. de las Casas, E. Chane-Sane, F....
Pith/arXiv arXiv 2026
-
[48]
Z. Liu, Z. Yang, Y. Chen, C. Lee, M. Shoeybi, B. Catanzaro, and W. Ping. Acereason- nemotron 1.1: Advancing math and code reasoning through sft and rl synergy.arXiv preprint arXiv:2506.13284, 2025
Pith/arXiv arXiv 2025
-
[49]
I. Loshchilov and F. Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017
Pith/arXiv arXiv 2017
-
[50]
M. Lübbering, T. Ruland, R. Rutmann, F. Stollenwerk, D. Fitzek, M. Fromm, A. Weber, R. Sifa, N. Flores-Herr, J. Köhler, et al. Modalities, a pytorch-native framework for large-scale llm training and research.arXiv preprint arXiv:2602.08387, 2026
arXiv 2026
-
[51]
R. K. Mahabadi, S. Satheesh, S. Prabhumoye, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-cc-math: A 133 billion-token-scale high quality math pretraining dataset.arXiv preprint arXiv:2508.15096, 2025
Pith/arXiv arXiv 2025
-
[52]
P. H. Martins, J. Alves, P. Fernandes, N. M. Guerreiro, R. Rei, A. Farajian, M. Klimaszewski, D. M. Alves, J. Pombal, N. Boizard, M. Faysse, P. Colombo, F. Yvon, B. Haddow, J. G. C. de Souza, A. Birch, and A. F. T. Martins. Eurollm-9b: Technical report.arXiv preprint arXiv:2506.04079, 2025
Pith/arXiv arXiv 2025
-
[53]
P. H. Martins, P. Fernandes, J. Alves, N. M. Guerreiro, R. Rei, D. M. Alves, J. Pombal, A. Farajian, M. Faysse, M. Klimaszewski, P. Colombo, B. Haddow, J. G. C. de Souza, A. Birch, and A. F. T. Martins. Eurollm: Multilingual language models for europe.arXiv preprint arXiv:2409.16235, 2024
Pith/arXiv arXiv 2024
-
[54]
Matton, T
A. Matton, T. Sherborne, D. Aumiller, E. Tommasone, M. Alizadeh, J. He, R. Ma, M. Voisin, E. Gilsenan-McMahon, and M. Gallé. On leakage of code generation evaluation datasets. In 41 Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, editors,Findings of the Association for Com- putational Linguistics: EMNLP 2024, pages 13215–13223, Miami, Florida, USA, Nov. 2024. A...
2024
-
[55]
Mistral small 3.1.https://mistral.ai/news/mistral-small-3-1/, 2025
Mistral AI. Mistral small 3.1.https://mistral.ai/news/mistral-small-3-1/, 2025. Model release, March 2025
2025
-
[56]
Morrison, S
J. Morrison, S. Adhikesaven, A. Bhagia, M. Zaharia, N. A. Smith, and S. Min. Train separately, merge together: Modular post-training with mixture-of-experts, 2026
2026
-
[57]
Muennighoff, L
N. Muennighoff, L. Soldaini, D. Groeneveld, K. Lo, J. Morrison, S. Min, W. Shi, E. P. Walsh, O. Tafjord, N. Lambert, Y. Gu, S. Arora, A. Bhagia, D. Schwenk, D. Wadden, A. Wettig, B. Hui, T. Dettmers, D. Kiela, A. Farhadi, N. A. Smith, P. W. Koh, A. Singh, and H. Ha- jishirzi. OLMoE: Open mixture-of-experts language models. InThe Thirteenth International C...
2025
-
[58]
Mt reasoning: Automatic translations into 2 languages of glaive ai reasoning dataset, 2025
MultiSynt. Mt reasoning: Automatic translations into 2 languages of glaive ai reasoning dataset, 2025. An automatic translations into 2 languages of Glaive AI reasoning dataset
2025
-
[59]
Basant, A
NVIDIA, :, A. Basant, A. Khairnar, A. Paithankar, A. Khattar, A. Renduchintala, A. Malte, A. Bercovich, A. Hazare, A. Rico, A. Ficek, A. Kondratenko, A. Shaposhnikov, A. Bukharin, A. Taghibakhshi, A. Barton, A. S. Mahabaleshwarkar, A. Shen, A. Tao, A. Guan, A. Shors, A. Mandarwal, A. Mehta, A. Venkatesan, A. Sharabiani, A. Aithal, A. Poojary, A. Dattagupt...
2025
-
[60]
NVIDIA, :, A. Blakeman, A. Basant, A. Khattar, A. Renduchintala, A. Bercovich, A. Ficek, A. Bjorlin, A. Taghibakhshi, A. S. Deshmukh, A. S. Mahabaleshwarkar, A. Tao, A. Shors, A. Aithal, A. Poojary, A. Dattagupta, B. Buddharaju, B. Chen, B. Ginsburg, B. Wang, B. Norick, B. Butterfield, B. Catanzaro, C. del Mundo, C. Dong, C. Harvey, C. Parisien, D. Su, D....
Pith/arXiv arXiv 2025
-
[61]
NVIDIA, :, A. Blakeman, A. Grattafiori, A. Basant, A. Gupta, A. Khattar, A. Renduchintala, A. Vavre, A. Shukla, A. Bercovich, A. Ficek, A. Shaposhnikov, A. Kondratenko, A. Bukharin, A. Milesi, A. Taghibakhshi, A. Liu, A. Barton, A. S. Mahabaleshwarkar, A. Klein, A. Zuker, A. Geifman, A. Shen, A. Bhiwandiwalla, A. Tao, A. Guan, A. Mandarwal, A. Mehta, A. A...
arXiv 2025
-
[62]
Deutsche telekom and NVIDIA launch industrial AI cloud.https://blogs.nvidia .com/blog/germany-industrial-ai-cloud-launch/, Nov
NVIDIA. Deutsche telekom and NVIDIA launch industrial AI cloud.https://blogs.nvidia .com/blog/germany-industrial-ai-cloud-launch/, Nov. 2025. Accessed 2026-07-08
2025
-
[63]
S. Oepen, N. Arefev, M. Aulamo, M. Bañón, M. Buljan, L. Burchell, L. Charpentier, P. Chen, M. Fedorova, O. de Gibert, B. Haddow, J. Hajič, J. Helcl, A. Kutuzov, V. Laippala, Z. Li, R. Luukkonen, B. Malik, V. Mikhailov, A. Myntti, D. O’Brien, L. Poláková, S. Pyysalo, G. R. Sánchez, J. Siewert, P. Stepachev, J. Tiedemann, T. Vahtola, D. Variš, F. Vitiugin, ...
Pith/arXiv arXiv 2026
-
[64]
T. Olmo, :, A. Ettinger, A. Bertsch, B. Kuehl, D. Graham, D. Heineman, D. Groeneveld, F. Brahman, F. Timbers, H. Ivison, J. Morrison, J. Poznanski, K. Lo, L. Soldaini, M. Jordan, M. Chen, M. Noukhovitch, N. Lambert, P. Walsh, P. Dasigi, R. Berry, S. Malik, S. Shah, S. Geng, S. Arora, S. Gupta, T. Anderson, T. Xiao, T. Murray, T. Romero, V. Graf, A. Asai, ...
Pith/arXiv arXiv 2026
-
[65]
The open source ai definition, version 1.0.https://opensource.org /ai/open-source-ai-definition, 2024
Open Source Initiative. The open source ai definition, version 1.0.https://opensource.org /ai/open-source-ai-definition, 2024. Accessed: 2026-07-06. 44
2024
-
[66]
OpenEuroLLM: A series of foundation models for transparent ai in europe, 2025
OpenEuroLLM Consortium. OpenEuroLLM: A series of foundation models for transparent ai in europe, 2025. Project website, accessed 2026-06-25
2025
-
[67]
G. Penedo. Finewiki, 2025. Source: Wikimedia Enterprise Snapshot API (https://api.enterprise.wikimedia.com/v2/snapshots). Text licensed under CC BY-SA 4.0 with attribution to Wikipedia contributors
2025
-
[68]
Penedo, H
G. Penedo, H. Kydlíček, A. Lozhkov, M. Mitchell, C. Raffel, L. Von Werra, T. Wolf, et al. The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849, 2024
2024
-
[69]
SYNTH: An open generalist synthetic dataset for training small reason- ing models.https://huggingface.co/datasets/PleIAs/SYNTH, 2025
Pleias and AI Alliance. SYNTH: An open generalist synthetic dataset for training small reason- ing models.https://huggingface.co/datasets/PleIAs/SYNTH, 2025. Dataset comprising 79,648,272 text samples (over 41 billion words) derived from the synthetic amplification of 58,698 Wikipedia and Wikibooks articles. Licensed under CC-BY-SA 4.0
2025
-
[70]
J. Poznanski, A. Rangapur, J. Borchardt, J. Dunkelberger, R. Huff, D. Lin, C. Wilhelm, K. Lo, andL.Soldaini. olmocr: Unlockingtrillionsoftokensinpdfswithvisionlanguagemodels.arXiv preprint arXiv:2502.18443, 2025
arXiv 2025
-
[71]
Qwen3.5: Towards native multimodal agents, February 2026
Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026
2026
-
[72]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. Technical report, OpenAI, 2019
2019
-
[73]
Raffel, N
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of Machine Learning Research, 21(140):1–67, 2020
2020
-
[74]
M. M. Ramos, D. M. Alves, H. Gisserot-Boukhlef, J. Alves, P. H. Martins, P. Fernandes, J. Pombal, N. M. Guerreiro, R. Rei, N. Boizard, A. Farajian, M. Klimaszewski, J. G. C. de Souza, B. Haddow, F. Yvon, P. Colombo, A. Birch, and A. F. T. Martins. Eurollm-22b: Technical report.arXiv preprint arXiv:2602.05879, 2026
arXiv 2026
-
[75]
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark, 2023
2023
-
[76]
Romanou, N
A. Romanou, N. Foroutan, A. Sotnikova, S. H. Nelaturu, S. Singh, R. Maheshwary, M. Al- tomare, Z. Chen, M. Haggag, S. A, A. Amayuelas, A. H. Amirudin, D. Boiko, M. Chang, J. Chim, G. Cohen, A. K. Dalmia, A. Diress, S. Duwal, D. Dzenhaliou, D. Florez, F. Farestam, J. M. Imperial, S. Islam, P. Isotalo, M. Jabbarishiviari, B. F. Karlsson, E. Khalilov, C. Kla...
2025
-
[77]
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean. Outra- geously large neural networks: The sparsely-gated mixture-of-experts layer.arXiv preprint arXiv:1701.06538, 2017. 45
Pith/arXiv arXiv 2017
-
[78]
W. Shi, A. Bhagia, K. Farhat, N. Muennighoff, J. Morrison, E. Walsh, D. Schwenk, S. Long- pre, J. Poznanski, A. Ettinger, et al. Flexolmo: Open language models for flexible data use. Advances in Neural Information Processing Systems, 38:165943–165974, 2026
2026
-
[79]
Soldaini, R
L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkinson, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y. Elazar, V. Hofmann, A. Jha, S. Kumar, L. Lucy, X. Lyu, N. Lambert, I. Magnus- son, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. Peters, A. Ravichander, K. Richardson, Z. Shen, E. Strubell, N. Subramani, O. Tafjord, E. Walsh, L. Zettlemoyer, N. Smit...
2024
-
[80]
D. Su, K. Kong, Y. Lin, J. Jennings, B. Norick, M. Kliegl, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-CC: Transforming Common Crawl into a refined long-horizon pre- training dataset. In W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.