Pith. sign in

REVIEW 3 major objections 3 minor 103 references

A Sovereign, Open-Source Foundation Model for German and English

T0 review · 3 major / 3 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read Soofi S, a fully open German–English model activating 3.2 of its 31.6B parameters per token, claims the top aggregates among fully open models in its comparison and matches dense 14–27B rivals while serving long contexts 8–9x faster.

desk verdict An unusually honest, well-disclosed German-English pretraining report whose headline benchmark claims are conditional on a contamination audit that covers only the QA-base slice and never screens the web or MT-German channels most likely to hide eval items. read the letter →

arxiv 2607.09424 v3 pith:UHY7DV6U submitted 2026-07-10 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords mixtureofexpertshybridMamba-TransformerGerman-EnglishmodelsovereignAIopen-sourceLLMpretrainingdatatransparencylong-contextinferencebenchmarkcontamination
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a sovereign, genuinely open foundation model can sit on the same capability-per-active-parameter frontier as the strongest international releases. Soofi S 30B-A3B activates only ~3.2 of its ~31.6 billion parameters per token, and because 23 of its 52 layers are recurrent Mamba-2 layers rather than attention, its per-sequence cache stays near-constant as context grows — which the authors measure as an 8–9x decode-throughput advantage over dense 14–27B models at 40K context. Trained on ~27 trillion tokens with German deliberately up-weighted (7.2% in the diverse phase, 15.3% in the annealing phase), it posts the highest English and German aggregate scores among fully open models in the comparison and matches or beats every European sovereign baseline on its German suite. The paper also discloses that GPQA evaluation items leaked into the training data, removes that benchmark from all aggregates for every model, and documents the corrective safeguards. If the results hold, other language communities gain a fully documented, rebuildable template for capable and efficient models beyond English.

What carries the argument

The carrying mechanism is the hybrid Mamba–Transformer MoE stack: 52 layers interleaving 23 Mamba-2 sequence-mixing layers (fixed-size recurrent state), 23 sparse MoE layers with 128 routed and 2 shared experts (6 active per token), and 6 Grouped-Query-Attention layers, which are the only layers maintaining a key–value cache. This yields ~3.2B active parameters per token and an incremental cache footprint of ~6 KB per token per sequence — 11–53x smaller than dense comparators — which is what keeps decode throughput flat as context grows. Around this sits a three-phase Warmup–Stable–Decay curriculum: ~20T tokens of diverse quality-tiered pretraining, ~6.6T of high-quality annealing in which G

What would settle it

Have an independent team screen the released per-source data accounting and corrected QA-base dataset against the evaluation items of every benchmark in the reported suites — English and German, including the machine-translated -DE variants — using n-gram and paraphrase-resistant overlap. A single remaining training-data copy of any reported benchmark's test item would falsify the aggregates as capability measurements. Separately replicate the serving protocol (batch 32, TP=1, single B200, latency-subtraction formula, 4K–256K contexts): failure to reproduce the ~4.8k aggregate decode TPS/GPU a

Watch

Extended reading notes

Core claim

The paper claims that Soofi S 30B-A3B — a fully open, sovereign German–English base model trained end-to-end on a German HPC cloud — is the strongest fully open model in its comparison on both English and German benchmarks, matches dense 14–27B international models on aggregate performance (English 77.3, German 85.3, both excluding the leaked GPQA benchmark) while activating only 3.2 of its 31.6 billion parameters per token, and matches or outperforms every European sovereign baseline in the comparison on every German benchmark in its suite. The near-constant inference cache is the mechanism: 23 Mamba-2 layers carry most sequence mixing with a fixed-size state, so only 6 attention layers acc

Load-bearing premise

The capability claims rest on the completeness of the contamination audit in Section 4.3: if any evaluation item from a reported benchmark — especially a machine-translated -DE variant — remains in the ~27T-token mixture, the headline aggregates measure memorization rather than capability; the audit was conducted by the training team only after outsiders discovered the GPQA leak, and the n-gram screening safeguard was added for future runs, not applied to this one.

Editorial extensions

If this is right

  • German capability can be bought with data allocation: raising German to 15.3% of the annealing mixture lifts the German aggregate by 4.6 points over the architecture-identical reference while the English aggregate rises 0.6 points, showing bilingual depth need not trade away English.
  • Fully open releases — weights, per-source data accounting, hyperparameters, training and evaluation code — can reach the capability-per-active-parameter frontier of weight-only international releases, giving other communities a rebuildable template rather than a checkpoint.
  • At high concurrency and long context, serving cost tracks cache size and memory bandwidth more than parameter count: the design sustains ~4.8k decode tokens/second/GPU at 40K context, with throughput essentially flat from 4K to 256K and a window extended to 1M tokens.
  • A model can be built end-to-end on sovereign European infrastructure (~253,000 GPU-hours for the ~27T-token run) without relinquishing benchmark competitiveness, addressing deployment under local data-protection rules.
  • Contamination is a measurable, correctable failure: the incident report shows name-based split selection can leak benchmark items into training, trajectory monitoring does not reveal such leaks, and removing the affected benchmark for all models symmetrically preserves the relative rankings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The architecture-identical comparison isolates the data recipe as the transferable asset; another language community could plausibly apply the same three-phase, native-language up-weighting curriculum to its own language pair and reproduce gains of the same shape without new architecture work.
  • Because the audit was reactive — completed only after external discovery — and the n-gram screening safeguard applies to future runs, the reported aggregates are best read as provisional upper bounds on capability until an independent overlap check clears every reported suite, including the machine-translated -DE items.
  • The near-constant-cache result reframes how models should be reported: two models with equal benchmark scores can differ by an order of magnitude in serving cost, so aggregate scores alone understate the deployment value of hybrid architectures at long context.
  • The acknowledged long-context weakness — collapse on common-word extraction beyond 32K, diagnosed as a data-mixture gap rather than a backbone limit — makes a concrete next step available: adding retrieval- and aggregation-style synthetic data in the 32K–1M window should close the gap while leaving the rest of the long-context profile intact.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper presents Soofi S 30B-A3B, a 31.6B-parameter mixture-of-experts hybrid Mamba-Transformer base model with ~3.2B active parameters, pretrained on roughly 26.68T tokens with deliberately up-weighted German. The authors document a three-phase curriculum (20T diverse pretraining, ~6.6T high-quality annealing, ~0.1T long-context extension), release full per-source token accounting, training hyperparameters, intermediate checkpoints, and evaluation code, and evaluate the model against 15 (elsewhere 16 or 17) open and open-weight baselines. The central claims are that Soofi S is the strongest fully open model in the evaluation on English and German aggregates, matches dense 14–27B international models on aggregate performance at a fraction of the active parameter cost, achieves best-in-comparison code aggregates, and sustains 8–9× the aggregate decode throughput of dense baselines at 40K context. The paper also discloses a benchmark-contamination incident involving GPQA in the QA-base pretraining constituent and describes a remediation that removes GPQA from reported aggregates and adds forward-looking screening practices.

Significance. If the capability claims survive scrutiny, this is a significant contribution: a fully documented, sovereign, German-English pretraining run with unusually complete data accounting, an architecture-identical baseline that cleanly isolates the data recipe, and a credible serving-efficiency advantage from the hybrid Mamba-MoE design. The release of weights, selected checkpoints, exact per-source token counts, hyperparameters, training/evaluation code, and even a discarded final-annealing stage is exemplary for reproducibility. The long-context weakness on common-word extraction is disclosed honestly. However, the headline capability claims are entirely benchmark-based, and the paper's own contamination disclosure establishes that benchmark material entered the training mixture through at least one pathway. The completeness of the contamination audit is therefore load-bearing for the central claims.

major comments (3)
  1. [Section 4.3 (GPQA Contamination Disclosure)] The audit's scope is a load-bearing limitation. The re-audit explicitly covers 'all QA-base constituents' — roughly 3.3B tokens, about 0.05% of the Phase 2 pool (Section 3.3) — and only the failure mode of evaluation-only benchmarks whose sole published split is mislabeled 'train'. The remaining ~99.95% of the ~26.68T-token corpus, including the deliberately up-sampled English web tiers (Nemotron-CC, 11.6T effective tokens in Phase 1) and the 571B-token German machine translation of ClimbMix, is not screened against the evaluation suite. This matters doubly for the German benchmarks: the -DE evaluation sets are German renderings of English items, so MT-German web text containing the underlying English items is a direct near-duplicate channel to the German eval sets that carry the flagship +5.9 German aggregate margin. The paper's own Figure 15 concedes that paraphrased contamination is i
  2. [Section 3.3 and Tables 4–5] The manuscript states that QA-base contains 'paraphrased training splits of 25 standard NLP benchmarks in English and German', and that the model trained on QA-base. The evaluation suite includes benchmarks from the same families (code, math, QA, knowledge). Training on paraphrased train splits of a benchmark family can inflate downstream scores on that family even when no evaluation item is duplicated. The contamination disclosure in §4.3 addresses only evaluation-set leakage of four specific datasets, not this broader train-split exposure. The paper does not list the 25 benchmarks, nor does it analyze which of the reported English or German eval tasks have train-split overlap with QA-base. Because several of the largest reported margins are on code and math tasks (HumanEval +10.8, MBPP-DE +13.4, Minerva +24.2), this is potentially a direct confound for the 'strongest fully open model'
  3. [Section 4.3, 'Remediation'] The n-gram screening of final training mixtures against the evaluation suite is described as a forward-looking practice ('final training mixtures are screened ... before training'), not as a result for this run. Removing GPQA from the reported aggregates is a necessary correction but does not repair the possibility that other evaluation material entered through unscreened web or MT-German data, nor does it address the train-split exposure identified above. The claims in the Contributions section and Conclusion — 'strongest fully open model', 'matches dense 14–27B models', 'first European sovereign model to sit on the same capability-per-active-parameter frontier' — are therefore stronger than the evidence currently supports. A revision should either supply screening results for this run or substantially qualify these claims, e.g., by stating that they hold 'barring undetected contaminati
minor comments (3)
  1. [Abstract and Conclusion] The number of comparison models is inconsistent: the Abstract says 'among 17 open base models', Section 4 says 'against 15 open-source and open-weight base models', and the Conclusion says 'unified evaluation of 16 open base models'. Please reconcile these counts.
  2. [Section 4.4, Eq. (1)] Equation (1) subtracts t(1) from t(1024), which removes prefill cost only if the per-token decode time is approximately linear in output length. The paper should state this linearity assumption explicitly, since the 'TTFT-like' t(1) values are reported separately and the aggregation of prefill and decode into a single TPS figure may be sensitive to the chosen output-length range.
  3. [Section 2.2 and Appendix E] The long-context comparison with Nemotron 3 Nano is a strength, but the RULER CWE collapse beyond 32K is a substantial capability gap that is only visible in Appendix E. Consider foregrounding this limitation in the main text rather than only in an appendix.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular steps: the paper's central claims are external measurements and controlled comparisons, not derivations that reduce to their own inputs; the contamination-audit scope is an evaluation-validity risk, not a construction-level circularity.

full rationale

The paper's central claims are empirical measurements on a released model, not predictions derived from fitted parameters. The architecture is adopted without modification from an external reference (Nemotron 3 Nano), and the serving-efficiency numbers are measured under the described latency-subtraction protocol (Eq. 1) against external baselines. The German-data proxy ablation in Appendix C is a design input that informed the final mixture; the final checkpoint's benchmark scores are reported as outcomes, not as predictions of the proxy model, and the proxy and final evaluation are not conflated into a single fitted quantity. The GPQA contamination section is a disclosure and re-audit rather than a derivation: the paper removes the contaminated benchmarks (GPQA-Diamond and GPQA-Diamond-DE) from all reported aggregates, so the headline comparisons are recomputed symmetrically without them. The claim that only four QA-base constituents share the split-name failure mode is a factual audit statement about data provenance, not an equation that reduces to its inputs; it could be wrong, but its wrongness would be an empirical error, not circularity. The audit's narrow scope — covering QA-base but not screening the web and MT-German corpus against the full evaluation suite for this run — is a genuine and load-bearing evaluation-validity risk for the benchmark-based claims, and should be weighed as a correctness concern, but it does not make the benchmark results equivalent to the training inputs by construction. Self-citations such as JQL, KletterMix, Propella, and Modalities are data-source or implementation references; they are not used to justify the headline capability comparisons. Therefore no circular step is established under the required evidentiary standard.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claims rest on unmeasured premises: (1) cross-model comparability of harness measurements; (2) completeness of the self-performed contamination audit; (3) validity of the translated German benchmarks; (4) representativeness of the single-GPU throughput protocol. Free parameters — mixture shares, epoch multipliers, phase budgets, suite membership — are chosen by hand or proxy ablation, not derived. There is no mathematical derivation to audit; the claims are empirical and inherit the uncertainty of benchmark and throughput measurements.

free parameters (4)
  • German share of pretraining mixture = 7.2% (Phase 1), 15.32% (Phase 2)
    The defining design choice; selected via 100B-token proxy ablations (Appendix C) on bits-per-byte and rank-choice accuracy. The headline German aggregates are outcomes of this ablation-chosen allocation.
  • Per-source epoch multipliers = 0-10 epochs per source (web HQ x3, Specialized x5, QA-base x10)
    Epoch multipliers set effective token counts and therefore the mixture; chosen by hand per source (Tables 6, 8). The QA-base x10 multiplier put contaminated GPQA items through ten training repeats.
  • Phase token budgets = 20T stable / 5T decay / 1.58T constant / 0.30T discarded / 0.10T long-context
    Chosen by hand following WSD practice; the discarded 0.30T final-annealing stage shows budget choices were evaluated empirically rather than derived.
  • Evaluation-suite membership for headline aggregates = 77.3 EN / 85.3 DE aggregates; LBPP excluded from code aggregates; GPQA + held-out group withdrawn
    The 'strongest fully open model' ranking is an average over an author-chosen suite; benchmark groups were removed post-hoc (Section 4.3), and small margins (e.g., +0.6 English aggregate over Nemotron) are sensitive to suite composition.
assumptions (4)
  • domain assumption Benchmark scores measured with lm-evaluation-harness under identical prompts and few-shot settings are directly comparable across base models with different training corpora.
    Section 4: every comparative claim rests on this assumption; it ignores differential training-data overlap with benchmarks, which the GPQA incident (Section 4.3) shows is non-zero for this model.
  • ad hoc to paper The post-hoc contamination audit is complete: only GPQA, TruthfulQA, BLiMP, and Inverse Scaling leaked evaluation material into training.
    Section 4.3: self-performed after external community members discovered the GPQA leak; the central benchmark claims are invalid if other evaluation items are in the mixture. The n-gram screening safeguard was added for future runs, not this one.
  • domain assumption The machine-translated German benchmark variants (-DE) are valid, leakage-free measures of German capability.
    Section 4: the German aggregate (85.3) and the +5.9 margin over Apertus rest substantially on HellaSwag-DE, PIQA-DE, ARC-Challenge-DE, and code -DE variants; their construction, translation provenance, and screening against the training corpus are not documented.
  • domain assumption Equation 1's latency-subtraction protocol adequately isolates aggregate decode throughput from prefill cost on a single B200 at batch 32.
    Section 4.4: the 8-9x decode advantage over dense models is computed as B(O_long - O_short)/[t(O_long) - t(O_short)]; validity depends on prefill removal and on TP=1, fixed-256K being representative of production serving.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Sovereign, Open-Source Foundation Model for German and English." pith.science (2026). https://pith.science/paper/UHY7DV6U

@misc{pith2026260709424,
  author       = {Pith},
  title        = {Pith review of: A Sovereign, Open-Source Foundation Model for German and English},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UHY7DV6U}},
  note         = {Machine review of arXiv:2607.09424}
}
read the original abstract

We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.

Figures

Figures reproduced from arXiv: 2607.09424 by the authors.

Figure 1
Figure 1. Long-context serving efficiency. Soofi S combines frontier-level capability with the highest measured aggregate long-context decode TPS, and unlike full-attention dense baselines maintains high throughput as context grows. Panel (a) plots Capability Index versus measured aggregate decode TPS/GPU at 40K context and batch 32. The Capability Index averages five benchmark groups, i.e., Code, GSM8K, GPQA-Diamond, English… view at source ↗
Figure 1
Figure 1. English and German aggregate performance versus decode throughput. Soofi S obtains the highest English and German aggregate scores among the fully open base models in this comparison. Soofi S and Nemotron 3 Nano, which share the same hybrid Mamba–MoE architecture, reach the highest measured decode throughput at 4.8k TPS/GPU. The left panel reports the mean of our English benchmark suite. The right panel reports the … view at source ↗
Figure 2
Figure 2. Training dynamics over the full ∼27T-token run. The quantity is plotted against the number of consumed tokens (in trillions); the pretraining-to-annealing transition occurs at ∼20T tokens. The solid line is a rolling median over 1,000 steps and the faint trace is the raw per-step signal. by existing software from the day of release, without bespoke integration work by downstream users. Second, serving efficiency: th… view at source ↗
Figures from the paper (29 more)
Figure 3
Figure 3. Figure 3: Effective-token mixture across the three training phases. A single flow diagram tracing seven data categories (English Web, Academic & Wiki, SFT, Reasoning, Code, Math, and German) from left to right across the phases. Phase 1 (diverse pretraining), 23,051.13B effectiv…
Figure 4
Figure 4. Figure 4: Evaluation overview for the open-source comparison. [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 4
Figure 4. Figure 4: Evaluation overview for the open-source comparison. [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]
Figure 5
Figure 5. Figure 5: Code generation results (pass@1) against large open-source models on English and German [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 5
Figure 5. Figure 5: Code generation results (pass@1) against large open-source models on English and German [PITH_FULL_IMAGE:figures/full_fig_p019_5.png]
Figure 6
Figure 6. Figure 6: Mathematics results against large open-source models on English and German bench [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 6
Figure 6. Figure 6: Mathematics results against large open-source models on English and German bench [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]
Figure 7
Figure 7. Figure 7: Knowledge (left) and reasoning/science (right) benchmarks against large open-source [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 7
Figure 7. Figure 7: Knowledge (left) and reasoning/science (right) benchmarks against large open-source [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]
Figure 8
Figure 8. Figure 8: German benchmark results against large open-source models. Soofi S ranks first on the [PITH_FULL_IMAGE:figures/full_fig_p021_8.png]
Figure 8
Figure 8. Figure 8: German benchmark results against large open-source models. Soofi S ranks first on [PITH_FULL_IMAGE:figures/full_fig_p022_8.png]
Figure 9
Figure 9. Figure 9: Base model evaluation overview for Soofi S against Nemotron (same architecture) and large open-weight models (Qwen, Ministral, and Gemma). Aggregates are the harness-level English and German suite means. Code EN averages HumanEval and MBPP, Code DE averages HumanEval-D…
Figure 9
Figure 9. Figure 9: Base model evaluation overview for Soofi S against Nemotron (same architecture) and large open-weight models (Qwen, Ministral, and Gemma). Aggregates are the harness-level English and German suite means. Code EN averages HumanEval and MBPP, Code DE averages HumanEval-D…
Figure 10
Figure 10. Figure 10: Code generation results (pass@1) against Nemotron 3 Nano and large open-weight models [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]
Figure 10
Figure 10. Figure 10: Code generation results (pass@1) against Nemotron 3 Nano and large open-weight models [PITH_FULL_IMAGE:figures/full_fig_p025_10.png]
Figure 11
Figure 11. Figure 11: Mathematics results against Nemotron 3 Nano and large open-weight models on English [PITH_FULL_IMAGE:figures/full_fig_p025_11.png]
Figure 11
Figure 11. Figure 11: Mathematics results against Nemotron 3 Nano and large open-weight models on English [PITH_FULL_IMAGE:figures/full_fig_p026_11.png]
Figure 12
Figure 12. Figure 12: Knowledge (left) and reasoning/science (right) benchmarks against Nemotron 3 Nano and [PITH_FULL_IMAGE:figures/full_fig_p026_12.png]
Figure 12
Figure 12. Figure 12: Knowledge (left) and reasoning/science (right) benchmarks against Nemotron 3 Nano [PITH_FULL_IMAGE:figures/full_fig_p027_12.png]
Figure 13
Figure 13. Figure 13: German benchmark results against Nemotron 3 Nano and large open-weight models. [PITH_FULL_IMAGE:figures/full_fig_p027_13.png]
Figure 13
Figure 13. Figure 13: German benchmark results against Nemotron 3 Nano and large open-weight models. [PITH_FULL_IMAGE:figures/full_fig_p028_13.png]
Figure 14
Figure 14. Figure 14: Per-benchmark score difference between Soofi S and the architecture-identical [PITH_FULL_IMAGE:figures/full_fig_p028_14.png]
Figure 14
Figure 14. Figure 14: Per-benchmark score difference between Soofi S and the architecture-identical [PITH_FULL_IMAGE:figures/full_fig_p029_14.png]
Figure 15
Figure 15. Figure 15: Temporal distribution of the 193.1M Genios articles [ [PITH_FULL_IMAGE:figures/full_fig_p055_15.png]
Figure 15
Figure 15. Figure 15: Benchmark trajectories across annealing checkpoints. GPQA-Diamond rises gradually [PITH_FULL_IMAGE:figures/full_fig_p030_15.png]
Figure 16
Figure 16. Figure 16: RULER accuracy averaged over the selected subtasks, by input context length (4K–1M). [PITH_FULL_IMAGE:figures/full_fig_p055_16.png]
Figure 16
Figure 16. Figure 16: Long-context decode-throughput scaling. Measured aggregate decode TPS/GPU as a function of input context length under the TP=1, single-B200, batch-32 latency-subtraction protocol. The same measurements also expose the prefill/TTFT side of the hybrid architecture: t(Os…
Figure 17
Figure 17. Figure 17: Temporal distribution of the 193.1M Genios articles [ [PITH_FULL_IMAGE:figures/full_fig_p057_17.png]
Figure 18
Figure 18. Figure 18: RULER accuracy averaged over the selected subtasks, by input context length (4K–1M). [PITH_FULL_IMAGE:figures/full_fig_p057_18.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

103 extracted references · 31 linked inside Pith

  1. [1]

    Gqa: Training generalized multi-query transformer models from multi-head checkpoints

    J.Ainslie, J.Lee-Thorp, M.DeJong, Y.Zemlyanskiy, F.Lebrón, andS.Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895–4901, 2023

  2. [2]

    M. Ali, M. Brack, M. Lübbering, E. Wendt, A. G. Khan, R. Rutmann, A. Jude, M. Kraus, A. A. Weber, F. Stollenwerk, D. Kaczér, F. Mai, L. Flek, R. Sifa, N. Flores-Herr, J. Koehler, P. Schramowski, M. Fromm, and K. Kersting. Judging quality across languages: A multilin- gual approach to pretraining data filtering with language models. In C. Christodoulopoulo...

  3. [3]

    M. Ali, M. Fromm, K. Thellmann, J. Ebert, A. A. Weber, R. Rutmann, C. Jain, M. Lübbering, D. Steinigen, J. Leveling, K. Klug, J. S. Buschhoff, L. Jurkschat, H. Abdelwahab, B. J. Stein, K.-H. Sylla, P. Denisov, N. Brandizzi, Q. Saleem, A. Bhowmick, L. Helmer, C. John, P. O. Suarez, M. Ostendorff, A. Jude, L. Manjunath, S. Weinbach, C. Penke, O. Filatov, F....

  4. [4]

    Almazrouei, H

    E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, M. Debbah, É. Goffinet, D. Hesslow, J. Launay, Q. Malartic, D. Mazzotta, B. Noune, B. Pannier, and G. Penedo. The falcon series of open language models.arXiv preprint arXiv:2311.16867, 2023. 36

  5. [5]

    Apertus, A

    P. Apertus, A. Hernández-Cano, A. Hägele, A. H. Huang, A. Romanou, A.-J. Solergibert, B. Pasztor, B. Messmer, D. Garbaya, E. F. Ďurech, I. Hakimi, J. G. Giraldo, M. Ismayilzada, N. Foroutan, S. Moalla, T. Chen, V. Sabolčec, Y. Xu, M. Aerni, B. AlKhamissi, I. A. Mar- iñas, M. H. Amani, M. Ansaripour, I. Badanin, H. Benoit, E. Boros, N. Browning, F. Bösch, ...

  6. [6]

    Austin, A

    J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al. Program synthesis with large language models.arXiv preprint arXiv:2108.07732, 2021

  7. [7]

    Biderman, H

    S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, A. Skowron, L. Sutawika, and O. Van Der Wal. Pythia: A suite for analyzing large language models across training and scaling. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors,Proceedings of...

  8. [8]

    BLOOM: A 176b-parameter open-access multilingual language model

    BigScience Workshop. BLOOM: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022

Show all 103 references
  1. [9]

    Black, S

    S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. Mc- Donell, J. Phang, M. Pieler, U. S. Prashanth, S. Purohit, L. Reynolds, J. Tow, B. Wang, and S. Weinbach. GPT-NeoX-20B: An open-source autoregressive language model. In A. Fan, S. Ilic, ...

  2. [10]

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S.Gray, B.Ch...

  3. [11]

    Burchell, O

    L. Burchell, O. de Gibert, N. Arefyev, M. Aulamo, M. Bañón, P. Chen, M. Fedorova, L. Guillou, B. Haddow, J. Hajič, J. Helcl, E. Henriksson, M. Klimaszewski, V. Komulainen, A. Kutuzov, J. Kytöniemi, V. Laippala, P. Mæhlum, B. Malik, F. Mehryary, V. Mikhailov, N. Moghe, A. Myntt...

  4. [12]

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021

  5. [13]

    Clark, I

    P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018

  6. [14]

    Cobbe, V

    K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021

  7. [15]

    D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu, et al. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1...

  8. [16]

    Dao and A

    T. Dao and A. Gu. Transformers are ssms: Generalized models and efficient algorithms through structured state space duality.arXiv preprint arXiv:2405.21060, 2024

  9. [17]

    Deutsche Telekom AG. Germany’s first AI factory for industry officially goes into operation in Munich.https://www.telekom.com/en/media/media-information/archive/germany-s-f irst-ai-factory-for-industry-1101670, Feb. 2026. Industrial AI Cloud, Munich; accessed 2026-07-08

  10. [18]

    S. Diao, Y. Yang, Y. Fu, X. Dong, D. Su, M. Kliegl, Z. Chen, P. Belcak, Y. Suhara, H. Yin, et al. Nemotron-climb: Clustering-based iterative data mixture bootstrapping for language model pre-training.Advances in Neural Information Processing Systems, 38, 2026

  11. [19]

    Fedus, B

    W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022

  12. [20]

    Fujii, Y

    K. Fujii, Y. Tajima, S. Mizuki, M. Kawamura, H. Shimada, T. Shiotani, K. Saito, M. Oi, T. Nakamura, T. Okamoto, S. Ishida, K. Hattori, Y. Ma, H. Takamura, R. Yokota, J. Sakuma, and N. Okazaki. Rewriting pre-training data boosts llm performance in math and code.arXiv preprint a...

  13. [21]

    L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy. The pile: An 800gb dataset of diverse text for language modeling.arXiv preprint arXiv:2101.00027, 2020

  14. [22]

    L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou. The language...

  15. [23]

    GENIOS – german business and press database.https://www.genios.de/browse/Alle, 2025

    GBI-Genios Deutsche Wirtschaftsdatenbank GmbH. GENIOS – german business and press database.https://www.genios.de/browse/Alle, 2025. Commercially licensed corpus of 38 German newspaper and trade-press archives; obtained under a license that does not permit redistribution

  16. [24]

    Gienapp, C

    L. Gienapp, C. Schröder, S. Schweter, C. Akiki, F. Schlatt, A. Zimmermann, P. Genêt, and M. Potthast. The german commons - 154 billion tokens of openly licensed text for german language models.arXiv preprint arXiv:2510.13996, 2025

  17. [25]

    Reasoning-v1-20m, 2025

    GlaiveAI. Reasoning-v1-20m, 2025. A synthetic reasoning dataset containing 22mil+ general reasoning questions and responses generated using deepseek-ai/DeepSeek-R1-Distill-Llama- 70B

  18. [26]

    Gonzalez-Agirre, M

    A. Gonzalez-Agirre, M. Pàmies, J. Llop, I. Baucells, S. D. Dalt, D. Tamayo, J. J. Saiz, F. Es- puña, J. Prats, J. Aula-Blasco, M. Mina, I. Pikabea, A. Rubio, A. Shvets, A. Sallés, I. La- cunza, J. Palomar, J. Falcão, L. Tormo, L. Vasquez-Reina, M. Marimon, O. Pareras, V. Ruiz-...

  19. [27]

    Groeneveld, I

    D. Groeneveld, I. Beltagy, E. Walsh, A. Bhagia, R. Kinney, O. Tafjord, A. Jha, H. Ivison, I. Magnusson, Y. Wang, S. Arora, D. Atkinson, R. Authur, K. Chandu, A. Cohan, J. Dumas, Y. Elazar, Y. Gu, J. Hessel, T. Khot, W. Merrill, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. ...

  20. [28]

    Gurgurov and T

    D. Gurgurov and T. Röhr. ReasonXL: A multilingual cross-domain reasoning corpus.https: //huggingface.co/datasets/toroe/Soofi-Think-SFT-10B-multilingual, 2026

  21. [29]

    Gurgurov, T

    D. Gurgurov, T. Röhr, and S. Ostermann. Nemotron-multilingual-reasoning: A multilingual science reasoning dataset.https://huggingface.co/datasets/DGurgurov/Nemotron-Multi lingual-Reasoning, 2025. Derived from nvidia/Llama-Nemotron-Post-Training-Dataset with machine-translated ...

  22. [30]

    Hägele, E

    A. Hägele, E. Bakouch, A. Kosson, L. B. Allal, L. Von Werra, and M. Jaggi. Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024

  23. [31]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding, 2021

  24. [32]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Stein- hardt. Measuring mathematical problem solving with the math dataset. In J. Vanschoren and S. Yeung, editors,Proceedings of the Neural Information Processing Systems Track on Datasets and ...

  25. [33]

    Hoffmann, S

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. Rae, and L. Sifre. ...

  26. [34]

    Hsieh, S

    C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, Y. Zhang, and B. Ginsburg. Ruler: What’s the real context size of your long-context language models?, 2024

  27. [35]

    S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, et al. Minicpm: Unveiling the potential of small language models with scalable training strategies. arXiv preprint arXiv:2404.06395, 2024

  28. [36]

    Idahl, B

    M. Idahl, B. Droste, B. Plüster, and J. P. Harries. propella-1: Multi-property document annotation for llm data curation at scale, 2026

  29. [37]

    Idahl, J

    M. Idahl, J. Tiedemann, S. Pyysalo, D. Salinas, T. Galica, S. Qian, T. N. Mateiu, Z. Li, A. Lokrantz, F. Vitiugin, A. F. T. Martins, J. Kanerva, F. Ginter, M. Lindemann, T. Isbis- ter, B. Moell, J. Lindh, J. Hajič, J. Jitsev, A. Kutuzov, S. Oepen, and G. Ramírez-Sánchez. Multi...

  30. [38]

    A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bres- sand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. Le Scao, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed. Mistral 7b.arXiv preprint arXiv:2310.06825, 2023

  31. [39]

    A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. Bou Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. Le Scao, T. Gervet, ...

  32. [40]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Rad- ford, J. Wu, and D. Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361, 2020

  33. [41]

    Krajewski, J

    J. Krajewski, J. Ludziejewski, K. Adamczewski, M. Pióro, M. Krutul, S. Antoniak, K. Ciebiera, K. Król, T. Odrzygóźdź, P. Sankowski, et al. Scaling laws for fine-grained mixture of experts. arXiv preprint arXiv:2402.07871, 2024

  34. [42]

    Kraus, R

    M. Kraus, R. Härle, S. Sztwiertnia, A. G. Khan, M. Ali, M. Fromm, and K. Kersting. Kletter- mix: Climbing toward high-quality german pretraining data.arXiv preprint arXiv:2606.03773, 2026

  35. [43]

    Kwiatkowski, J

    T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov. Natural questions: A benchmark for question answering resea...

  36. [44]

    Kydlíček, G

    H. Kydlíček, G. Penedo, and L. von Werra. Finepdfs.https://huggingface.co/datasets/ HuggingFaceFW/finepdfs, 2025

  37. [45]

    Lewkowycz, A

    A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra. Solving quantitative reasoning problems with language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belg...

  38. [46]

    Liesenfeld, M

    A. Liesenfeld, M. Dingemanse, D. Blankvoort, N. Kalra, and A. R. Golkhandan. European open source ai definitions.https://osai-index.eu/osai-definitions, 2026. Accessed: 2026-07-06

  39. [47]

    A. H. Liu, K. Khandelwal, S. Subramanian, V. Jouault, A. Rastogi, A. Sadé, A. Jeffares, A. Jiang, A. Cahill, A. Gavaudan, A. Sablayrolles, A. Héliou, A. You, A. Ehrenberg, A. Lo, A. Eliseev, A. Calvi, A. Sooriyarachchi, B. Bout, B. Rozière, B. D. Monicault, C. Lanfranchi, C. B...

  40. [48]

    Z. Liu, Z. Yang, Y. Chen, C. Lee, M. Shoeybi, B. Catanzaro, and W. Ping. Acereason- nemotron 1.1: Advancing math and code reasoning through sft and rl synergy.arXiv preprint arXiv:2506.13284, 2025

  41. [49]

    Loshchilov and F

    I. Loshchilov and F. Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017

  42. [50]

    Lübbering, T

    M. Lübbering, T. Ruland, R. Rutmann, F. Stollenwerk, D. Fitzek, M. Fromm, A. Weber, R. Sifa, N. Flores-Herr, J. Köhler, et al. Modalities, a pytorch-native framework for large-scale llm training and research.arXiv preprint arXiv:2602.08387, 2026

  43. [51]

    R. K. Mahabadi, S. Satheesh, S. Prabhumoye, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-cc-math: A 133 billion-token-scale high quality math pretraining dataset.arXiv preprint arXiv:2508.15096, 2025

  44. [52]

    P. H. Martins, J. Alves, P. Fernandes, N. M. Guerreiro, R. Rei, A. Farajian, M. Klimaszewski, D. M. Alves, J. Pombal, N. Boizard, M. Faysse, P. Colombo, F. Yvon, B. Haddow, J. G. C. de Souza, A. Birch, and A. F. T. Martins. Eurollm-9b: Technical report.arXiv preprint arXiv:250...

  45. [53]

    P. H. Martins, P. Fernandes, J. Alves, N. M. Guerreiro, R. Rei, D. M. Alves, J. Pombal, A. Farajian, M. Faysse, M. Klimaszewski, P. Colombo, B. Haddow, J. G. C. de Souza, A. Birch, and A. F. T. Martins. Eurollm: Multilingual language models for europe.arXiv preprint arXiv:2409...

  46. [54]

    Matton, T

    A. Matton, T. Sherborne, D. Aumiller, E. Tommasone, M. Alizadeh, J. He, R. Ma, M. Voisin, E. Gilsenan-McMahon, and M. Gallé. On leakage of code generation evaluation datasets. In 41 Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, editors,Findings of the Association for Com- putation...

  47. [55]

    Mistral small 3.1.https://mistral.ai/news/mistral-small-3-1/, 2025

    Mistral AI. Mistral small 3.1.https://mistral.ai/news/mistral-small-3-1/, 2025. Model release, March 2025

  48. [56]

    Morrison, S

    J. Morrison, S. Adhikesaven, A. Bhagia, M. Zaharia, N. A. Smith, and S. Min. Train separately, merge together: Modular post-training with mixture-of-experts, 2026

  49. [57]

    Muennighoff, L

    N. Muennighoff, L. Soldaini, D. Groeneveld, K. Lo, J. Morrison, S. Min, W. Shi, E. P. Walsh, O. Tafjord, N. Lambert, Y. Gu, S. Arora, A. Bhagia, D. Schwenk, D. Wadden, A. Wettig, B. Hui, T. Dettmers, D. Kiela, A. Farhadi, N. A. Smith, P. W. Koh, A. Singh, and H. Ha- jishirzi. ...

  50. [58]

    Mt reasoning: Automatic translations into 2 languages of glaive ai reasoning dataset, 2025

    MultiSynt. Mt reasoning: Automatic translations into 2 languages of glaive ai reasoning dataset, 2025. An automatic translations into 2 languages of Glaive AI reasoning dataset

  51. [59]

    Basant, A

    NVIDIA, :, A. Basant, A. Khairnar, A. Paithankar, A. Khattar, A. Renduchintala, A. Malte, A. Bercovich, A. Hazare, A. Rico, A. Ficek, A. Kondratenko, A. Shaposhnikov, A. Bukharin, A. Taghibakhshi, A. Barton, A. S. Mahabaleshwarkar, A. Shen, A. Tao, A. Guan, A. Shors, A. Mandar...

  52. [60]

    Blakeman, A

    NVIDIA, :, A. Blakeman, A. Basant, A. Khattar, A. Renduchintala, A. Bercovich, A. Ficek, A. Bjorlin, A. Taghibakhshi, A. S. Deshmukh, A. S. Mahabaleshwarkar, A. Tao, A. Shors, A. Aithal, A. Poojary, A. Dattagupta, B. Buddharaju, B. Chen, B. Ginsburg, B. Wang, B. Norick, B. But...

  53. [61]

    Blakeman, A

    NVIDIA, :, A. Blakeman, A. Grattafiori, A. Basant, A. Gupta, A. Khattar, A. Renduchintala, A. Vavre, A. Shukla, A. Bercovich, A. Ficek, A. Shaposhnikov, A. Kondratenko, A. Bukharin, A. Milesi, A. Taghibakhshi, A. Liu, A. Barton, A. S. Mahabaleshwarkar, A. Klein, A. Zuker, A. G...

  54. [62]

    Deutsche telekom and NVIDIA launch industrial AI cloud.https://blogs.nvidia .com/blog/germany-industrial-ai-cloud-launch/, Nov

    NVIDIA. Deutsche telekom and NVIDIA launch industrial AI cloud.https://blogs.nvidia .com/blog/germany-industrial-ai-cloud-launch/, Nov. 2025. Accessed 2026-07-08

  55. [63]

    Oepen, N

    S. Oepen, N. Arefev, M. Aulamo, M. Bañón, M. Buljan, L. Burchell, L. Charpentier, P. Chen, M. Fedorova, O. de Gibert, B. Haddow, J. Hajič, J. Helcl, A. Kutuzov, V. Laippala, Z. Li, R. Luukkonen, B. Malik, V. Mikhailov, A. Myntti, D. O’Brien, L. Poláková, S. Pyysalo, G. R. Sánc...

  56. [64]

    Olmo, :, A

    T. Olmo, :, A. Ettinger, A. Bertsch, B. Kuehl, D. Graham, D. Heineman, D. Groeneveld, F. Brahman, F. Timbers, H. Ivison, J. Morrison, J. Poznanski, K. Lo, L. Soldaini, M. Jordan, M. Chen, M. Noukhovitch, N. Lambert, P. Walsh, P. Dasigi, R. Berry, S. Malik, S. Shah, S. Geng, S....

  57. [65]

    The open source ai definition, version 1.0.https://opensource.org /ai/open-source-ai-definition, 2024

    Open Source Initiative. The open source ai definition, version 1.0.https://opensource.org /ai/open-source-ai-definition, 2024. Accessed: 2026-07-06. 44

  58. [66]

    OpenEuroLLM: A series of foundation models for transparent ai in europe, 2025

    OpenEuroLLM Consortium. OpenEuroLLM: A series of foundation models for transparent ai in europe, 2025. Project website, accessed 2026-06-25

  59. [67]

    G. Penedo. Finewiki, 2025. Source: Wikimedia Enterprise Snapshot API (https://api.enterprise.wikimedia.com/v2/snapshots). Text licensed under CC BY-SA 4.0 with attribution to Wikipedia contributors

  60. [68]

    Penedo, H

    G. Penedo, H. Kydlíček, A. Lozhkov, M. Mitchell, C. Raffel, L. Von Werra, T. Wolf, et al. The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849, 2024

  61. [69]

    SYNTH: An open generalist synthetic dataset for training small reason- ing models.https://huggingface.co/datasets/PleIAs/SYNTH, 2025

    Pleias and AI Alliance. SYNTH: An open generalist synthetic dataset for training small reason- ing models.https://huggingface.co/datasets/PleIAs/SYNTH, 2025. Dataset comprising 79,648,272 text samples (over 41 billion words) derived from the synthetic amplification of 58,698 W...

  62. [70]

    Poznanski, A

    J. Poznanski, A. Rangapur, J. Borchardt, J. Dunkelberger, R. Huff, D. Lin, C. Wilhelm, K. Lo, andL.Soldaini. olmocr: Unlockingtrillionsoftokensinpdfswithvisionlanguagemodels.arXiv preprint arXiv:2502.18443, 2025

  63. [71]

    Qwen3.5: Towards native multimodal agents, February 2026

    Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026

  64. [72]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. Technical report, OpenAI, 2019

  65. [73]

    Raffel, N

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of Machine Learning Research, 21(140):1–67, 2020

  66. [74]

    M. M. Ramos, D. M. Alves, H. Gisserot-Boukhlef, J. Alves, P. H. Martins, P. Fernandes, J. Pombal, N. M. Guerreiro, R. Rei, N. Boizard, A. Farajian, M. Klimaszewski, J. G. C. de Souza, B. Haddow, F. Yvon, P. Colombo, A. Birch, and A. F. T. Martins. Eurollm-22b: Technical report...

  67. [75]

    D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark, 2023

  68. [76]

    Romanou, N

    A. Romanou, N. Foroutan, A. Sotnikova, S. H. Nelaturu, S. Singh, R. Maheshwary, M. Al- tomare, Z. Chen, M. Haggag, S. A, A. Amayuelas, A. H. Amirudin, D. Boiko, M. Chang, J. Chim, G. Cohen, A. K. Dalmia, A. Diress, S. Duwal, D. Dzenhaliou, D. Florez, F. Farestam, J. M. Imperia...

  69. [77]

    Shazeer, A

    N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean. Outra- geously large neural networks: The sparsely-gated mixture-of-experts layer.arXiv preprint arXiv:1701.06538, 2017. 45

  70. [78]

    W. Shi, A. Bhagia, K. Farhat, N. Muennighoff, J. Morrison, E. Walsh, D. Schwenk, S. Long- pre, J. Poznanski, A. Ettinger, et al. Flexolmo: Open language models for flexible data use. Advances in Neural Information Processing Systems, 38:165943–165974, 2026

  71. [79]

    Soldaini, R

    L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkinson, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y. Elazar, V. Hofmann, A. Jha, S. Kumar, L. Lucy, X. Lyu, N. Lambert, I. Magnus- son, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. Peters, A. Ravichander, K. Richardson, Z...

  72. [80]

    D. Su, K. Kong, Y. Lin, J. Jennings, B. Norick, M. Kliegl, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-CC: Transforming Common Crawl into a refined long-horizon pre- training dataset. In W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, editors,Proceedings of the 63rd...

  73. [81]

    Suzgun, N

    M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. Le, E. Chi, D. Zhou, and J. Wei. Challenging BIG-bench tasks and whether chain-of-thought can solve them. In A. Rogers, J. Boyd-Graber, and N. Okazaki, editors,Findings of the Association for ...

  74. [82]

    G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovi- cova, A. Ramé, M. Rivière, L. Rouillard, T. Mesnard, G. Cideron, J. bastien Grill, S. Ramos, E. Yvinec, M. Casbon, E. Pot, I. Penchev, G. Liu, F. Visin, K. Kenealy, L. Beyer, X. Zhai, A....

  75. [83]

    Touvron, T

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample. LLaMA: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971, 2023

  76. [84]

    Touvron, L

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. Canton Ferrer, M. Chen, G. Cucurull, D. Es- iobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S...

  77. [85]

    Vendrow, E

    J. Vendrow, E. Vendrow, S. Beery, and A. Madry. Do large language model benchmarks test reliability?arXiv preprint arXiv:2502.03461, 2025

  78. [86]

    Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, etal. Mmlu-pro: Amorerobustandchallengingmulti-tasklanguageunderstandingbenchmark. Advances in Neural Information Processing Systems, 37:95266–95290, 2024

  79. [87]

    A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. L...

  80. [88]

    Zhang, S

    S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer. OPT: Open pre-trained transformer language models.arXiv preprint arXi...

  81. [89]

    Zhong, R

    W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan. AGIEval: A human-centric benchmark for evaluating foundation models. In K. Duh, H. Gomez, and S. Bethard, editors,Findings of the Association for Computational Linguistics: NAACL 2024, pages 22...

  82. [90]

    C. Zhou, H. Lyu, X. Lin, H. Zhao, J. Guo, X. Zhang, S. Xue, Q. Ma, J. Zhou, Y. Wang, and Z. Liu. Ultradata-math, 2026. 47 A Author Contributions A.1 Training

  83. [91]

    Pretraining stack development and evaluation: Max Lübbering, Richard Rutmann, Timm Ruland, David Fitzek, Mehdi Ali

  84. [92]

    Model architecture, training methodology and framework-correctness validation: Timm Ru- land, David Fitzek, Max Lübbering, Richard Rutmann

  85. [93]

    Compute infrastructure, cluster benchmarking and interconnect tuning: David Fitzek, Timm Ruland, Richard Rutmann, Max Lübbering

  86. [94]

    Distributed-training scaling and memory/throughput optimization: David Fitzek, Timm Ru- land, Max Lübbering, Richard Rutmann

  87. [95]

    Training execution, stability analysis, emergency debugging, framework bug fixes and experi- ment tracking: Timm Ruland, David Fitzek, Max Lübbering, Richard Rutmann A.2 Data

  88. [96]

    Base model data acquisition: Michael Fromm, Alex Jude, Abbas Khan, Ruben Härle, Maurice Kraus, Jan Pfister, Daniil Gurgurov

  89. [97]

    Pretraining data mixture: Michael Fromm

  90. [98]

    Data curation infrastructure and experimentation: Michael Fromm, Alex Jude, Abbas Khan, Ruben Härle, Maurice Kraus, Richard Rutmann, Mehdi Ali, Max Lübbering, Maximilian Idahl

  91. [99]

    Data preprocessing / tokenization pipeline: Richard Rutmann, Max Lübbering, Alex Jude

  92. [100]

    Mid- and long-context data curation and experimentation: Michael Fromm, Alex Jude, Ab- bas Khan, Ruben Härle, Maurice Kraus, Sebastian Sztwiertnia, Tom Röhr, Sebastian von Rohrscheidt A.3 Evaluation

  93. [101]

    Evaluation methodology and infrastructure: Maximilian Idahl, Benedikt Droste, Alex Jude, Abbas Khan A.4 Other

  94. [102]

    Mentorship, advising, program management, and broader strategy: Nicolas Flores-Herr, Si- mon Gottschalk, Jörg Bienert, Kristian Kersting, Andreas Hotho, Alexander Löser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Björn Plüster, Patrick Putzky

  95. [103]

    Raw” is the source token count in billions; “Ep

    Technical leadership and cross-workstream contributions: Mehdi Ali, Michael Fromm, Max Lübbering, Sebastian Sztwiertnia, Tom Röhr 48 B Detailed Pretraining Data Composition This appendix gives the full per-source token accounting underlying the mixture flow diagram in Section ...

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.