{"id":"25faefcf-3567-4977-9d5c-dd6e673cd73f","arxiv_id":"2608.13517","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"Mimir v1 is a 1B-parameter model trained only on permissible data that reports strong Danish benchmark results, but likely training/evaluation overlap undermines the state-of-the-art claim.","lead":"This paper releases Mimir v1, a 1 billion parameter language model trained from scratch on 161 permissible datasets. It reports competitive English and math scores and a new state of the art for Danish, but several Danish evaluation benchmarks appear in the training data, which weakens the headline claim.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Train/eval overlap is unaddressed: Table 10 includes ccdv/govreport-summarization, giannor/*dala* variants, and oliverkinch/multi-wiki-qa-high-quality-subset, which are exact or near counterparts of Table 11 benchmarks; without a decontamination check the reported SOTA scores may reflect…","rationale":"The reader and I identify the same load-bearing premise: evaluation benchmarks must be disjoint from training data. I would keep the REJECT verdict because the paper as written does not rule out contamination. I flag this as partial rather than full agreement because the exact dataset identifiers do not always match: training uses 'tv2r it' variants and a 'high-quality-subset' while evaluation uses 'gen v3' and the full Multi Wiki QA set, so the overlap has to be measured, not assumed. The concern extends beyond Danish: ccdv/govreport-summarization is an exact source match for the English GovReport benchmark, which makes the contamination risk broader than the reader's statement. The model release, the transparent dataset list, and the synthetic transplant effort are real contributions, and if the overlap test comes back clean the comparisons would be impressive. But the absence of any decontamination statement, combined with direct source-name matches in both English and Danish tables, leaves the central claim unsupported. Secondary issues such as mixed evaluation harnesses and missing error bars reinforce caution but are not the primary reason for rejection.","tokens_in":15184,"tokens_out":8810,"duration_ms":84417,"concrete_test":"For each apparent overlap (ccdv/govreport-summarization vs GovReport, oliverkinch/multi-wiki-qa-high-quality-subset vs Multi Wiki QA, and giannor/gec dala tv2r it plus giannor/dala tv2r it vs giannor/dala plus giannor/dala gen v3), download the exact training rows actually sampled from Table 10 and the evaluation rows from Table 11, normalize case, whitespace, and Unicode, and compute exact-match and near-duplicate (MinHash, Jaccard >= 0.8) overlap between prompt and target pairs. If any overlap is nonzero, recompute the affected benchmark scores after excluding those training rows; if the Danish or GovReport margins collapse, the central generalization claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—frontier-level performance at 1B and a new Danish state of the art—rests entirely on Tables 7, 8, and 9, and those tables are only evidence of generalization if the evaluation sets in Table 11 were excluded from the 161 training datasets in Table 10. The paper never states or documents such an exclusion. Several entries make the assumption insecure. The clearest exact case is English: training dataset #91, ccdv/govreport-summarization, is the same source as the GovReport evaluation set (ccdv/govreport). For Danish, training includes #66 oliverkinch/multi-wiki-qa-high-quality-subset, a curated subset of the Multi Wiki QA evaluation set (oliverkinch/multi-wiki-qa), and #43 giannor/gec dala tv2r it and #61 giannor/dala tv2r it, which are DaLA/GEC-DaLA variants of the evaluation sets giannor/dala and giannor/dala gen v3. The exact splits may differ ('tv2r it' versus 'gen v3', 'high-quality-subset' versus full), but the paper provides no decontamination analysis, and the token volumes involved (193M, 68.5M, 41.9M, 4.4M) are more than enough to memorize small evaluation benchmarks. Because the Danish SOTA claim is the headline result and depends on DaLA, GEC-DaLA, and WikiQA, this is the load-bearing weakness: if overlap is confirmed, the reported averages are memorization, not generalization.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Mimir v1, a 1-billion-parameter model based on the HRM-Text architecture, trained from scratch on a mixture of 161 datasets totaling about 70.5B tokens per epoch, all claimed to be permissible post-training data. The authors report benchmark results for English, Math & Code, and Danish, claiming that Mimir outperforms the original HRM-Text 1B, competes with larger models such as Qwen 3.5 4B and Gemma 4 E2B, and sets a new state of the art for Danish. The appendices list the complete training corpus and evaluation datasets, and the model is released on the Hugging Face Hub.","tokens_in":15525,"tokens_out":5901,"duration_ms":55168,"significance":"If the reported results are valid, the contribution is significant: it would demonstrate that a 1B model trained entirely on permissible data can be competitive with models several times its size, and it would provide a useful open recipe for low-resource language modeling. The paper is transparent about the training corpus, hyperparameters, and model release, and the synthetic transplant idea is interesting. However, the central evidence consists entirely of benchmark tables, and those tables are undermined by the apparent overlap between training datasets and evaluation datasets. Because the Danish state-of-the-art claim rests on benchmarks whose exact or near-identical sources appear in the training list, the results cannot currently be interpreted as evidence of generalization. The absence of error bars or decontamination analysis further weakens the headline claims.","major_comments":[{"comment":"The training corpus in Table 10 contains entries whose Hugging Face identifiers match or are near-identical to evaluation datasets in Table 11. Specifically, row #43 (giannor/gec dala tv2r it) and row #61 (giannor/dala tv2r it) correspond to the GEC-DaLA (giannor/dala gen v3) and DaLA (giannor/dala) evaluation sets; row #66 (oliverkinch/multi-wiki-qa-high-quality-subset) is a subset of the Multi Wiki QA evaluation set (oliverkinch/multi-wiki-qa); and row #91 (ccdv/govreport-summarization) is a summarization variant of the GovReport evaluation source (ccdv/govreport). The manuscript never states that these evaluation splits were excluded from training, and it provides no decontamination analysis. Because the Danish state-of-the-art claim depends on DaLA, GEC-DaLA, and Multi Wiki QA scores, the reported numbers may reflect memorization rather than generalization. The central claim of the paper is therefore unsupported as written; the authors must either demonstrate that these training entries are disjoint from the evaluation benchmarks or retrain without them.","section":"§2 Table 10; §5 Table 11"},{"comment":"No error bars, confidence intervals, or significance tests are reported for any benchmark. Several headline comparisons are close enough that sampling noise could change the ranking; for example, the English average is 69.0 for Mimir versus 69.3 for Qwen 3.5 4B, and on MATH Mimir scores 45.8 while HRM-Text scores 56.0. All numbers come from single greedy decoding runs. Without uncertainty quantification or repeated-seed evaluation, the claims to 'outperform' competitors on individual tasks and to be 'close to the best' on others are not statistically established.","section":"§5 Evaluation Setup"}],"minor_comments":[{"comment":"The author list gives 'Lukas Galke Poech' while the reference to Dynaword gives 'Lukas Galke'; the spelling should be reconciled.","section":"Author list and References"},{"comment":"The phrase 'fitting 4 contexts of length 4096 each' is grammatically incomplete; the intended meaning is that the per-accelerator batch comprises 4 sequences of length 4096.","section":"§4 Training"},{"comment":"Section 2.1 states that the Sapient repository 'bundles 107 sub-collections', while Table 1 reports 71 datasets in the Sapient mixed category; the relationship between these counts should be clarified.","section":"§2.1 and Table 1"},{"comment":"The evaluation protocol differs between Mimir, evaluated with Hugging Face Transformers, and the baselines, evaluated with vLLM via Inspect AI; the authors state the results are comparable, but no side-by-side verification of the exact tables is provided.","section":"§5 Evaluation Setup"},{"comment":"The claim that all training data are 'permissible' is asserted at the aggregate level, but the paper does not provide per-dataset license information or a documented permissibility audit for the 161 datasets.","section":"§2 Datasets"}],"recommendation":"reject","confidential_remarks":"The apparent train/eval overlap is the decisive issue. If the authors can later provide a rigorous decontamination analysis demonstrating that the listed training sources are disjoint from the evaluation sets, the paper might be reconsidered as a technical report, but the current version's headline claims are not credible. The paper may be more suitable for a venue that emphasizes reproducibility over broad significance claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nMimir v1 is a legitimately useful artifact: a 1B HRM model trained from scratch on a carefully documented 161-dataset, 70.5B-token mixture of permissible data, with the model and training details released. The synthetic transplant idea — regenerating non-permissible Sapient subsets under audit — is the most original piece, and the transparent appendix listing every dataset with token counts is exactly what reproducibility looks like. The English and Math results are plausible and partly encouraging, especially the HumanEval improvement over HRM-Text.\n\nThat said, the reader and the stress-test are right that the Danish state-of-the-art claim is load-bearing and unsecured. Table 10 includes giannor/gec_dala_tv2r_it, giannor/dala_tv2r_it, oliverkinch/multi-wiki-qa-high-quality-subset, and ccdv/govreport-summarization, which are the same sources as the DaLA, GEC-DaLA, Multi Wiki QA, and GovReport eval sets in Table 11. The splits are not identical ('tv2r_it' vs 'gen_v3', 'high-quality-subset' vs full), but the paper gives no decontamination analysis and no explicit statement that these overlaps were excluded. The token volumes (68M–193M for the Danish ones) are far more than enough to memorize small benchmarks. Without that check, the Danish averages in Table 9 are consistent with memorization rather than generalization.\n\nThe evaluation setup also has secondary weaknesses: Mimir is served with HF Transformers while baselines use vLLM+Inspect AI, and there are no error bars or significance tests. Those are minor if the contamination issue is fixed.\n\nThis is not a fatal or sloppy paper — the full dataset table is better reporting than most — but the central claim is exactly where the evidence is weakest. The fix is a proper contamination analysis (n-gram overlap with eval sets, or retraining eval-excluded variants) plus clear statements of which splits were held out.\n\nWho gets value from this paper: anyone working on low-resource or permissive-data LM training, and people building open Danish models. It deserves a serious referee, and I would accept it for review with the clear expectation that the decontamination analysis determines the outcome. If overlap is confirmed, the Danish SOTA claim goes away; if not, the paper becomes a solid contribution.","headline":"Mimir v1 is a real open-data model with careful documentation, but the Danish SOTA claim is unsecured by apparent train/eval overlap and needs a decontamination analysis before it can be trusted.","tokens_in":16074,"tokens_out":2305,"would_cite":false,"duration_ms":22346,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Mimir v1, a 1-billion-parameter model trained from scratch on 161 permissible post-training datasets, outperforms its 1B predecessor and matches or beats models two to five times larger, setting a new Danish state of the art.","keywords":["large language models","HRM architecture","permissible data","Danish benchmarks","instruction tuning","synthetic data","low-resource languages","from-scratch training"],"falsifier":"Compare the evaluation splits of giannor/dala, giannor/gec dala tv2r it, and oliverkinch/multi-wiki-qa against the training versions of those datasets; if any evaluation query or target appears in the training data, the reported Danish scores are inflated by memorization and the state-of-the-art claim fails for those benchmarks.","tokens_in":14985,"feed_emoji":"🇩🇰","tokens_out":11391,"duration_ms":98203,"temperature":0.7,"pith_summary":"Mimir v1 is a 1-billion-parameter language model trained from scratch using only permissible post-training data: openly licensed sources, agreement-supplied data, and EU text-and-data-mining-excepted material. The paper claims this model outperforms the original HRM-Text 1B, competes with larger models such as Qwen 3.5 4B and Gemma 4 E2B across 20 English, math, code, and Danish benchmarks, and sets a new state of the art for Danish. If the claim holds, small research groups and low-resource language communities can build competitive models without massive, often non-permissible pretraining corpora.","feed_headline":"1B model on permissible data beats 4B rivals","feed_subtitle":"Mimir v1 also sets a Danish benchmark record; a small team trained it in under three weeks.","key_machinery":"The mechanism is the HRM-Text architecture: a transformer with 32 layers, hidden size 1536, 12 attention heads per layer, and hierarchical reasoning configured as 2 H-cycles and 3 L-cycles, trained with truncated backpropagation over five steps. This structure is paired with a data recipe built from 161 permissible datasets, including synthetic 'transplant' datasets generated with a larger teacher model and audited before inclusion. The H/L-cycle design is what lets the model learn reasoning behavior from instruction and post-training data without a conventional massive pretraining corpus, and the synthetic transplants are what let the authors replace non-permissible components of the original HRM-Text corpus while keeping the task mix intact.","core_discovery":"The central claim is that the Hierarchical Reasoning Model (HRM) architecture makes it possible to train a competitive 1B model from scratch on post-training data alone, provided the data is diverse and permissible. Mimir v1 is trained on 161 datasets totaling 70.48 billion tokens per epoch, with a corpus that is 68.6% English and 24.7% Danish, and it uses synthetic 'transplant' datasets to replace the non-permissible portions of the original HRM-Text data. In the paper's evaluations, Mimir v1 exceeds HRM-Text 1B on nearly every benchmark, leads the 1B weight class on BoolQ, Winogrande, DROP, GSM8K, and HumanEval, trails Qwen 3.5 4B by only 0.3 points on the English average, and records the best Danish scores on DaLA, GEC, WikiQA, and the overall Danish suite.","pith_inferences":["Re-running the Danish suite on evaluation splits provably disjoint from the training corpus would settle whether the reported state-of-the-art scores reflect generalization rather than memorization, since three Danish evaluation sets share identifiers with training datasets.","Because 68.5% of the training tokens are English, the Danish-first performance may rely on transfer from English; a Danish-only or Danish-heavy training ablation would reveal how much of the Danish score comes from the 24.7% Danish share.","The model adopts another family's tokenizer and chat template, so the architecture's unique contribution is not isolated; an ablation with a different tokenizer would clarify what HRM's H/L-cycle design adds on its own.","Extending the same data recipe to 2-4B parameters is a natural next test; the 1B results suggest such a model would surpass Qwen 3.5 4B on the English average and widen the Danish lead."],"forward_implications":["Because Mimir v1 is trained only on permissible data and released openly, it provides a licensing-clean 1B base model for Danish and English applications.","The reported improvement over HRM-Text 1B is large, including a 36.7% gain on the Math & Code average, showing that the HRM recipe itself, not the original data, drives much of the performance.","The synthetic transplant strategy gives data teams a legal path to recreate proprietary-style training mixes from sources that are openly licensed or audited.","With about 1.65 million training steps on eight accelerators in under three weeks, the recipe is feasible for national and university-scale projects rather than only large industrial labs."],"supporting_citations":[{"why":"Supplies the HRM-Text architecture and the original post-training recipe that Mimir v1 extends with permissible data.","marker":"Wang et al. [2026]"},{"why":"Defines the Danish Foundation Models project's permissible-data philosophy that motivates the corpus constraints.","marker":"Enevoldsen et al. [2023]"},{"why":"Provides the Common Pile public-domain English corpus used to generate synthetic training tasks.","marker":"Kandpal et al. [2025]"},{"why":"Provides Danish Dynaword text used for synthetic Danish training data.","marker":"Enevoldsen et al. [2025]"},{"why":"Supplies OpenMathInstruct-2, the largest math-and-reasoning training source in the corpus.","marker":"Toshniwal et al. [2024]"},{"why":"Supplies Tulu 3 SFT mixtures used for English instruction following.","marker":"Lambert et al. [2024]"},{"why":"Supplies Dolci instruction data, the largest English instruction contribution after the Sapient corpus.","marker":"Team Olmo [2025]"},{"why":"Supplies AceReason-1.1-SFT reasoning traces used in the math and reasoning category.","marker":"Liu et al. [2025]"}],"fun_headline_variants":["Mimir v1: 1B params, only permissible data, tops 4B","Open 1B model, permissible data only, beats 4B rivals","1B param model sets Danish record, beats 4B rivals","Permissible-data 1B model rivals 4B, wins Danish","1B from scratch, permissible data, outdoes 4B"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The Danish state-of-the-art claim presumes that the evaluation benchmarks measure generalization rather than memorization, which requires that their test items were excluded from the training corpus; the paper does not state that the overlapping DaLA, GEC-DaLA, and Multi Wiki QA datasets were held out.","fun_headline_variants_meta":{"raw":{"variants":["Mimir v1: 1B params, only permissible data, tops 4B","Open 1B model, permissible data only, beats 4B rivals","1B param model sets Danish record, beats 4B rivals","Permissible-data 1B model rivals 4B, wins Danish","1B from scratch, permissible data, outdoes 4B"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000532,"raw_usage":{"total_tokens":2542,"prompt_tokens":908,"completion_tokens":1634,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":1534}},"tokens_in":524,"tokens_out":1634,"duration_ms":12029,"temperature":1.0,"reasoning_tokens":1534,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:13:02.610437+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the evaluation splits of giannor/dala, giannor/gec dala tv2r it, and oliverkinch/multi-wiki-qa against the training versions of those datasets; if any evaluation query or target appears in the training data, the reported Danish scores are inflated by memorization and the state-of-the-art claim fails for those benchmarks.","supporting_citations":[],"review_version":1}