{"id":"1d317589-e729-4da2-a9c1-d6cec1ecc2e2","arxiv_id":"2608.13545","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A K-5-only pretraining corpus and 5B model show that language model capabilities track the knowledge boundary of the training data, and standard post-training methods do not cross it.","lead":"The authors built an 88-billion-token training corpus containing only U.S. elementary school material (kindergarten through grade 5) and trained a 5-billion-parameter language model on it. The model handles grade-school questions well but fails on more advanced ones, and scaling, fine-tuning, or in-context examples did not push it beyond that boundary.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post-hoc MathCAMPS filtering (C.2.1), not the pretraining boundary, may drive the Beyond-K-5 ceiling: re-including 6.EE.B.7 and low-entropy items could shrink the gap.","rationale":"Good-faith summary: the paper builds a genuinely useful resource—a precision-filtered K-5 corpus and matched UNFILTERED control—and the qualitative gap in Table 1 and Figures 5-6 supports a real boundary. The strongest claim, however, is the ceiling statement in Section 4.4, and its evidence is dominated by MathCAMPS. The reader identified controlled-exposure leakage as the weakest assumption; I find the post-hoc evaluation filter more load-bearing because it directly determines the outcome of every intervention experiment. The authors document the filters transparently and give content-validity justifications, so this is not an accusation of cherry-picking; it is an identification of a confound that should be tested. If the re-run including 6.EE.B.7 and low-entropy items still shows no Beyond-K-5 gains, the central claim holds and the paper is strong. If not, the abstract and Section 4.4 overreach. The paper itself qualifies the ICL result (Section 6: emergent behaviors may be less pronounced at smaller scale), which supports narrowing the claim regardless. The controlled-exposure premise is well supported by multiple independent checks, so I do not elevate it above the evaluation-filter concern. Verdict remains conditional: addressable, not fundamental.","tokens_in":27780,"tokens_out":6819,"duration_ms":71791,"concrete_test":"Re-run the scaling (Figure 7), post-training (Figure 8), and ICL (Figure 9) experiments on the unfiltered MathCAMPS Beyond-K-5 split—i.e., restore 6.EE.B.7, the low-entropy standards, and the verbatim-answer questions—and report pass@1 and pass@1024 per grade for LITTLELEARNER and UNFILTERED. If LITTLELEARNER's Beyond-K-5 accuracy rises toward UNFILTERED, or if SFT+GRPO or few-shot ICL now show gains on these restored items, then the C.2.1 filters are load-bearing and the ceiling claim must be restricted to the filtered subset. As a secondary check, report the grade-6 result with and without 6.EE.B.7 alone.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that scaling, SFT+GRPO, and ICL cannot move LITTLELEARNER beyond K-5—is tested almost entirely on MathCAMPS, whose Beyond-K-5 split is constructed by the post-hoc filters of Section C.2.1. Those filters remove standards with fewer than 30 unique gold answers, questions with the gold string in the prompt, and the entire grade-6 standard 6.EE.B.7. The 6.EE.B.7 exclusion is explicitly motivated by the observation that its MathCAMPS items are structurally identical to grade-3 add/sub word problems; including it would 'double-count grade-3 add/sub competence in the grade-6 band.' This is exactly the kind of crossover item on which a K-5-trained model with post-training or ICL might plausibly improve. By purging the easiest Beyond-K-5 items, the evaluation filter may be responsible for a substantial part of the apparent ceiling. The paper's other boundary evidence (Jeopardy, CoMTA, CLEAR) is qualitative or correlational and does not test the three interventions. Section 4.4's assertion that 'the pretraining filter, rather than the intervention, sets the effective capability ceiling' is therefore not independent of the evaluation filter.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LITTLECURRICULUM, an 88B-token pretraining corpus derived from FineWeb-Edu and filtered to U.S. elementary-school (K-5) content via a multi-stage pipeline (Age-of-Acquisition pre-filtering, LLM-as-a-judge annotation with FastText/ModernBERT classifiers, symbolic filtering, and frequency sampling), together with LITTLELEARNER, a 5B-parameter model trained from scratch on this corpus. The filter is validated on CommonCoreText (0% Beyond-K-5 retention), on the external WeeBit corpus (2.48% retention), and via a 126-term n-gram audit of the retained corpus (0.09% of passages). LITTLELEARNER performs comparably to an unfiltered control within K-5 across language-complexity (CLEAR), math-familiarity (CoMTA), fact-retrieval (Jeopardy), and math-reasoning (MathCAMPS) measures, while degrading on Beyond-K-5 content. Three intervention studies—model scaling from 0.6B to 5B, SFT+GRPO post-training, and few-shot in-context learning—improve in-scope performance but do not close the Beyond-K-5 gap, and the paper concludes in Section 4.4 that the pretraining filter, rather than any tested intervention, sets the effective capability ceiling. The authors release the corpus and model as a developmentally restricted sandbox for controlled-exposure research.","tokens_in":28039,"tokens_out":15184,"duration_ms":146583,"significance":"The work is significant as a community resource: if the K-5 boundary is as clean as claimed, LITTLECURRICULUM and LITTLELEARNER enable controlled studies of knowledge acquisition, transfer, post-training, and in-context learning that are otherwise confounded by unknown pretraining exposure. Strengths include the release of corpus and model, a precision-first filter with staged validation, an external WeeBit check and a corpus-level n-gram audit, matched baselines for the main comparisons, honest disclosure of boundary fuzziness (e.g., the non-monotonic skill orderings in Section C.2.2), and careful pass@k analysis showing that the Beyond-K-5 gap persists at high sampling budgets. The headline negative result—that scaling, SFT+GRPO, and ICL amplify in-scope capability but do not extend the boundary—is interesting, falsifiable, and useful even if later work qualifies it. The main risk is that the intervention results rest on a single benchmark whose Beyond-K-5 split is partly constructed post hoc in Section C.2.1.","major_comments":[{"comment":"The central claim of Section 4.4—that 'the pretraining filter, rather than the intervention, sets the effective capability ceiling'—is evaluated on MathCAMPS, but the Beyond-K-5 split used in Figures 7-9 is constructed by the post-hoc filters of Section C.2.1. In particular, standard 6.EE.B.7 is excluded explicitly because its items are 'structurally indistinguishable from grade-3 add/sub word problems,' which is precisely the class of crossover items where a K-5-pretrained model with post-training or ICL might most plausibly show out-of-scope gains, and standards with fewer than 30 unique gold answers (including perfect squares, cubes, and cube roots, which are grade-8 content) are dropped. Excluding these items from the aggregate Beyond-K-5 score makes the ceiling look harder than the pretraining filter alone would produce, so the attribution in Section 4.4 is not separable from the evaluation filter. The authors should report the aggregate Beyond-K-5 numbers for the scaling, post-training, and ICL experiments with 6.EE.B.7 re-included and, ideally, with the low-entropy standards re-included, as a sensitivity analysis.","section":"§4.4 and §C.2.1"},{"comment":"The three intervention claims are tested on a single benchmark: MathCAMPS. The Jeopardy, CLEAR, and CoMTA experiments characterize the base model's boundary but do not test scaling, post-training, or in-context learning, so the claim that 'each lever ... provides limited gains in Beyond-K-5' is supported by only one evaluation instrument. I ask the authors to run at least one additional Beyond-K-5 evaluation (for example, the Jeopardy or CoMTA protocol) for at least the post-training condition, or, failing that, to narrow the summary and abstract claims to the MathCAMPS setting. This is not a request for exhaustive evaluation, but the single-benchmark basis is thin for the paper's headline negative result.","section":"§4, Figures 7-9"},{"comment":"The UNFILTERED control and the scaling runs are said to share LITTLELEARNER's training recipe, but the token budgets, number of epochs, and total compute for each model are not reported. If UNFILTERED, or the 0.6B and 1.3B models, saw different amounts of data than LITTLELEARNER, the divergences in Figures 3-8 cannot be attributed solely to the K-5 filter, because data quantity is a confound. The paper should state the exact token counts, step budgets, and data mixtures for every model in the comparisons.","section":"§3.3, §B, Figure 7"},{"comment":"CommonCoreText is used during pipeline construction—for rule-based metric selection in Section A.1 and for the LLMJ 'precision-on-validation' prompt optimization in Section A.2—so the 0% Beyond-K-5 retention on CommonCoreText is a development-set result rather than an independent validation; the Figure 2 caption discloses this role, but the main text of Section 3.1.6 presents it as validation. The genuinely independent checks (WeeBit at 2.48% retention with 0.05% genuinely out-of-scope after manual inspection, and the 126-n-gram corpus scan at 0.09%) are reassuring but narrow relative to the claim that the 88B-token corpus respects the K-5 boundary. I recommend an additional corpus-level audit, such as a broader grade-6+ vocabulary or classifier-based scan of the released corpus, or an explicit statement of what the 126-n-gram audit can and cannot detect.","section":"§3.1.6, §A.5, §A.6"}],"minor_comments":[{"comment":"The sentence 'This divergence can be attributed to pretraining data, since UNFILTERED shares LITTLELEARNER's training recipe but Gemma 2B's BPB' is garbled and should be rephrased.","section":"§3.3.1"},{"comment":"The text claims 'no difference between LITTLELEARNER post-trained on K-5 data versus Beyond-K-5 data' in Beyond-K-5 performance, but Figure 8's legend ('Pretrain / Post-training K-5 / Unfiltered') does not clearly show the K-5-versus-unfiltered GRPO ablation; the figure should label each condition explicitly so the claimed null result is visible.","section":"§4.2 and Figure 8"},{"comment":"The paper does not state how many samples were LLMJ-annotated or the accuracy of the trained FastText and ModernBERT classifiers on held-out LLMJ labels; these numbers are needed to assess classifier quality and to reproduce the pipeline.","section":"§3.1.2"},{"comment":"The frequency-sampling stage is underspecified: the minimum frequency threshold, the blocklist size, and the downsampling rule are not reported, even though this stage removes documents from the final corpus and could affect the boundary.","section":"§3.1.5 and §A.4"},{"comment":"The phrase 'diverge framing-dependent at Grade 7' is ungrammatical and should be rewritten, for example as 'diverge depending on the framing of the problem at Grade 7.'","section":"§C.2.2"},{"comment":"Figure 1 labels 'ln(0) = -1' as an out-of-scope example, but ln(0) is undefined rather than equal to -1; since the paper is about mathematical boundaries, this is a distracting error.","section":"Figure 1"},{"comment":"The naming of the post-training data variants is inconsistent between the text ('K-5 data versus Beyond-K-5 data') and the figure legends ('K-5 / Unfiltered'); the authors should use one terminology throughout.","section":"§4.2 and §3.2"}],"recommendation":"major_revision","confidential_remarks":"The post-hoc MathCAMPS filtering in Section C.2.1 is the main risk to the paper's headline claim; I would not accept Section 4.4 without the sensitivity analysis described in Major Comment 1. The authors should also state, for transparency, whether the 6.EE.B.7 exclusion and the low-entropy-standard exclusion were fixed before the intervention experiments were run or introduced afterward in response to the results. The corpus and model release are valuable, and the WeeBit and n-gram checks are genuine strengths, but the single-benchmark basis of the intervention claims currently makes the scope of the conclusion larger than the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful part. This paper delivers a real resource: an 88B-token K-5 corpus and a 5B model trained on it, with an unusually careful filter validation. The multi-stage pipeline is checked on CommonCoreText, WeeBit, and an n-gram scan; the precision-first design is sensible, and the matched UNFILTERED baseline makes the within- versus beyond-K-5 gap credible. The qualitative examples are striking. For anyone who wants a tractable substrate to study knowledge boundaries, LittleCurriculum/LittleLearner is worth having.\n\nThe scaling experiment is also honest: within K-5, scale helps; beyond it, mostly not. That part is clear.\n\nNow the soft spot, and it is substantial. The claim that post-training and ICL cannot move the model beyond K-5 rests almost entirely on MathCAMPS, and the Beyond-K-5 split of that benchmark is constructed post hoc in C.2.1. Excluding 6.EE.B.7 because it resembles grade-3 add/sub word problems removes exactly the kind of crossover item on which a K-5-trained model with more computation or better prompting might plausibly improve. Removing low-entropy items and questions with embedded gold answers is defensible, but it further narrows the beyond-K-5 set. So the observation that interventions do not lift the ceiling may partly be a property of the evaluation filter rather than the pretraining boundary. The paper's own Section 4.4 assertion that 'the pretraining filter, rather than the intervention, sets the effective capability ceiling' is stronger than the evidence supports.\n\nAlso, the ICL test is narrow: the few-shot exemplars are K-5 style, hand-authored, and only three per standard; the conclusion should be phrased as 'the ICL setups we tried,' not a blanket statement.\n\nThe abstract is mostly okay, but the body is more careful. The code, URL, and commit hashes are not yet in the preprint; that matters for reproducibility but is fixable.\n\nNet: the dataset and model are a real contribution, and the moderate claims are fine. The strong ceiling claim needs either an unfiltered evaluation set or an analysis showing the excluded items do not change the conclusion. That is a referee-level revision, not a desk reject.\n\nWho is this for? People working on data-centric AI, knowledge boundaries, and RL/post-training credit assignment. It deserves serious peer review. I would cite the resource if I worked on controlled exposure.","headline":"A genuinely useful controlled-exposure resource whose strongest claim about the ceiling on post-training and ICL is still entangled with the post-hoc MathCAMPS evaluation filters.","tokens_in":28622,"tokens_out":2581,"would_cite":true,"duration_ms":27337,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a model's pretraining knowledge boundary, set by a K-5-filtered corpus, acts as a hard ceiling that scaling, post-training, and in-context learning cannot break through.","keywords":["pretraining data filtering","knowledge boundary","curriculum-constrained corpus","controlled exposure","post-training","in-context learning","model scaling","K-5 education"],"falsifier":"Train a matched 5B model on a version of the released corpus spiked with a small, measured fraction of grade-8 mathematics content, for example 1%, then evaluate it on the paper's grade-8 math questions. If accuracy stays at the original model's floor, the ceiling is not caused by content exposure alone; if it rises, the paper's attribution of the ceiling to the pretraining filter is falsified.","tokens_in":27575,"feed_emoji":"🎓","tokens_out":10196,"duration_ms":101559,"temperature":0.7,"pith_summary":"The paper tries to establish that a language model's pretraining distribution, not later interventions, sets the boundary of what it can do. To show this, the paper constructs an 88B-token corpus restricted to U.S. kindergarten through grade 5 (K-5) material and trains a 5B-parameter model, LittleLearner, entirely on it. Across scaling, supervised fine-tuning followed by GRPO post-training, and few-shot in-context learning, the model improves within its K-5 scope but does not meaningfully move beyond it. If right, this gives researchers a controlled sandbox for studying knowledge acquisition, because prior exposure is known exactly rather than inferred after the fact.","feed_headline":"Pretraining, not scale or post-training, sets this model's ceiling","feed_subtitle":"A K-5-restricted 5B model stays at grade 5 even after scaling, reinforcement learning, and few-shot examples.","key_machinery":"The mechanism that carries the argument is the multi-stage LittleCurriculum filter, which turns a large web-scale educational corpus into a sharply bounded K-5 corpus. It works in layers: an age-of-acquisition prefilter with frequency-based imputation removes texts whose vocabulary is learned after age 12; a lightweight text classifier and then a stronger classifier, trained on LLM-as-judge annotations grounded in curriculum standards, assign grade bands; a symbolic regular-expression stage removes mathematical notation such as equations, exponents, and integrals; and a final frequency-sampling stage drops documents concentrated with beyond-K-5 terms. Validation on a held-out curriculum-aligned benchmark and an external grade-labeled corpus is what makes the boundary interpretable and grounds the ceiling claims.","core_discovery":"The paper's central claim is that a deliberately restricted pretraining corpus creates a sharp, interpretable capability boundary, and that the boundary is set by the filter that built the corpus rather than by the model or by subsequent training. LittleLearner answers grade-school questions comparably to an unfiltered control, but on questions beyond grade 5 it collapses, producing plausible but wrong answers such as describing an out-of-scope quantum mechanics thought experiment as a literal cat with fabricated attributes. Scaling the model from 0.6B to 5B parameters improves in-scope and boundary performance but leaves grade-8 math at floor; SFT followed by GRPO post-training, even when the post-training data are unfiltered, does not close the gap; and few-shot demonstrations with hand-written reasoning traces give at most a small in-scope boost. The paper reads the consistent pattern as evidence that the pretraining filter, not the intervention, sets the effective capability ceiling in these tested settings.","pith_inferences":["The same design can be reused to test whether other post-training paradigms, such as long-horizon reinforcement learning, verifier-guided search, or retrieval, can genuinely introduce new capabilities rather than elicit latent ones, since the model's prior is fully known.","The K-5 boundary is a content boundary, not a human-developmental claim; the model's skill ordering differs from curriculum progression, so any prerequisite reasoning should be tested against the filter's actual content rather than grade-level labels.","The ceiling result may be scale-dependent: at frontier scales with stronger in-context learning, the same filter might show more boundary flexibility, so the sandbox should be re-run at larger sizes before generalizing to all language models.","A direct way to test the mechanism is to spike a small fraction of out-of-scope math content into a filtered corpus and measure whether grade-8 accuracy rises; if it stays at floor, the ceiling is not solely about content exposure."],"forward_implications":["A model trained under K-5 exposure performs on par with an unfiltered model on in-scope math and factual questions, so the filtering does not destroy elementary competence.","Increasing model parameters from 0.6B to 5B yields strong in-scope gains but leaves fully out-of-scope grade-8 performance at floor, so scale does not by itself extend the knowledge boundary.","SFT plus GRPO post-training lifts in-scope math accuracy for both the restricted and unfiltered models but does not close the beyond-K-5 gap, even when the restricted model is post-trained on unrestricted data.","Few-shot in-context learning with hand-written chain-of-thought examples can steer output format but does not unlock new out-of-scope reasoning for this 5B model.","For the tested settings, the pretraining filter, not the later intervention, is the effective capability ceiling."],"supporting_citations":[{"why":"Supplies the web-scale educational corpus that the filtering pipeline turns into the K-5 corpus.","marker":"[1]"},{"why":"Provides the age-of-acquisition word norms used in the first filtering stage to remove advanced vocabulary.","marker":"[33]"},{"why":"Defines the curriculum standards that anchor the annotations and set the K-5 boundary.","marker":"[35]"},{"why":"Implements the lightweight text classifier used as the first learned filtering stage.","marker":"[38]"},{"why":"Implements the stronger classifier applied after the coarse stage to assign grade bands.","marker":"[39]"},{"why":"Provides an external grade-labeled corpus used to audit beyond-K-5 leakage.","marker":"[41]"},{"why":"Generates the grade-aligned synthetic math questions used to measure the capability boundary and ceiling.","marker":"[51]"},{"why":"Supports the interpretation that post-training amplifies pretrained behavior rather than creating new capabilities.","marker":"[27]"}],"fun_headline_variants":["K-5 model's ceiling set by pretraining filter, not scale","Pretraining filter, not post-training, caps model at grade 5","Scaling and RL can't push model beyond grade 5","LittleLearner: grade 5 ceiling from pretraining choice","Filtered pretraining, not model size, sets knowledge bounds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The filtering pipeline's measured leakage rates on validation sets, near-zero beyond-K-5 retention, 2.48% on an external corpus, and 0.09% phrase matches, are representative of the actual 88B-token corpus, so that the K-5 boundary is genuinely clean.","fun_headline_variants_meta":{"raw":{"variants":["K-5 model's ceiling set by pretraining filter, not scale","Pretraining filter, not post-training, caps model at grade 5","Scaling and RL can't push model beyond grade 5","LittleLearner: grade 5 ceiling from pretraining choice","Filtered pretraining, not model size, sets knowledge bounds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000771,"raw_usage":{"total_tokens":3419,"prompt_tokens":957,"completion_tokens":2462,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":2372}},"tokens_in":573,"tokens_out":2462,"duration_ms":17257,"temperature":1.0,"reasoning_tokens":2372,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:25:07.944942+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a matched 5B model on a version of the released corpus spiked with a small, measured fraction of grade-8 mathematics content, for example 1%, then evaluate it on the paper's grade-8 math questions. If accuracy stays at the original model's floor, the ceiling is not caused by content exposure alone; if it rises, the paper's attribution of the ceiling to the pretraining filter is falsified.","supporting_citations":[{"cited_title":"FineWeb-edu: the finest col- lection of educational content, 2024,","cited_arxiv_id":null,"evidence_quote":"Supplies the web-scale educational corpus that the filtering pipeline turns into the K-5 corpus."},{"cited_title":"Common core state standards","cited_arxiv_id":null,"evidence_quote":"Defines the curriculum standards that anchor the annotations and set the K-5 boundary."},{"cited_title":"On improving the accuracy of readability classification using insights from second language acquisition","cited_arxiv_id":null,"evidence_quote":"Provides an external grade-labeled corpus used to audit beyond-K-5 leakage."}],"review_version":1}