{"id":"5d16f535-3d52-40b3-ace7-5e91bc899167","arxiv_id":"2608.09093","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Deleting structural announcements makes following prose harder for LLMs to predict, swapping notation does nothing, and the paper proposes a pure-frame format that strips announcements into sidecars.","lead":"This paper measures a training variable no dataset card records: the notation or markup that survives in pre-training corpora, and reports that the announcement of a boundary, not the markdown symbols that carry it, is the cue models actually use. It ships a 'pure frame' format that deletes announcements into a reversible sidecar, and asks data cards to record extractor identity and a new survival statistic.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Training-time translation of the announcement deficit is an untested bet; reader-side results don't settle it.","rationale":"The paper's strongest claim is the reader-side finding, which is well supported across six readers and three pipelines. However, the shipped format (pure frame mixed on announcement presence) is justified by the reliability mechanism, which is explicitly a bet. The reader already conditions acceptance on the training-time test. I agree with that. No additional concern changes the verdict; the scientific posture is unusually honest, with pre-registration, falsifier ledger, and costed experiment. The one concrete test that would settle the concern is the three-arm training run or its cheaper slot-substitution precursor. Until then, CONDITIONAL remains the right verdict.","tokens_in":40766,"tokens_out":5547,"duration_ms":64093,"concrete_test":"Run the three-arm training experiment of §10 Test 3 at the pre-specified screening scale (29B tokens per arm): marked, pure, and paired (half-marked/half-pure by work). Evaluate with the paper's own readouts (probe announcement contrast, whole-book curve, generation probe) plus a downstream long-range benchmark such as ChapterBreak. Pre-register the decision rule that the paired arm must differ from the marked arm on at least one instrument; if it does not, the reliability mechanism is falsified and the format recommendation loses its warrant.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The pure-frame recommendation assumes the reader-side announcement deficit will force a trainable discourse-channel factorization. The paper explicitly labels this a bet in §9 and specifies but does not run the decisive three-arm test in §10. If the paired arm is indistinguishable from the marked arm, reliability is not the operative variable and the design collapses to a notation preference. This is not an internal inconsistency—the paper pre-commits to reporting the collapse—but it is the load-bearing empirical link for the central format claim. The reader-side R1/R3 results, however robust, cannot distinguish 'training pressure never existed' from 'these readers never formed the machinery', a limit the paper itself states in §11. So the central claim as a design recommendation rests on an untested assumption, exactly as the reader's weakest_assumption says.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces clean-window survival S(W), a deterministic, seedless count of how much of a token stream still demands boundary inference, and uses it to measure structural notation in thirteen public corpora. It reports a pre-registered converter study whose own prediction failed, a three-arm reading probe across five base models (plus a sixth replication) showing that deleting an announcement line raises next-token prediction difficulty while swapping the notation changes long-range information gain by a measured zero, and a writer-side generation probe examining whether base models re-impose markup. On the basis of these measurements it proposes a 'pure frame' format that deletes announcements into a reversible sidecar and recommends mixing marked and pure copies at training time, together with a data-card field for notation.","tokens_in":40910,"tokens_out":6178,"duration_ms":66868,"significance":"The paper's reader-side core is genuinely well controlled: the three-arm design holds content byte-identical across arms, the anchor set is identical across all readers, the notation-swap null is tight and pre-registered, and the announcement deficit replicates across six readers and three pipelines. The paper also deserves explicit credit for shipping reproducible artifacts, a falsifier ledger with pre-committed directions, a costed and pre-registered decisive test, and unusually candid disclosure of its own failed predictions and limitations. If the central claim is taken as the reader-side finding that the announcement, not the sigil, is the operative cue for long-range reading at fixed checkpoints, the evidence is strong. However, the paper's format recommendation is explicitly a bet about training-time behavior, and several census and converter-study results are weaker than the abstract suggests, so the central design claim currently rests on an untested empirical link.","major_comments":[{"comment":"The pure-frame recommendation is explicitly a bet: §9 says 'We are betting that the factorization exists and that reliability is what forces it,' and §10's Test 3 is described as 'Specified and costed; not run.' The reader-side evidence in §7 cannot distinguish 'the training pressure never existed' from 'these readers never formed the machinery,' a limit the paper itself states in §11. Because the abstract and §9 present the format as 'what those measurements imply,' the central design claim currently rests on an untested training-time link. I ask the authors either to run Test 3 or to reword the central claim so that the format proposal is explicitly conditional on the paired-arm outcome.","section":"§9–§10"},{"comment":"The census's flagship number is not reproducible as published: Table 2 reports re-measured S(8,192) for the olmOCR slice as 0.271 and 0.296 against the published 0.153, and seven of the ten published rows have no recorded sample size, document count, sampling rule, revision or date. Although the qualitative claim that olmOCR has the lowest survival survives (0.271 is still below DCLM's 0.558), the abstract's 'falls to 0.153' states a number the paper's own re-verification does not support. The corrected value, the limited window counts (23, 62, 138), and the provisional status of the table should appear prominently, including in the abstract.","section":"§5.1, Table 2; abstract"},{"comment":"The pre-registered converter null was measured on nine non-visual configurations, and the paper states in the same section that olmOCR and MinerU 'were provisioned and produced no rows.' Since the census's most contaminated slice is produced by a vision-language converter and §6 identifies that slice as the input to the long-context stage, the claim that modern converters under-mark cannot be extended to the vision-converted frontier without an additional assumption. The null should be explicitly scoped to non-visual converters, or the missing VLM rows should be run.","section":"§5.4"},{"comment":"The writer-side front's headline 'bounded null' is not supported by the primary greedy-decoding arm: the repetition guard fired on 84–100% of continuations in every cell, leaving 77 of 1,328 generations, and the paper itself calls this a selection effect. Only the temperature-1.0 secondary arm on 31 anchors supports the I3 exact zero, while I1, I2 and I4 are unevaluable in the primary arm. The front's conclusions should be rewritten around the secondary arm, with I1/I2/I4 reported as indeterminate rather than as components of a 'bounded null.'","section":"§8, Table 7"},{"comment":"There is an internal inconsistency in the announcement-reconstruction claim. §7 reports that R2 fired at 1.7B: the information-gain contrast +0.270 [+0.010, +0.521] excludes zero, which the pre-registered rule defines as larger readers reconstructing a deleted announcement. Yet §8 states that 'no reader recovers a deleted announcement from long-range context either,' and §11 lists R2 as a fired rule that 'cuts against this paper.' The §8 sentence should be corrected to say that reconstruction was observed at one reader and did not replicate at three larger ones; as written, it contradicts the paper's own pre-registered outcome.","section":"§7, §8, §11"}],"minor_comments":[{"comment":"The supply-side verdict 'institutional, not consumer' rests on one sampling-sensitive row: reading from the head of a shard changes Hansard's median document by almost a factor of eight and S(8,192) from 0.930 to 0.757. The abstract should carry the caveat that this conclusion depends on the seeded-shuffle sampling rule.","section":"§5.2"},{"comment":"The adjusted narrative precision of 0.967 lands in the registered SUPPORTED, WEAKENED band, not SUPPORTED AS STATED. The paper does say 'near-perfect reliability' rather than 'perfect,' but the abstract and §9 should make equally explicit that the reliability claim is scoped to narrative prose and does not hold for sustained argument.","section":"§5.5, Table 5"},{"comment":"The long-context stage argument cites Olmo 3's long-context pool as 'to a rounding error, entirely olmOCR-converted science PDFs.' Given the olmOCR census row's re-measurement divergence, this sentence should cite both the published and re-measured survival values, or at least flag which value is being used.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The reader-side core is well controlled and the paper is unusually transparent, which is why I am not recommending rejection. The revision must address the untested training-time link, the census reproducibility and olmOCR divergence, the absence of VLM converters in the converter study, and the R2 contradiction in §8, before the central format claim can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this paper gives the field a genuinely missing variable—notation—and a cheap deterministic instrument, clean-window survival, to measure it. The census of thirteen corpora and the controlled probe across five base models are new, and the pure-frame format with sidecar and validator is a concrete artifact. The paper is unusually honest: it pre-registers rules, reports nulls that ran against its own predictions, and ships code and data.\n\nWhat's strong: the reader-side result survives scaling. Deleting an announcement makes following prose harder to predict at six readers across three pipelines; swapping notation moves nothing. That dissociation is tight and the determinism gates pass. The census, despite provenance problems, establishes a real gradient: web corpora are clean and short, PDF-converted corpora are long and marked, books are long and low-density. The data-card proposal is modest and realistic.\n\nThe soft spots are real but mostly disclosed. Seven of ten census rows can't be regenerated from the original run, and the olmOCR row re-measures from 0.153 to 0.271/0.296—still lowest, but the headline number is shaky. The converter study has no vision-language converter, and the writer-side front is crippled by greedy decoding collapse; the sampled arm rescues the main null but not the scaling claims. The biggest issue is the one the paper itself names: the pure-frame recommendation depends on the assumption that the reader-side announcement deficit translates into a trainable capability. That is a bet, and the decisive three-arm training test hasn't been run. If the paired arm looks like the marked arm, the factorization collapses to a notation preference. The paper pre-commits to reporting that, but as it stands the format claim is untested at the point where it matters.\n\nWho it's for: people working on pre-training data curation and long-context training. It deserves a serious referee—there's enough new measurement here to engage with even if the training-time bet fails. I'd ask the authors to run the cheaper slot-substitution precursor or the three-arm test before the format recommendation is taken as established; the census and probe warrant publication on their own.\n\nRecommendation: send to peer review, with the training-time bet clearly flagged as unresolved.","headline":"A transparent, pre-registered paper that adds a missing variable to pre-training data curation; the format recommendation rests on a disclosed but untested training-time bet.","tokens_in":41409,"tokens_out":1802,"would_cite":true,"duration_ms":19902,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The announcement is the cue: deleting a structural heading makes following prose harder to predict at every model scale tested, while swapping the markup notation changes nothing.","keywords":["pre-training corpora","structural notation","announcement cue","clean-window survival","boundary inference","long-context training","pure frame","shortcut learning"],"falsifier":"Run the paper's three-arm training test over identical content: marked, pure, and paired arms at about 100 billion tokens per arm; if the paired arm ends indistinguishable from the marked arm on the held-out announcement probe, reliability was not the operative variable and the factorization collapses to a notation preference. The cheaper precursor is replacing 5 to 25 billion tokens of a 50-billion-token long-context stage with pure-frame text; if no instrument moves, the locus argument fails.","tokens_in":40553,"feed_emoji":"📑","tokens_out":8162,"duration_ms":82281,"temperature":0.7,"pith_summary":"Pre-training corpora are usually described by what documents contain, not by how their arrangement is written down. This paper argues that this notation is an unrecorded training variable, and that for long-range reading the cue a model actually uses is the announcement, a short standalone line saying a boundary is here, not the markup sigil that writes it. Across thirteen corpora the paper measures clean-window survival, the share of fixed-length windows with no structural markup, and finds the scarce resource is long unmarked text, which has largely left modern mixtures. Across five base models and two pipelines, deleting an announcement makes following prose measurably harder to predict, while swapping the notation changes nothing; base writers do not put deleted announcements back. The paper ships a format, the pure frame, that deletes every announcement into a reversible sidecar, mixed against marked copies over announcement presence, and argues this should be aimed at the long-context training stage.","feed_headline":"Deleting headings hurts language models; markup swap does nothing","feed_subtitle":"A census and reader probe show long unmarked prose is the scarce training resource; the paper ships a format that supplies it.","key_machinery":"The central instrument is clean-window survival, S(W), the deterministic fraction of non-overlapping W-token windows in a tokenized stream that contain no structural markup; a companion statistic is the distribution of clean runs, the longest markup-free spans. The paper pairs this with a conceptual distinction between the sigil, the markup character such as ##, and the announcement, the short line set off by whitespace that says a boundary is here. The hinge of the argument is the announcement's reliability: in narrative a structural sigil marks an authored boundary at precision 0.994, or 0.967 after documented typographic apparatus is removed, so the cue is near-perfectly predictive and therefore a shortcut, while the sigil itself is interchangeable. The pure frame is the machinery's output: paragraphs in authored order with every announcement deleted into a reversible sidecar, so a stream demands the boundary inference instead of receiving it.","core_discovery":"The paper's central claim is that across long prose, the cue a trained language model actually uses is the announcement, a short standalone line such as a chapter title saying that a boundary is here, and not the markup sigil that writes the announcement down. On a fifty-document ground-truth corpus, deleting the announcement makes the following prose measurably harder to predict at every reader scale tested, from 0.60B to 8.19B parameters, while swapping Markdown for a bare line moves nothing, inside a pre-registered bound under two percent of measured information gain. A census of thirteen corpora measured with the paper's clean-window survival statistic finds the scarce resource is long unmarked text: survival falls to 0.153 in a vision-converted PDF slice versus 0.889 in a standard web corpus, and the longest clean runs live in books, the row that left modern mixtures. The paper concludes that faithful flattening of markup is a no-op for boundary inference, and ships a format, the pure frame, that deletes every announcement into a reversible sidecar, mixed against marked copies over announcement presence rather than notation.","pith_inferences":["If the central claim holds, data-mixing methods that reweight documents as already built cannot reach notation; the only operators that can are deletion-style transforms such as the pure frame, so the actionable design space for long-context curation shifts from weighting to format operators.","The announcement cost concentrates in roughly the first thousand tokens after a boundary, which predicts a testable training signature: a model trained on pure-frame long text should show its largest gains on prose just after boundaries, not uniformly across a document.","The notation-invariance null implies that information-gain or attention-based selectors for long-context data, which ignore serialization, may be picking on the wrong coordinate; comparing such selectors over marked versus pure-frame versions of the same works would settle it cheaply.","If the format works as intended, model-generated prose should carry fewer explicit signpost announcements; a blind or automated comparison of announcement density between models trained on the two mixtures would be a direct product-level check."],"forward_implications":["Faithful flattening of markup does not restore the missing inference demand; only deleting the announcement line does, so corpus operators should act on announcements, not sigils.","The pure frame should be mixed with marked copies over announcement presence, so the same underlying boundary is sometimes announced and sometimes not, attacking cue reliability at the level that matters.","The long-context training stage, where the sequence is the document, is the right target: it is under one percent of the token budget and currently fed the most heavily marked material in the census.","Data cards should record extractor identity, conversion target, and clean-window survival at training context lengths, a field the paper argues one flagship PDF corpus already partially carries.","The scarce resource is long unmarked text, not unmarked text: institutional proceedings such as parliamentary records supply it, consumer transcripts and library scans do not, and books left modern mixtures."],"supporting_citations":[{"why":"Shows circuit formation in pretraining is more sensitive to reliability than frequency, the mechanism connecting cue reliability to shortcut formation.","marker":"[5]"},{"why":"Establishes that shortcut adoption depends on how reliably and how easily a feature predicts, the basis for treating the announcement as the operative cue.","marker":"[27]"},{"why":"Provides the theoretical treatment of shortcut learning that supports the claim that a faithful, near-perfectly predictive cue is what models take.","marker":"[28]"},{"why":"Documents an eighteen-property audit of major corpora with no notational category, evidence that notation is absent from the record.","marker":"[33]"},{"why":"Records a frontier pipeline that deliberately removed Markdown, and gives the token split between the 8,192-token main stage and the long-context stage.","marker":"[50]"},{"why":"Supplies the line filter that discarded headings from the old web corpus, the largest accidental destroyer of structural markup.","marker":"[65]"},{"why":"Is the prior finding that long-range models fail to use chapter boundaries, which the reader-side probe extends.","marker":"[73]"},{"why":"Shows the long-context stage is fed almost entirely vision-converted science PDFs, the row with lowest clean-window survival.","marker":"[76]"}],"fun_headline_variants":["Announcement, not markup, is the cue LMs use for boundaries","LM boundary cue: the announcement, not the sigil","Long unmarked prose is the scarce training resource","Deleting headings hurts LM; markup swap does nothing","Pure frame: delete announcements, ship reversible sidecar"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recommendation stands or falls on the bet that the reader-side difficulty from deleted announcements translates into a trainable capability, so that a model trained on announcement-free text learns the boundary inference; the paper says the decisive training run has not been run.","fun_headline_variants_meta":{"raw":{"variants":["Announcement, not markup, is the cue LMs use for boundaries","LM boundary cue: the announcement, not the sigil","Long unmarked prose is the scarce training resource","Deleting headings hurts LM; markup swap does nothing","Pure frame: delete announcements, ship reversible sidecar"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001134,"raw_usage":{"total_tokens":4795,"prompt_tokens":1117,"completion_tokens":3678,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":733,"completion_tokens_details":{"reasoning_tokens":3610}},"tokens_in":733,"tokens_out":3678,"duration_ms":26625,"temperature":1.0,"reasoning_tokens":3610,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:40:40.463233+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's three-arm training test over identical content: marked, pure, and paired arms at about 100 billion tokens per arm; if the paired arm ends indistinguishable from the marked arm on the held-out announcement probe, reliability was not the operative variable and the factorization collapses to a notation preference. The cheaper precursor is replacing 5 to 25 billion tokens of a 50-billion-token long-context stage with pure-frame text; if no instrument moves, the locus argument fails.","supporting_citations":[{"cited_title":"propella-1: Multi- property document annotation for LLM data curation at scale.arXiv preprint arXiv:2602.12414,","cited_arxiv_id":null,"evidence_quote":"Documents an eighteen-property audit of major corpora with no notational category, evidence that notation is absent from the record."},{"cited_title":"ChapterBreak: A challenge dataset for long-range language models","cited_arxiv_id":null,"evidence_quote":"Is the prior finding that long-range models fail to use chapter boundaries, which the reader-side probe extends."}],"review_version":1}