{"id":"b79aa752-7bd4-4fa2-a176-63cccfbe084e","arxiv_id":"2608.12072","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"A new randomized negative-control design, Level II-A, turns 'the past explains it' into a magnitude-qualified testable claim by randomizing a delay after endpoint commitment, validated on synthetic anticipatory EEG data.","lead":"This paper introduces Level II-A, a statistical framework that tests whether a pre-event measurement can be fully explained by past information, by randomizing a delay after the measurement is locked and using that delay as a negative control. The authors demonstrate on simulated data how the design can support either a qualified negative finding or a bounded clean null, with public benchmark code, in an anticipatory EEG worked case.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Retained-sample delay-neutrality (R3*) is the load-bearing unverifiable premise: a negative residual ordering can be manufactured by selection after all declared audits pass, so the insufficiency claim is conditional on a condition the design cannot certify.","rationale":"The reader's weakest-assumption analysis and the stress-test pass converge on the same point: the retained-sample delay-neutrality condition (R3*) and the frozen-comparator independence condition (S5) are sufficient, not verifiable, premises on which the real-world interpretation of a Level II-A result depends. The paper itself repeatedly acknowledges this, labels all conclusions as conditional, and makes no empirical claim about human EEG. The synthetic benchmark is scoped to a declared injection family and to a pipeline whose audits are shown to catch the leakage, standard-selection, and balanced-collider generators tested. The concern is therefore a real limitation of the framework's transferability rather than an internal inconsistency in the proposal. The proposed computational test would determine whether the audit battery can be defeated by an unmeasured-selection mechanism that is not in the declared failure surface. If it can, the framework would still be coherent as a conditional design, but its practical ability to deliver an unconditional empirical verdict would be weaker than the abstract's 'testable claim' language suggests. This does not undermine the paper's mathematical core, which is correct within its stated conditions, nor its value as a design-based inference template. Hence the reader's ACCEPT verdict remains appropriate, with the caveat that the load-bearing nature of R3* should be foregrounded in any empirical application.","tokens_in":44160,"tokens_out":5628,"duration_ms":59434,"concrete_test":"Use the public benchmark pipeline (Zenodo 21887583, version 1.2.1) to add a selection generator that satisfies all declared audits. Let U be an unrecorded latent trial-level variable with U ~ N(0,1), let the forward-only endpoint be A_pre = f(X) + U + noise, and set inclusion probability P(S=1 | τ, Z, R, U) to depend on U·(τ - E[τ|R]) while keeping marginal retention balanced within each stratum and keeping the declared endpoint-by-delay collider diagnostic below its threshold. Omit U from Z, R, and G_frz. Run the full classifier with M = 1200 datasets under a true past-adapted forward-only process. If the supported-negative rate materially exceeds the declared false-support control, this demonstrates that R3* failure is not screened by the audit battery and that the insufficiency inference is not uniquely identified by the design.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inferential step is Lemma S1 in the SI: to carry the pre-selection no-ordering implication to the frozen residual, the design requires (R3*) S ⟂ τ_L | Z, R, G_frz (Eq. 3) and the frozen-comparator independence condition (S5) τ_L ⟂ Z | R, G_frz. The authors state explicitly that these are sufficient conditions, not implied by randomisation, and that audits and sensitivity analyses only qualify them within a declared envelope. The load-bearing problem is that they are not merely unestablished; they can fail in ways that pass the entire declared battery. Figure 2's collider path Z → S ← τ is only one failure mode. A latent common cause U of the endpoint and inclusion, or a selection rule depending on an unrecorded endpoint-by-delay interaction, can generate a retained-sample endpoint-delay association while marginal retention is balanced per stratum and the declared collider diagnostic, which targets the recorded endpoint-by-delay interaction, does not fire. Under such a process, the classifier's supported negative departure or clean affirmative null would misattribute selection to class insufficiency, or vice versa. Because the design's real-world conclusion is exactly that attribution, the framework's central claim is not self-certifying. This limitation is scoped in Section 3.2 and SI Section 1.4, but it is the point at which the synthetic benchmark cannot substitute for assurance in an empirical deployment.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Level II-A, a design-based inference framework for testing whether a predictive model class is informationally sufficient for a committed endpoint, using post-endpoint randomisation as a negative-control probe. In the worked anticipatory-EEG instantiation, a pre-event summary A_pre is committed at time t1, after which the delay to the imperative event is randomised within a declared stratum. Under the declared past-adapted factorisation, any temporally admissible pre-assignment statistic is conditionally independent of the assigned delay given the stratum, yielding a no-ordering implication. The authors show that carrying this exclusion to the retained, frozen-residual analysis requires three additional conditions: retained-sample delay-neutrality (R3*, Eq. 3), retained-support positivity, and frozen-comparator independence (S5). A non-compensatory decision rule classifies outcomes as supported negative departure, forward-only adequate affirmative null, diagnostic failure, selection-limited, opposite-direction, or inconclusive, with the affirmative null calibrated by simulation-based false-adequacy boundaries of 15 µV/s (assignment-isolation route) and 30 µV/s (sequential e-value route). The paper reports no human EEG data and relies on a seven-generator synthetic benchmark with a certified machine-readable run.","tokens_in":44518,"tokens_out":5746,"duration_ms":61511,"significance":"If the framework holds up, it is a valuable contribution: it turns the common explanatory claim that 'the past explains the endpoint' into a magnitude-qualified, falsifiable statement, and it does so without requiring a mechanistic model. The central derivation, Proposition 1 and Lemma S1, is logically sound given the stated assumptions, and the authors are unusually explicit about what those assumptions do and do not establish. Particular strengths are the non-compensatory decision architecture, the route-specific calibration separating assignment isolation from sequential e-value inference under carryover, the leakage-safe causal preprocessing requirements, the prospective locking of the endpoint and comparator, the comprehensive synthetic benchmark with archived code and certified run outputs, and the explicit acknowledgement that R3* and S5 are sufficient conditions that cannot be globally certified.","major_comments":[],"minor_comments":[{"comment":"The sentence 'leakage-safe preprocessing, a frozen label-blind comparator and retained-sample qualifications carry the exclusion to the confirmatory residual' could be read as implying that R3* and S5 are operational achievements; consider adding the qualifier 'under sufficient conditions that are qualified, not established, by the audits' to match the careful wording in Section 3.2.","section":"Abstract"},{"comment":"The manuscript already states prominently that randomisation alone does not establish (R3*) and that (S5) is not implied by the pre-selection factorisation; consider placing a short 'envelope of interpretation' box near Eq. (3) so that the conditional nature of the central claim is visually impossible to miss in secondary citations.","section":"Section 3.2, Eq. (3) and SI §1.4"},{"comment":"In panel (a), the annotation 'slope --      →   null' is unclear; I recommend writing 'slope is not significantly negative → classified as forward-only adequate' in plain text.","section":"Figure 3"},{"comment":"The outcome name 'forward-only adequate' is potentially misleading because it denotes a bounded affirmative null within the evaluated linear injection family, not global adequacy; consider renaming it 'qualified affirmative null' in the table and text to match the careful interpretation given in Section 5.2.","section":"Table 1"},{"comment":"The certified false-adequacy boundaries (15 and 30 µV/s) are easily confused with the resolution floor β_min from Eq. (14); I suggest stating explicitly in Section 6.1 that δ*_r,d are classifier-level operating-characteristic boundaries and are distinct from the per-participant resolution floor.","section":"Section 6.1"},{"comment":"There is a typographical error in the sentence 'Because the factorisation holds for every bounded measurable test function φ and every integrable prospectively fixed h, division by this positive conditional probability yields...' where the period is missing after 'yields'; the sentence should end with a period.","section":"SI §1.3, proof of Lemma S1"}],"recommendation":"accept","confidential_remarks":"This is a methodological Perspective with no empirical data, which is appropriate for the journal's scope as a design-based inference contribution. The authors provide a public, versioned code archive and a certified benchmark run, which is a strength. The main risk for acceptance is not technical error but overstatement in secondary citations; the manuscript's own conditional framing is careful enough that I do not see a need for major revision before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, the core idea is genuinely new: randomise a delay after the endpoint has been committed, then use that delay as a negative-control probe for whether past-adapted information was sufficient. The lemmas are elementary, but the synthesis — committed endpoint, frozen label-blind comparator, non-compensatory decision rule, and a false-adequacy benchmark calibrated by simulation — is a real contribution to design-based inference. Second, the paper is unusually honest about its own limits. The authors explicitly flag that R3* (retained-sample delay-neutrality) and S5 (frozen-comparator independence) are sufficient conditions that cannot be established globally, and that audits and sensitivity analyses only qualify them within an envelope. That is not a buried assumption; it is in the main text and the SI.\n\nWhat the paper does well: the math is correct given the stated assumptions, the decision architecture is transparent and the seven scenario families in the synthetic benchmark are comprehensive. The archived code and certified run matter. The paper also avoids overclaiming — no human EEG data, and the conclusions are explicitly conditional.\n\nThe real soft spot is exactly what the stress-test note says: a negative ordering can be manufactured by selection after all declared audits pass, via a latent common cause or an unrecorded endpoint-by-delay interaction that the collider diagnostic does not catch. This means the field-level insufficiency claim is conditional on a condition that cannot be certified. That is a genuine limit, and it is proportionate to say so. But the paper already says so, repeatedly. It does not present the design as self-certifying. The inference is conditional, and the authors want the reader to know it.\n\nA stronger paper might have developed a sharper sensitivity analysis for R3* — for instance, bounding the magnitude of selection-induced slope as a function of a plausible selection mechanism — rather than the current scalar gate, which the authors themselves note is not an affirmative-null gate. But that is a refinement, not a fatal gap.\n\nWho is this for? Methodologists working on negative controls, randomised experiments, or preregistration in neuroscience; also anyone planning anticipatory EEG studies where commitment-before-randomisation can be implemented. It deserves a serious referee. I would send it to review and expect a decent conversation about how to make R3* more auditable in practice, but the core framework is sound and useful.\n\nMy recommendation: engage with it. Cite it if you work on design-based inference. I would bring it to the reading group.","headline":"A careful, well-scoped methods paper whose novel negative-control design is real; the unverifiable retained-sample neutrality is an honest epistemic limit, not a hidden flaw.","tokens_in":45004,"tokens_out":1914,"would_cite":true,"duration_ms":20991,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62G10","62K10","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that randomising the delay after committing a pre-event EEG endpoint lets a researcher test, rather than assume, whether past-adapted information is sufficient.","keywords":["anticipatory EEG","contingent negative variation","temporal expectation","post-endpoint randomisation","negative controls","randomisation inference","design-based inference","informational sufficiency"],"falsifier":"Apply the full locked pipeline to a synthetic past-adapted generator in which inclusion is driven only by the endpoint-by-delay collider while marginal retention is balanced and all operational audits pass; the classifier must return selection-limited in essentially every dataset, as its own benchmark reports, so any independent run returning directional support would falsify the non-compensatory rule.","tokens_in":43960,"feed_emoji":"🧠","tokens_out":9172,"duration_ms":79613,"temperature":0.7,"pith_summary":"This paper attempts to establish that informational sufficiency—the claim that past-adapted information exhausts the systematic variation in a pre-event EEG endpoint—can be tested as a design-based, magnitude-qualified hypothesis rather than assumed from model fit. The construction commits a pre-event endpoint, a terminal CNV-like amplitude, before randomly assigning the delay to the imperative event; under the past-adapted factorisation, that later-assigned delay should not systematically order the committed endpoint or its frozen residual. A qualified material negative ordering is argued to support conditional insufficiency of the whole past-adapted class, while an adequately sensitive null supports a bounded affirmative conclusion calibrated by the pipeline's false-adequacy rate. The synthetic benchmark certifies false-adequacy boundaries of 15 uV/s for the assignment-isolation route and 30 uV/s for the sequential e-value route, in both directions. The design is presented as transferable to any setting where a statistic can be irrevocably committed before an exogenous label is generated.","feed_headline":"Post-endpoint randomisation makes 'the past explains it' testable","feed_subtitle":"A delay drawn after the endpoint is fixed becomes a probe; a surviving ordering rejects past-adapted sufficiency.","key_machinery":"The carrying mechanism is the past-adapted factorisation together with the post-endpoint randomisation boundary it describes. For each trial j, the later-assigned delay $\\tau_L^{(j)}$ is independent of the pre-assignment filtration $\\mathcal{F}_{t_1}^{(j)}$ given the declared randomisation stratum $R_j$; Proposition 1 converts this into the no-ordering statement $E[A_{\\mathrm{pre}} \\mid \\sigma(\\tau_L), R_j] = E[A_{\\mathrm{pre}} \\mid R_j]$ for any committed pre-assignment statistic. Around that core, a frozen label-blind comparator produces held-out residuals, the retained-sample condition $(R3^*)$ and frozen-comparator independence carry the exclusion into the analysed sample, and the non-compensatory decision rule with the false-adequacy benchmark turns the contrast into a magnitude-qualified classification.","core_discovery":"The central claim is that the post-endpoint randomised delay turns the past-adapted explanatory assumption into a testable no-ordering implication. Under the past-adapted factorisation, for each trial j the assigned delay $\\tau_L^{(j)}$ is conditionally independent of the pre-assignment filtration $\\mathcal{F}_{t_1}^{(j)}$ given the randomisation stratum $R_j$, and Proposition 1 derives that any committed pre-assignment statistic, in particular the endpoint $A_{\\mathrm{pre}}$, satisfies $E[A_{\\mathrm{pre}} \\mid \\sigma(\\tau_L), R_j] = E[A_{\\mathrm{pre}} \\mid R_j]$ almost surely before selection. The paper further shows that this pre-selection exclusion can be carried to the analysed sample through retained-sample delay-neutrality, retained-support positivity, and frozen-comparator independence, so that a qualified material negative slope of the frozen residual by assigned delay supports conditional insufficiency of the past-adapted class, while an adequately sensitive null supports a bounded affirmative conclusion. The discovery is therefore a design-based inference framework, Level II-A, that can reject or bound the adequacy of an entire class of past-adapted explanations without requiring a mechanism to be specified.","pith_inferences":["The same locked-score design could audit machine-learning pipelines: fix a model output before a randomised release, routing, or deployment label, and let a surviving association diagnose leakage or selection rather than model insufficiency.","The certified boundaries cover only the additive endpoint-level linear-injection family; nonlinear, time-varying, or subgroup-specific departures would require new calibration, so the affirmative null should be read as linear adequacy over the evaluated grid.","The gap between the 15 and 30 uV/s boundaries suggests the sequential route's conservative fold mixture costs sensitivity, and different e-value combining rules might narrow it in a testable extension.","For EEG, high-throughput committed statistics such as decoder outputs may achieve finer resolution than late CNV under the same design, making the affirmative null more informative in practice."],"forward_implications":["A clean, adequately sensitive null licenses a bounded affirmative conclusion: no linear assigned-delay ordering of the frozen residual at or above the certified magnitude remains for the declared endpoint, comparator, delay support, and regime.","A qualified material negative ordering supports conditional insufficiency of the entire past-adapted class for the tested contrast, without identifying any mechanism.","A qualified positive material slope blocks both directional support and a clean affirmative null and is routed to the opposite-direction diagnostic rather than to confirmatory class rejection.","No adequacy claim is licensed below the certified boundaries: 15 uV/s for assignment isolation and 30 uV/s for the sequential e-value route, in both directions.","Beyond EEG, the same boundary works wherever a statistic can be irrevocably committed before an exogenous label; without a declared alternative predicting ordering by that label, it remains only a pipeline-integrity negative control."],"supporting_citations":[{"why":"Supplies the negative-control exposure logic that makes the post-assignment delay a probe for confounding and bias.","marker":"[12]"},{"why":"Extends the negative-control toolkit used to interpret the assigned-delay label as a randomised negative-control exposure.","marker":"[13]"},{"why":"Underpins the severity-style interpretation of an adequately sensitive null as bounded positive evidence.","marker":"[19]"},{"why":"Provides the conditional randomisation test with finite replicates used by the assignment-isolation route.","marker":"[41]"},{"why":"Justifies the plus-one correction for Monte Carlo randomisation p-values in the assignment-isolation route.","marker":"[42]"},{"why":"Frames the structural selection-bias problem that retained-sample delay-neutrality and the collider diagnostic must address.","marker":"[49]"},{"why":"Quantifies collider-stratification bias, anchoring the endpoint-by-delay collider diagnostic used to block selection-limited outcomes.","marker":"[50]"},{"why":"Supplies sharp bounds used by the selection-sensitivity gate to map audited retention imbalance onto the slope scale.","marker":"[44]"},{"why":"Supplies worst-case non-monotone bounds used alongside the sharp bounds in the selection-sensitivity gate.","marker":"[45]"}],"fun_headline_variants":["Post-endpoint randomisation turns 'the past explains it' into a testable claim","A randomised delay after the endpoint probes whether past-adapted accounts suffice","Level II-A: making 'the past explains it' testable","Post-endpoint randomisation: a negative-control probe for past-adapted sufficiency"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire interpretation rests on retained-sample delay-neutrality and frozen-comparator independence—that after conditioning on the endpoint, covariates, stratum, and the frozen comparator object, neither retention nor the comparator itself creates an endpoint-delay association; these conditions cannot be established globally and are only qualified by audits and sensitivity analyses.","fun_headline_variants_meta":{"raw":{"variants":["Post-endpoint randomisation turns 'the past explains it' into a testable claim","A randomised delay after the endpoint probes whether past-adapted accounts suffice","Level II-A: making 'the past explains it' testable","Post-endpoint randomisation: a negative-control probe for past-adapted sufficiency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001042,"raw_usage":{"total_tokens":4446,"prompt_tokens":1072,"completion_tokens":3374,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":688,"completion_tokens_details":{"reasoning_tokens":3292}},"tokens_in":688,"tokens_out":3374,"duration_ms":23328,"temperature":1.0,"reasoning_tokens":3292,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:17:40.508627+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the full locked pipeline to a synthetic past-adapted generator in which inclusion is driven only by the endpoint-by-delay collider while marginal retention is balanced and all operational audits pass; the classifier must return selection-limited in essentially every dataset, as its own benchmark reports, so any independent run returning directional support would falsify the non-compensatory rule.","supporting_citations":[],"review_version":1}