REVIEW 6 minor 29 references
Testing the limits of past-adapted explanations by post-endpoint randomisation: anticipatory EEG as a worked case
T0 review · 0 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper argues that randomising the delay after committing a pre-event EEG endpoint lets a researcher test, rather than assume, whether past-adapted information is sufficient.
desk verdict A careful, well-scoped methods paper whose novel negative-control design is real; the unverifiable retained-sample neutrality is an honest epistemic limit, not a hidden flaw. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the past-adapted factorisation together with the post-endpoint randomisation boundary it describes. For each trial j, the later-assigned delay $\tau_L^{(j)}$ is independent of the pre-assignment filtration $\mathcal{F}_{t_1}^{(j)}$ given the declared randomisation stratum $R_j$; Proposition 1 converts this into the no-ordering statement $E[A_{\mathrm{pre}} \mid \sigma(\tau_L), R_j] = E[A_{\mathrm{pre}} \mid R_j]$ for any committed pre-assignment statistic. Around that core, a frozen label-blind comparator produces held-out residuals, the retained-sample condition $(R3^*)$ and frozen-comparator independence carry the exclusion into the analysed sample, and the non-compensatory decision rule with the false-adequacy benchmark turns the contrast into a magnitude-qualified classification.
What would settle it
Apply the full locked pipeline to a synthetic past-adapted generator in which inclusion is driven only by the endpoint-by-delay collider while marginal retention is balanced and all operational audits pass; the classifier must return selection-limited in essentially every dataset, as its own benchmark reports, so any independent run returning directional support would falsify the non-compensatory rule.
Extended reading notes
Core claim
The central claim is that the post-endpoint randomised delay turns the past-adapted explanatory assumption into a testable no-ordering implication. Under the past-adapted factorisation, for each trial j the assigned delay $\tau_L^{(j)}$ is conditionally independent of the pre-assignment filtration $\mathcal{F}_{t_1}^{(j)}$ given the randomisation stratum $R_j$, and Proposition 1 derives that any committed pre-assignment statistic, in particular the endpoint $A_{\mathrm{pre}}$, satisfies $E[A_{\mathrm{pre}} \mid \sigma(\tau_L), R_j] = E[A_{\mathrm{pre}} \mid R_j]$ almost surely before selection. The paper further shows that this pre-selection exclusion can be carried to the analysed sample through retained-sample delay-neutrality, retained-support positivity, and frozen-comparator independence, so that a qualified material negative slope of the frozen residual by assigned delay supports conditional insufficiency of the past-adapted class, while an adequately sensitive null supports a bounded affirmative conclusion. The discovery is therefore a design-based inference framework, Level II-A, that can reject or bound the adequacy of an entire class of past-adapted explanations without requiring a mechanism to be specified.
Load-bearing premise
The entire interpretation rests on retained-sample delay-neutrality and frozen-comparator independence—that after conditioning on the endpoint, covariates, stratum, and the frozen comparator object, neither retention nor the comparator itself creates an endpoint-delay association; these conditions cannot be established globally and are only qualified by audits and sensitivity analyses.
Editorial extensions
If this is right
- A clean, adequately sensitive null licenses a bounded affirmative conclusion: no linear assigned-delay ordering of the frozen residual at or above the certified magnitude remains for the declared endpoint, comparator, delay support, and regime.
- A qualified material negative ordering supports conditional insufficiency of the entire past-adapted class for the tested contrast, without identifying any mechanism.
- A qualified positive material slope blocks both directional support and a clean affirmative null and is routed to the opposite-direction diagnostic rather than to confirmatory class rejection.
- No adequacy claim is licensed below the certified boundaries: 15 uV/s for assignment isolation and 30 uV/s for the sequential e-value route, in both directions.
- Beyond EEG, the same boundary works wherever a statistic can be irrevocably committed before an exogenous label; without a declared alternative predicting ordering by that label, it remains only a pipeline-integrity negative control.
Reading between the lines
- The same locked-score design could audit machine-learning pipelines: fix a model output before a randomised release, routing, or deployment label, and let a surviving association diagnose leakage or selection rather than model insufficiency.
- The certified boundaries cover only the additive endpoint-level linear-injection family; nonlinear, time-varying, or subgroup-specific departures would require new calibration, so the affirmative null should be read as linear adequacy over the evaluated grid.
- The gap between the 15 and 30 uV/s boundaries suggests the sequential route's conservative fold mixture costs sensitivity, and different e-value combining rules might narrow it in a testable extension.
- For EEG, high-throughput committed statistics such as decoder outputs may achieve finer resolution than late CNV under the same design, making the affirmative null more informative in practice.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Level II-A, a design-based inference framework for testing whether a predictive model class is informationally sufficient for a committed endpoint, using post-endpoint randomisation as a negative-control probe. In the worked anticipatory-EEG instantiation, a pre-event summary A_pre is committed at time t1, after which the delay to the imperative event is randomised within a declared stratum. Under the declared past-adapted factorisation, any temporally admissible pre-assignment statistic is conditionally independent of the assigned delay given the stratum, yielding a no-ordering implication. The authors show that carrying this exclusion to the retained, frozen-residual analysis requires three additional conditions: retained-sample delay-neutrality (R3*, Eq. 3), retained-support positivity, and frozen-comparator independence (S5). A non-compensatory decision rule classifies outcomes as supported negative departure, forward-only adequate affirmative null, diagnostic failure, selection-limited, opposite-direction, or inconclusive, with the affirmative null calibrated by simulation-based false-adequacy boundaries of 15 µV/s (assignment-isolation route) and 30 µV/s (sequential e-value route). The paper reports no human EEG data and relies on a seven-generator synthetic benchmark with a certified machine-readable run.
Significance. If the framework holds up, it is a valuable contribution: it turns the common explanatory claim that 'the past explains the endpoint' into a magnitude-qualified, falsifiable statement, and it does so without requiring a mechanistic model. The central derivation, Proposition 1 and Lemma S1, is logically sound given the stated assumptions, and the authors are unusually explicit about what those assumptions do and do not establish. Particular strengths are the non-compensatory decision architecture, the route-specific calibration separating assignment isolation from sequential e-value inference under carryover, the leakage-safe causal preprocessing requirements, the prospective locking of the endpoint and comparator, the comprehensive synthetic benchmark with archived code and certified run outputs, and the explicit acknowledgement that R3* and S5 are sufficient conditions that cannot be globally certified.
minor comments (6)
- [Abstract] The sentence 'leakage-safe preprocessing, a frozen label-blind comparator and retained-sample qualifications carry the exclusion to the confirmatory residual' could be read as implying that R3* and S5 are operational achievements; consider adding the qualifier 'under sufficient conditions that are qualified, not established, by the audits' to match the careful wording in Section 3.2.
- [Section 3.2, Eq. (3) and SI §1.4] The manuscript already states prominently that randomisation alone does not establish (R3*) and that (S5) is not implied by the pre-selection factorisation; consider placing a short 'envelope of interpretation' box near Eq. (3) so that the conditional nature of the central claim is visually impossible to miss in secondary citations.
- [Figure 3] In panel (a), the annotation 'slope -- → null' is unclear; I recommend writing 'slope is not significantly negative → classified as forward-only adequate' in plain text.
- [Table 1] The outcome name 'forward-only adequate' is potentially misleading because it denotes a bounded affirmative null within the evaluated linear injection family, not global adequacy; consider renaming it 'qualified affirmative null' in the table and text to match the careful interpretation given in Section 5.2.
- [Section 6.1] The certified false-adequacy boundaries (15 and 30 µV/s) are easily confused with the resolution floor β_min from Eq. (14); I suggest stating explicitly in Section 6.1 that δ*_r,d are classifier-level operating-characteristic boundaries and are distinct from the per-participant resolution floor.
- [SI §1.3, proof of Lemma S1] There is a typographical error in the sentence 'Because the factorisation holds for every bounded measurable test function φ and every integrable prospectively fixed h, division by this positive conditional probability yields...' where the period is missing after 'yields'; the sentence should end with a period.
Circularity Check
No significant circularity: the no-ordering implication is derived from declared premises, and the benchmark boundaries are simulation-based operating characteristics, not fitted inputs.
full rationale
The paper's central derivation is deductive and self-contained. Proposition 1 and Lemma S1 (Eqs. 1-2, Eq. 3, SI Lemma S1) derive the retained-sample exclusion of the assigned delay from declared conditions: the past-adapted factorisation, the operational scheduler law, temporal integrity, retained-sample delay-neutrality (R3*), and frozen-comparator independence (S5). These are stated as sufficient structural conditions, and the paper explicitly acknowledges that they are not established globally ('Condition (3) provides a sufficient criterion for retained-sample exclusion; alternative sufficient conditions may exist, and randomisation alone leaves (R3*) unestablished'). This is a limitation of the design, not circular reasoning: the inference is conditional on premises that the design audits but cannot prove. The false-adequacy boundaries (15 uV/s and 30 uV/s) are Monte Carlo operating characteristics of the full locked pipeline under known injected departures (Section 6.1, SI S9), not parameters fitted to the confirmatory target, so the 'prediction' is not equivalent to an input by construction. The only self-reference is citation of the authors' own public benchmark pipeline (ref. [56]) as the source of the certified run, which is a reproducibility artefact and not a load-bearing theoretical premise; the benchmark is externally inspectable and archived. No uniqueness theorem, ansatz, or known-result renaming is imported from the authors' prior work. The inferential limitation concerning R3* and possible selection or collider failures is explicitly scoped and is a correctness/design limitation rather than a circular step.
Assumptions & free parameters
free parameters (2)
- Resolution floor beta_min (with kappa, sigma_resid, n_ret in Eq. 14) =
kappa = 2 (illustrative), sigma_resid ~ 4 uV, n_ret ~ 300; beta_min ~ 2.7 uV/s in the worked example
- Benchmark constants: injected magnitude grid, residual scale, participant count, delay support =
grid {5,10,15,20,30,40,50,60,75,90} uV/s; sigma ~ 1 uV; P=24; Nmin=10; T0=20 ms
assumptions (6)
- standard math Standard measure-theoretic conditional independence and martingale properties
- domain assumption Existence of the pre-assignment filtration F_t1 and the declared scheduler law P_sch
- ad hoc to paper Past-adapted factorisation (S1): tau_L independent of F_t1 given R_j
- ad hoc to paper Retained-sample delay-neutrality (R3*, Eq. 3): S independent of tau_L given (Z, R, G_frz)
- ad hoc to paper Frozen-comparator independence (S5): tau_L independent of Z given (R, G_frz)
- domain assumption Operational scheduler and concealment (R1)
Cite this review
Pith. "Pith review of Testing the limits of past-adapted explanations by post-endpoint randomisation: anticipatory EEG as a worked case." pith.science (2026). https://pith.science/paper/5SVLYAM6
@misc{pith2026260812072,
author = {Pith},
title = {Pith review of: Testing the limits of past-adapted explanations by post-endpoint randomisation: anticipatory EEG as a worked case},
year = {2026},
howpublished = {\url{https://pith.science/paper/5SVLYAM6}},
note = {Machine review of arXiv:2608.12072}
}
abstract
A predictive model can fit its data even when its information set is insufficient; fit alone cannot establish sufficiency. This Perspective introduces Level II-A, a new design-based inference framework to test this distinction, illustrated in anticipatory EEG using contingent negative variation. A pre-event endpoint is committed before the delay to the imperative event is randomised. That later-assigned delay thereby becomes a negative-control probe of whether past-adapted information was sufficient for an already committed result. Under the past-adapted factorisation, accounts using only pre-commitment information cannot systematically order the endpoint by that delay. Leakage-safe preprocessing, a frozen label-blind comparator and retained-sample qualifications carry the exclusion to the confirmatory residual. A qualified material negative ordering supports conditional insufficiency without identifying a mechanism; an adequately sensitive null supports a bounded affirmative conclusion calibrated by the pipeline's false-adequacy rate. A non-compensatory rule separates these from diagnostic failure, selection-limited, opposite-direction and inconclusive outcomes. No human EEG data are analysed. In the synthetic benchmark, grid-based false-adequacy boundaries are $15\,\mu\mathrm{V\,s^{-1}}$ for assignment isolation and $30\,\mu\mathrm{V\,s^{-1}}$ for the sequential e-value route, in both directions. The design transfers wherever endpoint commitment precedes an exogenous label, probing sufficiency only where a declared alternative predicts ordering by it. It turns "the past explains it" from a working explanatory assumption into a magnitude-qualified, testable claim.
Figures
Reference graph
Works this paper leans on
-
[1]
C. E. Frangakis and D. B. Rubin. Principal stratification in causal inference.Biometrics, 58(1):21–29,
-
[2]
X. Lu, L. Shi, H. Liu, and P. Ding. Conditional cross-fitting for unbiased machine-learning-assisted covariate adjustment in randomized experiments, 2025. URLhttps://arxiv.org/abs/2508.15664
arXiv 2025
-
[3]
D. A. Freedman. On regression adjustments in experiments with several treatments.Annals of Applied Statistics, 2:176–196, 2008. doi:10.1214/07-AOAS143
-
[4]
W. Lin. Agnostic notes on regression adjustments to experimental data: reexamining freedman’s cri- tique. Annals of Applied Statistics , 7:295–318, 2013. doi:10.1214/12-AOAS583
-
[5]
A. Zhao and P. Ding. Covariate-adjusted Fisher randomization tests for the average treatment effect. Journal of Econometrics , 225:278–294, 2021. doi:10.1016/j.jeconom.2021.04.007
-
[6]
K. D. Harris and K. J. Miller. Conditional randomization tests for behavioral and neural time series,
-
[7]
V. Vovk and R. Wang. E-values: calibration, combination and applications.Annals of Statistics , 49: 1736–1754, 2021. doi:10.1214/20-AOS2020
-
[8]
G. Shafer. Testing by betting: a strategy for statistical and scientific communication.Journal of the Royal Statistical Society: Series A , 184:407–431, 2021. doi:10.1111/rssa.12647
Show all 29 references
-
[9]
Ramdas, P
A. Ramdas, P. Grünwald, V. Vovk, and G. Shafer. Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38:576–601, 2023. doi:10.1214/23-STS894
2023 doi
-
[10]
Grünwald, R
P. Grünwald, R. de Heide, and W. M. Koolen. Safe testing.Journal of the Royal Statistical Society: Series B, 86:1091–1128, 2024. doi:10.1093/jrsssb/qkae011
2024 doi
-
[11]
Phipson and G
B. Phipson and G. K. Smyth. Permutationp-values should never be zero: calculating exactp-values when permutations are randomly drawn.Statistical Applications in Genetics and Molecular Biology , 9: Article 39, 2010. doi:10.2202/1544-6115.1585
2010
-
[12]
P. Hall. The Bootstrap and Edgeworth Expansion . Springer, New York, NY, 1992. doi:10.1007/978-1 -4612-4384-7
1992 doi
-
[13]
T. J. DiCiccio and B. Efron. Bootstrap confidence intervals. Statistical Science, 11:189–228, 1996. doi:10.1214/ss/1032280214
1996
-
[14]
Efron and R
B. Efron and R. J. Tibshirani. An Introduction to the Bootstrap . Chapman & Hall, New York, NY,
-
[15]
Lipsitch, E
M. Lipsitch, E. Tchetgen Tchetgen, and T. Cohen. Negative controls: a tool for detecting confounding and bias in observational studies.Epidemiology, 21:383–388, 2010. doi:10.1097/EDE.0b013e3181d61e eb
2010 doi
-
[16]
X. Shi, W. Miao, and E. Tchetgen Tchetgen. A selective review of negative control methods in epidemi- ology. Current Epidemiology Reports, 7:190–202, 2020. doi:10.1007/s40471-020-00243-4
2020 doi
-
[17]
T. J. VanderWeele and P. Ding. Sensitivity analysis in observational research: introducing the E-value. Annals of Internal Medicine , 167:268–274, 2017. doi:10.7326/M16-2607
2017 doi
-
[18]
L. H. Smith and T. J. VanderWeele. Bounding bias due to selection.Epidemiology, 30:509–516, 2019. doi:10.1097/EDE.0000000000001032
2019 doi
-
[19]
D. S. Lee. Training, wages, and sample selection: estimating sharp bounds on treatment effects.Review of Economic Studies , 76:1071–1102, 2009. doi:10.1111/j.1467-937X.2009.00536.x
2009 arXiv
-
[20]
C. F. Manski. Nonparametric bounds on treatment effects.American Economic Review, 80:319–323, 1990
1990
-
[21]
P. R. Rosenbaum. Observational Studies. Springer, New York, NY, 2 edition, 2002. doi:10.1007/97 8-1-4757-3692-2
2002 doi
-
[22]
M. A. Hernán, S. Hernández-Díaz, and J. M. Robins. A structural approach to selection bias.Epidemi- ology, 15:615–625, 2004. doi:10.1097/01.ede.0000135174.63482.43
2004
-
[23]
Greenland
S. Greenland. Quantifying biases in causal models: classical confounding versus collider-stratification bias. Epidemiology, 14:300–306, 2003. doi:10.1097/01.EDE.0000042804.12056.6C
2003
-
[24]
W. G. Walter, R. Cooper, V. J. Aldridge, W. C. McCallum, and A. L. Winter. Contingent negative variation: an electric sign of sensori-motor association and expectancy in the human brain.Nature, 203: 380–384, 1964. doi:10.1038/203380a0
1964 doi
-
[25]
Negativeslowwavesasindicesofanticipation: the Bereitschaftspotential, the contingent negative variation, and the stimulus-preceding negativity
C.H.M.Brunia, G.J.M.vanBoxtel, andK.B.E.Böcker. Negativeslowwavesasindicesofanticipation: the Bereitschaftspotential, the contingent negative variation, and the stimulus-preceding negativity. In E. S. Kappenman and S. J. Luck, editors,The Oxford Handbook of Event-Related Poten...
2012
-
[26]
A. C. Nobre and F. van Ede. Anticipated moments: temporal structure in attention.Nature Reviews Neuroscience, 19:34–48, 2018. doi:10.1038/nrn.2017.141. 56
2018 doi
-
[1994]
doi:10.1201/9780429246593
-
[2002]
doi:10.1111/j.0006-341X.2002.00021.x
2002 arXiv
-
[2023]
URL https://arxiv.org/abs/2311.03554. 55
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.