Pith. sign in

REVIEW 6 minor 29 references

Testing the limits of past-adapted explanations by post-endpoint randomisation: anticipatory EEG as a worked case

T0 review · 0 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The paper argues that randomising the delay after committing a pre-event EEG endpoint lets a researcher test, rather than assume, whether past-adapted information is sufficient.

desk verdict A careful, well-scoped methods paper whose novel negative-control design is real; the unverifiable retained-sample neutrality is an honest epistemic limit, not a hidden flaw. read the letter →

arxiv 2608.12072 v1 pith:5SVLYAM6 submitted 2026-08-12 stat.ME q-bio.NC

classification stat.MEq-bio.NC MSC 62F0362G1062K1062P10
keywords anticipatoryEEGcontingentnegativevariationtemporalexpectationpost-endpointrandomisationcontrolsinferencedesign-basedinformationalsufficiency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper attempts to establish that informational sufficiency—the claim that past-adapted information exhausts the systematic variation in a pre-event EEG endpoint—can be tested as a design-based, magnitude-qualified hypothesis rather than assumed from model fit. The construction commits a pre-event endpoint, a terminal CNV-like amplitude, before randomly assigning the delay to the imperative event; under the past-adapted factorisation, that later-assigned delay should not systematically order the committed endpoint or its frozen residual. A qualified material negative ordering is argued to support conditional insufficiency of the whole past-adapted class, while an adequately sensitive null supports a bounded affirmative conclusion calibrated by the pipeline's false-adequacy rate. The synthetic benchmark certifies false-adequacy boundaries of 15 uV/s for the assignment-isolation route and 30 uV/s for the sequential e-value route, in both directions. The design is presented as transferable to any setting where a statistic can be irrevocably committed before an exogenous label is generated.

What carries the argument

The carrying mechanism is the past-adapted factorisation together with the post-endpoint randomisation boundary it describes. For each trial j, the later-assigned delay $\tau_L^{(j)}$ is independent of the pre-assignment filtration $\mathcal{F}_{t_1}^{(j)}$ given the declared randomisation stratum $R_j$; Proposition 1 converts this into the no-ordering statement $E[A_{\mathrm{pre}} \mid \sigma(\tau_L), R_j] = E[A_{\mathrm{pre}} \mid R_j]$ for any committed pre-assignment statistic. Around that core, a frozen label-blind comparator produces held-out residuals, the retained-sample condition $(R3^*)$ and frozen-comparator independence carry the exclusion into the analysed sample, and the non-compensatory decision rule with the false-adequacy benchmark turns the contrast into a magnitude-qualified classification.

What would settle it

Apply the full locked pipeline to a synthetic past-adapted generator in which inclusion is driven only by the endpoint-by-delay collider while marginal retention is balanced and all operational audits pass; the classifier must return selection-limited in essentially every dataset, as its own benchmark reports, so any independent run returning directional support would falsify the non-compensatory rule.

Watch

Extended reading notes

Core claim

The central claim is that the post-endpoint randomised delay turns the past-adapted explanatory assumption into a testable no-ordering implication. Under the past-adapted factorisation, for each trial j the assigned delay $\tau_L^{(j)}$ is conditionally independent of the pre-assignment filtration $\mathcal{F}_{t_1}^{(j)}$ given the randomisation stratum $R_j$, and Proposition 1 derives that any committed pre-assignment statistic, in particular the endpoint $A_{\mathrm{pre}}$, satisfies $E[A_{\mathrm{pre}} \mid \sigma(\tau_L), R_j] = E[A_{\mathrm{pre}} \mid R_j]$ almost surely before selection. The paper further shows that this pre-selection exclusion can be carried to the analysed sample through retained-sample delay-neutrality, retained-support positivity, and frozen-comparator independence, so that a qualified material negative slope of the frozen residual by assigned delay supports conditional insufficiency of the past-adapted class, while an adequately sensitive null supports a bounded affirmative conclusion. The discovery is therefore a design-based inference framework, Level II-A, that can reject or bound the adequacy of an entire class of past-adapted explanations without requiring a mechanism to be specified.

Load-bearing premise

The entire interpretation rests on retained-sample delay-neutrality and frozen-comparator independence—that after conditioning on the endpoint, covariates, stratum, and the frozen comparator object, neither retention nor the comparator itself creates an endpoint-delay association; these conditions cannot be established globally and are only qualified by audits and sensitivity analyses.

Editorial extensions

If this is right

  • A clean, adequately sensitive null licenses a bounded affirmative conclusion: no linear assigned-delay ordering of the frozen residual at or above the certified magnitude remains for the declared endpoint, comparator, delay support, and regime.
  • A qualified material negative ordering supports conditional insufficiency of the entire past-adapted class for the tested contrast, without identifying any mechanism.
  • A qualified positive material slope blocks both directional support and a clean affirmative null and is routed to the opposite-direction diagnostic rather than to confirmatory class rejection.
  • No adequacy claim is licensed below the certified boundaries: 15 uV/s for assignment isolation and 30 uV/s for the sequential e-value route, in both directions.
  • Beyond EEG, the same boundary works wherever a statistic can be irrevocably committed before an exogenous label; without a declared alternative predicting ordering by that label, it remains only a pipeline-integrity negative control.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same locked-score design could audit machine-learning pipelines: fix a model output before a randomised release, routing, or deployment label, and let a surviving association diagnose leakage or selection rather than model insufficiency.
  • The certified boundaries cover only the additive endpoint-level linear-injection family; nonlinear, time-varying, or subgroup-specific departures would require new calibration, so the affirmative null should be read as linear adequacy over the evaluated grid.
  • The gap between the 15 and 30 uV/s boundaries suggests the sequential route's conservative fold mixture costs sensitivity, and different e-value combining rules might narrow it in a testable extension.
  • For EEG, high-throughput committed statistics such as decoder outputs may achieve finer resolution than late CNV under the same design, making the affirmative null more informative in practice.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 6 minor

Summary. The paper proposes Level II-A, a design-based inference framework for testing whether a predictive model class is informationally sufficient for a committed endpoint, using post-endpoint randomisation as a negative-control probe. In the worked anticipatory-EEG instantiation, a pre-event summary A_pre is committed at time t1, after which the delay to the imperative event is randomised within a declared stratum. Under the declared past-adapted factorisation, any temporally admissible pre-assignment statistic is conditionally independent of the assigned delay given the stratum, yielding a no-ordering implication. The authors show that carrying this exclusion to the retained, frozen-residual analysis requires three additional conditions: retained-sample delay-neutrality (R3*, Eq. 3), retained-support positivity, and frozen-comparator independence (S5). A non-compensatory decision rule classifies outcomes as supported negative departure, forward-only adequate affirmative null, diagnostic failure, selection-limited, opposite-direction, or inconclusive, with the affirmative null calibrated by simulation-based false-adequacy boundaries of 15 µV/s (assignment-isolation route) and 30 µV/s (sequential e-value route). The paper reports no human EEG data and relies on a seven-generator synthetic benchmark with a certified machine-readable run.

Significance. If the framework holds up, it is a valuable contribution: it turns the common explanatory claim that 'the past explains the endpoint' into a magnitude-qualified, falsifiable statement, and it does so without requiring a mechanistic model. The central derivation, Proposition 1 and Lemma S1, is logically sound given the stated assumptions, and the authors are unusually explicit about what those assumptions do and do not establish. Particular strengths are the non-compensatory decision architecture, the route-specific calibration separating assignment isolation from sequential e-value inference under carryover, the leakage-safe causal preprocessing requirements, the prospective locking of the endpoint and comparator, the comprehensive synthetic benchmark with archived code and certified run outputs, and the explicit acknowledgement that R3* and S5 are sufficient conditions that cannot be globally certified.

minor comments (6)
  1. [Abstract] The sentence 'leakage-safe preprocessing, a frozen label-blind comparator and retained-sample qualifications carry the exclusion to the confirmatory residual' could be read as implying that R3* and S5 are operational achievements; consider adding the qualifier 'under sufficient conditions that are qualified, not established, by the audits' to match the careful wording in Section 3.2.
  2. [Section 3.2, Eq. (3) and SI §1.4] The manuscript already states prominently that randomisation alone does not establish (R3*) and that (S5) is not implied by the pre-selection factorisation; consider placing a short 'envelope of interpretation' box near Eq. (3) so that the conditional nature of the central claim is visually impossible to miss in secondary citations.
  3. [Figure 3] In panel (a), the annotation 'slope -- → null' is unclear; I recommend writing 'slope is not significantly negative → classified as forward-only adequate' in plain text.
  4. [Table 1] The outcome name 'forward-only adequate' is potentially misleading because it denotes a bounded affirmative null within the evaluated linear injection family, not global adequacy; consider renaming it 'qualified affirmative null' in the table and text to match the careful interpretation given in Section 5.2.
  5. [Section 6.1] The certified false-adequacy boundaries (15 and 30 µV/s) are easily confused with the resolution floor β_min from Eq. (14); I suggest stating explicitly in Section 6.1 that δ*_r,d are classifier-level operating-characteristic boundaries and are distinct from the per-participant resolution floor.
  6. [SI §1.3, proof of Lemma S1] There is a typographical error in the sentence 'Because the factorisation holds for every bounded measurable test function φ and every integrable prospectively fixed h, division by this positive conditional probability yields...' where the period is missing after 'yields'; the sentence should end with a period.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the no-ordering implication is derived from declared premises, and the benchmark boundaries are simulation-based operating characteristics, not fitted inputs.

full rationale

The paper's central derivation is deductive and self-contained. Proposition 1 and Lemma S1 (Eqs. 1-2, Eq. 3, SI Lemma S1) derive the retained-sample exclusion of the assigned delay from declared conditions: the past-adapted factorisation, the operational scheduler law, temporal integrity, retained-sample delay-neutrality (R3*), and frozen-comparator independence (S5). These are stated as sufficient structural conditions, and the paper explicitly acknowledges that they are not established globally ('Condition (3) provides a sufficient criterion for retained-sample exclusion; alternative sufficient conditions may exist, and randomisation alone leaves (R3*) unestablished'). This is a limitation of the design, not circular reasoning: the inference is conditional on premises that the design audits but cannot prove. The false-adequacy boundaries (15 uV/s and 30 uV/s) are Monte Carlo operating characteristics of the full locked pipeline under known injected departures (Section 6.1, SI S9), not parameters fitted to the confirmatory target, so the 'prediction' is not equivalent to an input by construction. The only self-reference is citation of the authors' own public benchmark pipeline (ref. [56]) as the source of the certified run, which is a reproducibility artefact and not a load-bearing theoretical premise; the benchmark is externally inspectable and archived. No uniqueness theorem, ansatz, or known-result renaming is imported from the authors' prior work. The inferential limitation concerning R3* and possible selection or collider failures is explicitly scoped and is a correctness/design limitation rather than a circular step.

Assumptions & free parameters 2 free parameters · 6 assumptions · 0 invented entities

The framework rests on standard measure-theoretic probability and on a set of explicitly stated design conditions (R1, R2, R3*, and frozen-comparator independence). These are acknowledged as conditions to be audited or qualified, not derived. No parameters are fitted to data; the only free choices are design constants (resolution floor, minimum participant count, benchmark grid) that are fixed prospectively.

free parameters (2)
  • Resolution floor beta_min (with kappa, sigma_resid, n_ret in Eq. 14) = kappa = 2 (illustrative), sigma_resid ~ 4 uV, n_ret ~ 300; beta_min ~ 2.7 uV/s in the worked example
    Equation (14) defines the threshold for a material slope. It is set prospectively from label-blind pilot estimates and is not fitted to any outcome. The choice of kappa and the pilot estimates are user-specified.
  • Benchmark constants: injected magnitude grid, residual scale, participant count, delay support = grid {5,10,15,20,30,40,50,60,75,90} uV/s; sigma ~ 1 uV; P=24; Nmin=10; T0=20 ms
    These constants define the synthetic benchmark and are chosen to stress-test the pipeline. They are not fitted to data and are clearly labeled as simulation parameters.
assumptions (6)
  • standard math Standard measure-theoretic conditional independence and martingale properties
    Used throughout the proofs (Proposition S1, Lemma S1) to derive the no-ordering consequence from independence.
  • domain assumption Existence of the pre-assignment filtration F_t1 and the declared scheduler law P_sch
    The design assumes trials are indexed within a probability space with a filtration representing all pre-t1 information and a registered scheduler law. This is a modeling assumption, but it is standard.
  • ad hoc to paper Past-adapted factorisation (S1): tau_L independent of F_t1 given R_j
    This is the defining statistical restriction of the past-adapted class. It is not proven; it is the property being tested. The framework assumes that if this factorisation holds, the no-ordering implication follows.
  • ad hoc to paper Retained-sample delay-neutrality (R3*, Eq. 3): S independent of tau_L given (Z, R, G_frz)
    A sufficient condition for the retained-sample exclusion. It is unprovable in general and is only qualified by audits and sensitivity analyses.
  • ad hoc to paper Frozen-comparator independence (S5): tau_L independent of Z given (R, G_frz)
    Additional condition needed because conditioning on the fitted comparator object can reintroduce dependence. The paper notes it is route-specific and must be operationally qualified.
  • domain assumption Operational scheduler and concealment (R1)
    The framework assumes the scheduler faithfully generates the delay after endpoint commitment using only declared inputs and conceals it from confirmatory analyses. Audits check this, but it remains an assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Testing the limits of past-adapted explanations by post-endpoint randomisation: anticipatory EEG as a worked case." pith.science (2026). https://pith.science/paper/5SVLYAM6

@misc{pith2026260812072,
  author       = {Pith},
  title        = {Pith review of: Testing the limits of past-adapted explanations by post-endpoint randomisation: anticipatory EEG as a worked case},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5SVLYAM6}},
  note         = {Machine review of arXiv:2608.12072}
}
abstract

A predictive model can fit its data even when its information set is insufficient; fit alone cannot establish sufficiency. This Perspective introduces Level II-A, a new design-based inference framework to test this distinction, illustrated in anticipatory EEG using contingent negative variation. A pre-event endpoint is committed before the delay to the imperative event is randomised. That later-assigned delay thereby becomes a negative-control probe of whether past-adapted information was sufficient for an already committed result. Under the past-adapted factorisation, accounts using only pre-commitment information cannot systematically order the endpoint by that delay. Leakage-safe preprocessing, a frozen label-blind comparator and retained-sample qualifications carry the exclusion to the confirmatory residual. A qualified material negative ordering supports conditional insufficiency without identifying a mechanism; an adequately sensitive null supports a bounded affirmative conclusion calibrated by the pipeline's false-adequacy rate. A non-compensatory rule separates these from diagnostic failure, selection-limited, opposite-direction and inconclusive outcomes. No human EEG data are analysed. In the synthetic benchmark, grid-based false-adequacy boundaries are $15\,\mu\mathrm{V\,s^{-1}}$ for assignment isolation and $30\,\mu\mathrm{V\,s^{-1}}$ for the sequential e-value route, in both directions. The design transfers wherever endpoint commitment precedes an exogenous label, probing sufficiency only where a declared alternative predicts ordering by it. It turns "the past explains it" from a working explanatory assumption into a magnitude-qualified, testable claim.

Figures

Figures reproduced from arXiv: 2608.12072 by the authors.

Figure 1
Figure 1. Post-endpoint randomisation boundary. The single-trial endpoint [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Schematic design and selection-threat graph for the committed-endpoint test. The de [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Synthetic validation of the Level II-A pipeline using simulated data only. The reference [PITH_FULL_IMAGE:figures/full_fig_p025_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Analysis-access schematic, not within-trial event chronology or executable classifier prece [PITH_FULL_IMAGE:figures/full_fig_p028_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

29 extracted references · 9 canonical work pages

  1. [1]

    C. E. Frangakis and D. B. Rubin. Principal stratification in causal inference.Biometrics, 58(1):21–29,

  2. [2]

    X. Lu, L. Shi, H. Liu, and P. Ding. Conditional cross-fitting for unbiased machine-learning-assisted covariate adjustment in randomized experiments, 2025. URLhttps://arxiv.org/abs/2508.15664

  3. [3]

    D. A. Freedman. On regression adjustments in experiments with several treatments.Annals of Applied Statistics, 2:176–196, 2008. doi:10.1214/07-AOAS143

  4. [4]

    W. Lin. Agnostic notes on regression adjustments to experimental data: reexamining freedman’s cri- tique. Annals of Applied Statistics , 7:295–318, 2013. doi:10.1214/12-AOAS583

  5. [5]

    Zhao and P

    A. Zhao and P. Ding. Covariate-adjusted Fisher randomization tests for the average treatment effect. Journal of Econometrics , 225:278–294, 2021. doi:10.1016/j.jeconom.2021.04.007

  6. [6]

    K. D. Harris and K. J. Miller. Conditional randomization tests for behavioral and neural time series,

  7. [7]

    Vovk and R

    V. Vovk and R. Wang. E-values: calibration, combination and applications.Annals of Statistics , 49: 1736–1754, 2021. doi:10.1214/20-AOS2020

  8. [8]

    G. Shafer. Testing by betting: a strategy for statistical and scientific communication.Journal of the Royal Statistical Society: Series A , 184:407–431, 2021. doi:10.1111/rssa.12647

Show all 29 references
  1. [9]

    Ramdas, P

    A. Ramdas, P. Grünwald, V. Vovk, and G. Shafer. Game-theoretic statistics and safe anytime-valid inference. Statistical Science, 38:576–601, 2023. doi:10.1214/23-STS894

  2. [10]

    Grünwald, R

    P. Grünwald, R. de Heide, and W. M. Koolen. Safe testing.Journal of the Royal Statistical Society: Series B, 86:1091–1128, 2024. doi:10.1093/jrsssb/qkae011

  3. [11]

    Phipson and G

    B. Phipson and G. K. Smyth. Permutationp-values should never be zero: calculating exactp-values when permutations are randomly drawn.Statistical Applications in Genetics and Molecular Biology , 9: Article 39, 2010. doi:10.2202/1544-6115.1585

  4. [12]

    P. Hall. The Bootstrap and Edgeworth Expansion . Springer, New York, NY, 1992. doi:10.1007/978-1 -4612-4384-7

  5. [13]

    T. J. DiCiccio and B. Efron. Bootstrap confidence intervals. Statistical Science, 11:189–228, 1996. doi:10.1214/ss/1032280214

  6. [14]

    Efron and R

    B. Efron and R. J. Tibshirani. An Introduction to the Bootstrap . Chapman & Hall, New York, NY,

  7. [15]

    Lipsitch, E

    M. Lipsitch, E. Tchetgen Tchetgen, and T. Cohen. Negative controls: a tool for detecting confounding and bias in observational studies.Epidemiology, 21:383–388, 2010. doi:10.1097/EDE.0b013e3181d61e eb

  8. [16]

    X. Shi, W. Miao, and E. Tchetgen Tchetgen. A selective review of negative control methods in epidemi- ology. Current Epidemiology Reports, 7:190–202, 2020. doi:10.1007/s40471-020-00243-4

  9. [17]

    T. J. VanderWeele and P. Ding. Sensitivity analysis in observational research: introducing the E-value. Annals of Internal Medicine , 167:268–274, 2017. doi:10.7326/M16-2607

  10. [18]

    L. H. Smith and T. J. VanderWeele. Bounding bias due to selection.Epidemiology, 30:509–516, 2019. doi:10.1097/EDE.0000000000001032

  11. [19]

    D. S. Lee. Training, wages, and sample selection: estimating sharp bounds on treatment effects.Review of Economic Studies , 76:1071–1102, 2009. doi:10.1111/j.1467-937X.2009.00536.x

  12. [20]

    C. F. Manski. Nonparametric bounds on treatment effects.American Economic Review, 80:319–323, 1990

  13. [21]

    P. R. Rosenbaum. Observational Studies. Springer, New York, NY, 2 edition, 2002. doi:10.1007/97 8-1-4757-3692-2

  14. [22]

    M. A. Hernán, S. Hernández-Díaz, and J. M. Robins. A structural approach to selection bias.Epidemi- ology, 15:615–625, 2004. doi:10.1097/01.ede.0000135174.63482.43

  15. [23]

    Greenland

    S. Greenland. Quantifying biases in causal models: classical confounding versus collider-stratification bias. Epidemiology, 14:300–306, 2003. doi:10.1097/01.EDE.0000042804.12056.6C

  16. [24]

    W. G. Walter, R. Cooper, V. J. Aldridge, W. C. McCallum, and A. L. Winter. Contingent negative variation: an electric sign of sensori-motor association and expectancy in the human brain.Nature, 203: 380–384, 1964. doi:10.1038/203380a0

  17. [25]

    Negativeslowwavesasindicesofanticipation: the Bereitschaftspotential, the contingent negative variation, and the stimulus-preceding negativity

    C.H.M.Brunia, G.J.M.vanBoxtel, andK.B.E.Böcker. Negativeslowwavesasindicesofanticipation: the Bereitschaftspotential, the contingent negative variation, and the stimulus-preceding negativity. In E. S. Kappenman and S. J. Luck, editors,The Oxford Handbook of Event-Related Poten...

  18. [26]

    A. C. Nobre and F. van Ede. Anticipated moments: temporal structure in attention.Nature Reviews Neuroscience, 19:34–48, 2018. doi:10.1038/nrn.2017.141. 56

  19. [1994]

    doi:10.1201/9780429246593

  20. [2002]

    doi:10.1111/j.0006-341X.2002.00021.x

  21. [2023]

    URL https://arxiv.org/abs/2311.03554. 55

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.