Pith. sign in

REVIEW 3 major objections 4 minor 69 references

In neural-mass simulation-based inference, a model can pass simulated-data recovery tests and still fail to cover the real recordings it is supposed to explain; the paper's hierarchical audit makes these two verdicts independent and identif

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-07-31 23:08 UTC pith:P45QFWBH

load-bearing objection A valuable audit framework for NMM-SBI, but the headline pass/fail contrast rests on an unstated conformal p-value threshold that needs formalizing. the 3 major comments →

arxiv 2607.24874 v1 pith:P45QFWBH submitted 2026-07-27 q-bio.QM

A Hierarchical Validity-Audit Framework for Neural Mass Models in Simulation-Based Inference: From Observational Coverage to Mechanistic Interpretation

classification q-bio.QM
keywords simulation-based inferenceneural mass modelsposterior validity auditobservational coveragesummary statisticsparameter identifiabilityEpileptorcanonical microcircuit
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that a neural mass model can look fully recoverable inside a simulator—tight posteriors, high calibration scores—and still fail when confronted with real recordings, because the model may not generate the observed dynamics, the summary features may have discarded the information needed for a specific parameter, or parameters may simply trade off against each other in the joint posterior. To keep these failure modes distinct, it introduces a three-stage audit: first test whether the model-prior-summary pipeline covers the real data in both the summary space and a separate fixed waveform-diagnostic space; then estimate target-specific posteriors for raw parameters, mechanistic ratios, and data-driven active directions, scoring them with a proper-score gain; finally examine joint-posterior coupling and cross-track consistency. Applied to real data, the audit finds the reduced single-source Epileptor does not cover the core seizure dynamics of seizure-onset-zone-local intracranial EEG (support p≈0.004), so none of its recovered coordinates can be read as patient-specific mechanisms, whereas the five-population canonical microcircuit conditionally covers mismatch-negativity data but exposes summary-induced information loss and a stable negative compensation between two gains. The payoff is a graded boundary on what 'successful SBI' may conclude.

Core claim

The central claim is that validity in neural-mass SBI must be established at three separate levels before interpretation: observational coverage, target-specific recoverability, and joint interpretability. The authors demonstrate that high posterior recovery under simulation (high proper-score gain, high R²) can coexist with failure to cover real observations in a fixed diagnostic space, and that a summary with excellent coverage (waveform PCA) can carry almost no parameter information while a poorer-coverage representation carries more. They isolate four diagnosable failure sources—model-configuration mismatch, summary-induced information loss, insufficient target information, and joint par

What carries the argument

The framework's carrying mechanism is a hierarchical audit with three stages and three target tracks. Stage one uses split-conformal local-support and local-predictive p-values in two parallel spaces—the SBI summary space and a fixed waveform-diagnostic space—so that compression cannot hide dynamical mismatch. Stage two defines targets as raw parameters, predefined mechanistic combinations (such as an excitation–inhibition gain ratio), and data-driven active directions, and scores each with a Proper Score Gain (the reduction in continuous ranked probability score relative to a condition-specific prior), adding a zero-waveform control and a waveform-complement branch to distinguish summary lo

Load-bearing premise

The load-bearing premise is that the waveform-diagnostic feature space was fixed before results were inspected and that the same p-value reading rules apply to both experiments: if the Epileptor's diagnostic set was chosen knowing it could not be satisfied, or if p≈0.176 is accepted as adequate while p≈0.004 is not, the headline Epileptor-fails/CMC-passes contrast collapses.

What would settle it

Re-run the Epileptor coverage audit with a diagnostic feature set that is registered before any real data are viewed—for example, features the restricted three-parameter model can plausibly generate—and pre-specify one p-value threshold for both experiments. If the Epileptor waveform support p-value rises to the CMC's ~0.18 range, or if the same threshold applied to the CMC classifies it as a failure, the paper's central demonstration is not robust.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • If a model fails the waveform-diagnostic coverage check, all downstream posterior estimates are restricted to within-simulator recoverability and cannot support patient-specific physiological claims.
  • Coverage and recoverability must be reported as separate quantities: a summary can cover the observed distribution (high conformal p-value) while discarding exactly the information needed to invert a target.
  • The summary-loss branches can identify when a learned summary should be augmented rather than abandoned, using a zero-waveform control to rule out pure capacity effects.
  • Joint-posterior coupling can reveal negative compensation between individually recoverable parameters, so single-target recovery is insufficient to certify independent interpretation.
  • The framework outputs graded rather than binary evidence, so users can see which conclusion levels are supported for each target.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same logic implies every SBI study that reports posterior means should also report a coverage diagnostic in a space not used for training; otherwise high calibration metrics can mask model misspecification.
  • The negative gain compensation detected in the CMC experiment is a testable signature: summaries that preserve it should be preferred, and simulations conditioned on real waveforms should reproduce the trade-off if it is genuine.
  • Applied to model comparison, the four-failure taxonomy could decide when added model complexity is justified: a candidate gains interpretability only if it passes coverage and target-invertibility audits, not merely if it fits simulated data better.
  • Requiring the diagnostic feature space to be fixed before real observations are examined, with a formal p-value threshold, would remove the main degree of freedom in the headline Epileptor failure.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces NMM-SBI Audit, a three-stage hierarchical validation framework for simulation-based inference with neural mass models. Stage 1 performs observation-layer calibration and then assesses coverage of real data in both the summary space used for SBI and a separate waveform diagnostic space, using split-conformal local-support and local-predictive p-values. Stage 2 trains target-specific posteriors for raw parameters, predefined mechanistic coordinates, and data-driven active directions, evaluating recoverability via proper-score gain, point-recovery metrics, and a zero-waveform-controlled summary-loss diagnostic. Stage 3 examines joint-posterior marginal stability, within-track parameter synchrony/coupling, and cross-track consistency (active–mechanism alignment, raw–mechanism distributional consistency, and local–global active consistency). The framework is applied to two real datasets: SOZ-local iEEG with a reduced single-source Epileptor model, and ERP CORE MMN with a fixed five-node CMC model. The authors conclude that the Epileptor configuration does not adequately cover core seizure dynamics (support p≈0.004), so its Step 2–3 results are only internal-recoverability evidence, whereas the CMC configuration conditionally covers the MMN data and supports limited interpretation of a few targets, while exposing summary information loss and instability in active-subspace alignment.

Significance. If the central claims hold, the framework makes a useful methodological contribution by separating four failure sources—configuration mismatch, summary-induced information loss, insufficient target information, and joint parameter compensation/cross-track coupling—and by explicitly warning against overinterpretation of within-simulator posterior recovery. The negative Epileptor result, the zero-waveform control, the structure-matched permutation reference for active–mechanism alignment, and the repeated caveat that Step 2–3 results are internal recoverability rather than real-data validation are all strengths. However, the headline contrast between the two applications rests on an unformalized interpretation of the conformal p-values in Stage 1, and the claim that the waveform diagnostic space was predefined is not substantiated by a preregistration artifact. These issues are fixable but currently prevent the strong binary conclusions from being fully supported.

major comments (3)
  1. [§3.3.1, Eq. (4)] The central pass/fail contrast—Epileptor fails while CMC conditionally passes—is read from two conformal p-values (0.004 vs 0.176) without a prespecified threshold, null model, or error-rate control. The text states that p≈0.5 is the most natural outcome but then treats 0.176 as 'comparatively natural behavior' and 0.004 as failure. A cutoff at 0.01 preserves the contrast, one at 0.2 makes both configurations fail, and a symmetric typicality rule around 0.5 makes CMC marginal at best. Because this classification determines which configurations are eligible for Step 2–3 interpretation, the headline conclusions are not yet established. Please add an explicit decision rule (e.g., a preregistered lower-tail cutoff with justification, or a calibrated reference distribution for the observed p-values), or present all downstream results purely as graded evidence without converting them into bina
  2. [§2.1, §2.2] It is unclear whether the real samples used in the dual-space coverage audit are the held-out test partition. Section 2.1 partitions real data 7:3 and calibrates the observation layer on the training split; Section 2.2 says 'real observed data are strictly held out' but does not explicitly state that the coverage p-values in Fig. 2 and §3.3.1 are computed exclusively on the test real samples. If the same real samples used to fit the observation layer appear in the coverage audit, the p-values are optimistically biased. Please specify the split used for every reported coverage p-value and confirm that calibration of the observation layer used only training real data, with all coverage statistics computed on the test real data.
  3. [§2.2, §3.2] The claim that the waveform diagnostic space D is 'predefined before the experiment' is used to rule out outcome-dependent construction of the diagnostic space, but no preregistration or a priori feature-selection protocol is provided. For the Epileptor, the 15-feature diagnostic set is the same set on which the model fails, so the failure is partly a statement about that feature choice. Even granting that D is fixed, the missing decision rule from the first major comment prevents the observed p-values from supporting the stated conclusions. I therefore ask for either a time-stamped preregistration or an explicit outcome-independence argument, together with a sensitivity analysis using alternative diagnostic feature sets. This is secondary to the missing threshold, but it affects how strongly the Epileptor negative result can be interpreted.
minor comments (4)
  1. [§3.3.2] In the point-recovery text, 'effective fast-system drivets' should be 'teff' or 'effective fast drive'. The label 'Dynamics feature' in Tables 3 and 4 is inconsistent with 'dynamical-feature' used elsewhere.
  2. [§2.2, §2.3] Several hyperparameters that affect the diagnostics are not specified: the kNN neighborhood size k, the number of reference centers Ncenter, the number of bootstrap resamples B, and the exact PCA dimensions for the waveform-complement branch are either omitted or only given in the experimental narrative. Please provide a complete parameter table for reproducibility.
  3. [§2.4.2(a), Tables 5–6] For the three-parameter Epileptor, the structure-matched permutation reference sets are small, so p_AM values are highly discrete (e.g., minimum 1/6). The tables report medians and ranges but not the reference-set sizes. Please report B_h and note the resulting granularity when interpreting 'no stable alignment'.
  4. [General] The manuscript does not state code or data availability. Given the reproducibility emphasis of the proposed audit framework, an availability statement for the analysis code, simulation pipelines, and processed data would strengthen the contribution.

Circularity Check

0 steps flagged

No significant circularity: the audit chain is self-contained; remaining concerns are threshold specification and preregistration transparency, not derivation-by-construction.

full rationale

The derivation chain is self-contained. Step 1 coverage is a split-conformal rank of held-out real observations against a calibration distribution from simulations; the p-values come from Eqs. (3)-(4) and could have gone either way (CMC p≈0.176 is read as adequate, Epileptor p≈0.004 as failure, but both are readings of the same statistic, not a quantity fitted to define the statistic). Step 2 PSG (Eqs. 7-9) compares posterior CRPS to a prior baseline on blind-test simulations, and the zero-waveform control isolates capacity effects; it is not a parameter fitted to the real data and then called a prediction. Step 3 within-track coupling, raw-mechanism W1 calibration, and active-mechanism permutation alignments are consistency checks with explicit null-like references; the absence of stable alignment is a falsifiable null result. No step defines its conclusion into its input: the diagnostic space D is a preselected test quantity, not a function of the coverage verdict, and the paper's own limitation statements restrict Steps 2-3 to within-simulation recoverability. The pass/fail threshold on conformal p-values is under-specified and the 'preregistered' status of D is not evidenced—these are transparency/correctness limitations, not circularity, because changing the threshold does not make any equation equivalent to its inputs. Any potentially author-overlapping citation (e.g., ref. [38]) is used only as an example summary architecture, not as load-bearing support for the audit claims.

Axiom & Free-Parameter Ledger

11 free parameters · 6 axioms · 1 invented entities

The framework's central claims do not depend on hidden constants, but the application results depend on chosen meta-parameters (L, τ, k, PCA dims, reference operating points) and on observation-layer nuisance parameters fitted to real training data. None of these is accompanied by sensitivity analysis beyond the 10-seed variability of the main pipeline.

free parameters (11)
  • Observation-layer gain α = not reported
    Bounded black-box search over Ψ_cand (§2.1, Eq. 2) fitted to reduce train-set discrepancy; applied in both experiments (App. 1.2, App. 2.2).
  • Background noise level ε(t) = not reported
    Automatically estimated by the observation-layer calibration (§2.1); part of the observation mapping noise term.
  • Temporal jitter δ (CMC only) = not reported
    Estimated automatically by the calibration layer (App. 2.2, Eq. 15).
  • Low-frequency background b(t) (CMC only) = not reported
    Estimated automatically by the calibration layer (App. 2.2, Eq. 15); parameterization unspecified.
  • Number of global active directions L = 2
    Hand-chosen dimensionality of the active subspace (§3.3.4: 'jointly treated as a two-dimensional global active subspace, with L = 2').
  • Practical coupling threshold τ = 0.3
    Chosen a priori for classifying joint-posterior coupling as moderate (§2.4.1).
  • kNN neighborhood size k = not specified
    k in the Local Support and Predictive Audits (Eqs. 3–5) is never given; a free meta-parameter of the coverage verdicts.
  • Number of reference centers Ncenter = not specified
    Active-track Jacobian estimation (§2.3) averages over Ncenter reference parameter centers; value not stated.
  • PCA dimensions for waveform-complement branch and waveform-PCA summary = 32 and 64
    Chosen representation dimensions (§3.2); affect which variance is retained in the summary-loss diagnostic.
  • Median nearest-neighbour bandwidth h = data-dependent
    Kernel bandwidth in Eqs. 5 and 25 set to the median NN distance; data-dependent and not reported per dataset.
  • Reference operating points in tF/S definition = Iref1=3.1, Iref2=0.45
    Chosen constants anchoring the fast/spike–wave drive-balance coordinate (§3.1.1, Eq. 28); the coordinate's scale depends on them.
axioms (6)
  • domain assumption Reference simulations under the simulator–prior are an exchangeable null for the real observations in each condition (conformal p-value calibration, Eq. 4).
    Split-conformal p-values in §2.2 assume calibration simulations and target observations are exchangeable under the null; for misspecified models the p-values are rank scores, which the paper acknowledges ('distinguished from the p value of an independent hypothesis test').
  • ad hoc to paper The waveform diagnostic space D is fixed before the experiment and independent of the audit outcome.
    §2.2: 'The function D is predefined before the experiment'; §3.2 calls the features 'preregistered,' but no registry or timestamp is provided and the Epileptor failure is defined relative to this space.
  • domain assumption SNPE posterior estimators are faithful enough for PSG, coupling, and Wasserstein comparisons.
    §3.2 selects SNPE as default; no SBC, C2ST, or MCMC reference check is run on the trained posteriors, yet §3.3.3 interprets joint-posterior correlations as parameter compensation.
  • domain assumption Local ridge-regression Jacobians approximate the simulator's sensitivity structure in neighborhoods.
    Active track in §2.3 and local–global test in §2.4.2 assume linearity of the summary w.r.t. standardized parameters within kNN neighborhoods; no diagnostic for the linearity assumption is provided.
  • domain assumption kNN geometry in 39–64 dimensional summary spaces is meaningful for local support.
    Local Support Audit (§2.2, Eq. 3) uses mean Euclidean distance to k neighbors; with dims 15–64 and 3,200/4,000 simulations, concentration-of-distances effects are not analyzed; k is not specified.
  • standard math Standard statistical machinery: fair CRPS estimator, bootstrap percentile intervals, 1-Wasserstein distance as divergence.
    Eqs. 7–9, 14–26 use textbook estimators; these are uncontroversial.
invented entities (1)
  • Mechanism coordinates teff, tF/S (Epileptor) and tE/I, tSP/DP, tS/I (CMC) no independent evidence
    purpose: Interpretable target axes for the mechanism track; designed 'specifically for this study' (§3.1.1–3.1.2).
    Defined a priori from model parameters with chosen reference operating points (Iref1=3.1, Iref2=0.45); they have no independent empirical handle, and the paper itself reports they do not stably align with data-driven active directions.

pith-pipeline@v1.3.0-alltime-deepseek · 35284 in / 25819 out tokens · 210939 ms · 2026-07-31T23:08:14.988011+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of A Hierarchical Validity-Audit Framework for Neural Mass Models in Simulation-Based Inference: From Observational Coverage to Mechanistic Interpretation." pith.science (2026). https://pith.science/paper/P45QFWBH

@misc{pith2026260724874,
  author       = {Pith},
  title        = {Pith review of: A Hierarchical Validity-Audit Framework for Neural Mass Models in Simulation-Based Inference: From Observational Coverage to Mechanistic Interpretation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P45QFWBH}},
  note         = {Machine review of arXiv:2607.24874}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Neural mass models describe population level neural activity using low-dimensional dynamical parameters and, through simulation-based inference, enable posterior estimation when explicit likelihoods are intractable, but good posterior recovery on simulated data does not guarantee that the model covers real observations, that summary statistics retain target information, or that multiple parameters stay independently interpretable. We introduce NMM-SBI Audit, a hierarchical validity-audit framework. It first evaluates whether a candidate model configuration covers the observed data. It then trains separate posterior estimators for parameter coordinates at multiple levels to assess summary-induced information loss. Multi-track joint posteriors are used to examine consistency across interpretations, with outcomes reported as graded evidence. We applied the framework to two real datasets: an SOZ-local iEEG-Epileptor model and an ERP CORE MMN-CMC model, using three summary representations. In the Epileptor experiment, waveform PCA simulations approximated observed signals, but persistent mismatch in preregistered dynamical diagnostics prevented interpreting recovered parameters as patient-specific mechanisms. The five population CMC configuration showed good conditional coverage, stable negative compensatory gain relationships, and coherent posteriors, while revealing summary information loss and active-subspace instability. NMM-SBI Audit thus distinguishes four failure sources: model-configuration mismatch, summary-induced information loss, insufficient target information, and joint parameter compensation/cross-track coupling. This prevents strong within-simulator recovery from being misread as valid physiological interpretation, providing a scalable constraint with explicit boundaries on supportable conclusions.

Figures

Figures reproduced from arXiv: 2607.24874 by Jiayuan He, Ning Jiang, Tianming Cai, Yuan Yang.

Figure 2
Figure 2. Figure 2: Dual-space coverage assessment for Epileptor/iEEG and CMC/MMN. (A) CMC/MMN assessment. The upper-left panel shows example real EEG waveforms exhibiting an evoked MMN response (black) together with simulated data generated by the CMC dynamical model (blue; 50 simulated traces shown for illustration). The lower-left panel shows the mean and standard deviation, across 10 random seeds, of the local-support p v… view at source ↗
Figure 3
Figure 3. Figure 3: PSG gains for targets across different tracks and summary representations in the Epileptor/iEEG experiment. Results for the original summary, zero-waveform branch, and waveform branch are reported across all random seeds. Between-group differences were assessed using two-sided paired t-tests matched by random seed. Within each target and summary representation, Holm correction was applied to the three pair… view at source ↗
Figure 4
Figure 4. Figure 4: PSG gains for targets across different tracks and summary representations in the CMC/MMN experiment. Results for the original summary, zero-waveform branch, and waveform branch are reported across all random seeds. Between-group differences were assessed using two-sided paired t-tests matched by random seed. Within each target and summary representation, Holm correction was applied to the three pairwise co… view at source ↗
Figure 5
Figure 5. Figure 5: Point-estimate recoverability of candidate targets across the three tracks in the Epileptor/iEEG and CMC/MMN experiments. (A) In the Epileptor/iEEG experiment, the evaluated targets included the raw targets I1, I2, and x0; the mechanism targets ts and tF/S; and the automatically identified active coordinates active 1 and active 2. (B) In the CMC/MMN experiment, the evaluated targets included five raw targe… view at source ↗
Figure 6
Figure 6. Figure 6: Within-track posterior synchrony and joint coupling in the Epileptor/iEEG experiment under the CNN-LSTM summary. Panels show the (A) raw, (B) mechanism, and (C) active tracks. The diagonal entries label the targets. The lower-triangular scatter plots show posterior means of target pairs across blind-test observations and random seeds, together with the Pearson correlation coefficient ρmean. Filled and open… view at source ↗
Figure 7
Figure 7. Figure 7: Within-track posterior synchrony and joint coupling in the Epileptor/iEEG experiment under the dynamical￾feature summary. Panels show the (A) raw, (B) mechanism, and (C) active tracks. The diagonal entries label the targets. The lower-triangular scatter plots show posterior means of target pairs across blind-test observations and random seeds, together with the Pearson correlation coefficient ρmean. Filled… view at source ↗
Figure 8
Figure 8. Figure 8: Within-track posterior synchrony and joint coupling in the Epileptor/iEEG experiment under the waveform￾PCA summary. Panels show the (A) raw, (B) mechanism, and (C) active tracks. The diagonal entries label the targets. The lower-triangular scatter plots show the posterior means of target pairs across blind-test observations and random seeds, together with the Pearson correlation coefficient ρmean. Filled … view at source ↗
Figure 9
Figure 9. Figure 9: Within-track posterior synchrony and joint coupling in the CMC/MMN experiment under the CNN-LSTM summary. Panels show the (A) raw, (B) mechanism, and (C) active tracks. The diagonal entries label the targets. The lower-triangular scatter plots show posterior means of target pairs across blind-test observations and random seeds, together with the Pearson correlation coefficient ρmean. Filled and open diamon… view at source ↗
Figure 10
Figure 10. Figure 10: Within-track posterior synchrony and joint coupling in the CMC/MMN experiment under the dynamical-feature summary. Panels show the (A) raw, (B) mechanism, and (C) active tracks. The diagonal entries label the targets. The lower-triangular scatter plots show the posterior means of target pairs across blind-test observations and random seeds, together with the Pearson correlation coefficient ρmean. Filled a… view at source ↗
Figure 11
Figure 11. Figure 11: Within-track posterior synchrony and joint coupling in the CMC/MMN experiment under the waveform-PCA summary. Panels show the (A) raw, (B) mechanism, and (C) active tracks. The diagonal entries label the targets. The lower-triangular scatter plots show posterior means of target pairs across blind-test observations and random seeds, together with the Pearson correlation coefficient ρmean. Filled and open d… view at source ↗
Figure 12
Figure 12. Figure 12: Cross-track consistency between the raw-posterior push-forward distribution and the directly estimated mechanism posterior. Blue indicates the distribution of the mechanism coordinate obtained by sampling from the raw joint posterior and applying the predefined mechanism function, whereas orange indicates the posterior estimated directly in the mechanism track. The upper curves show kernel density estimat… view at source ↗
Figure 1
Figure 1. Figure 1: Connection relationship of five nodes: left A1, left STG, right A1, right STG, and right IFG. 55 [PITH_FULL_IMAGE:figures/full_fig_p055_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

69 extracted references · 5 linked inside Pith

  1. [1]

    Multi-scale neural sources of EEG: genuine, equivalent, and representative

    Nunez P L, Nunez M D, Srinivasan R. Multi-scale neural sources of EEG: genuine, equivalent, and representative. A tutorial review[J]. Brain Topography, 2019, 32(2): 193-214

  2. [2]

    Towards large-scale, human-based, mesoscopic neurotechnologies[J]

    Chang E F. Towards large-scale, human-based, mesoscopic neurotechnologies[J]. Neuron, 2015, 86(1): 68-78

  3. [3]

    Multi-scale neural decoding and analysis[J]

    Lu H Y, Lorenc E S, Zhu H, et al. Multi-scale neural decoding and analysis[J]. Journal of neural engineering, 2021, 18(4): 045013

  4. [4]

    A neural mass model of spectral responses in electrophysiology[J]

    Moran R J, Kiebel S J, Stephan K E, et al. A neural mass model of spectral responses in electrophysiology[J]. NeuroImage, 2007, 37(3): 706-720

  5. [5]

    Bifurcation analysis of two coupled Jansen-Rit neural mass models[J]

    Ahmadizadeh S, Karoly P J, Nešić D, et al. Bifurcation analysis of two coupled Jansen-Rit neural mass models[J]. PloS one, 2018, 13(3): e0192842

  6. [6]

    A canonical microcircuit for neocortex[J]

    Douglas R J, Martin K A C, Whitteridge D. A canonical microcircuit for neocortex[J]. Neural computation, 1989, 1(4): 480-488

  7. [7]

    A dynamic causal model study of neuronal population dynamics[J]

    Marreiros A C, Kiebel S J, Friston K J. A dynamic causal model study of neuronal population dynamics[J]. Neuroimage, 2010, 51(1): 91-101. 48

  8. [8]

    Neural masses and fields in dynamic causal modeling[J]

    Moran R, Pinotsis D A, Friston K. Neural masses and fields in dynamic causal modeling[J]. Frontiers in computational neuroscience, 2013, 7: 57

  9. [9]

    Global dynamics of neural mass models[J]

    Cooray G K, Rosch R E, Friston K J. Global dynamics of neural mass models[J]. PLoS computational biology, 2023, 19(2): e1010915

  10. [10]

    Bayesian model comparison for simulation- based inference[J]

    Spurio Mancini A, Docherty M M, Price M A, et al. Bayesian model comparison for simulation- based inference[J]. RAS Techniques and Instruments, 2023, 2(1): 710-722

  11. [11]

    A comprehensive guide to simulation-based inference in computational biology[J]

    Wang X, Kelly R P, Jenner A L, et al. A comprehensive guide to simulation-based inference in computational biology[J]. arXiv preprint arXiv:2409.19675, 2024

  12. [12]

    Simulation-based inference: A practical guide[J]

    Deistler M, Boelts J, Steinbach P, et al. Simulation-based inference: A practical guide[J]. arXiv preprint arXiv:2508.12939, 2025

  13. [13]

    Benchmarking simulation-based infer- ence[C]//International conference on artificial intelligence and statistics

    Lueckmann J M, Boelts J, Greenberg D, et al. Benchmarking simulation-based infer- ence[C]//International conference on artificial intelligence and statistics. PMLR, 2021: 343-351

  14. [14]

    A crisis in simulation-based inference? Beware, your posterior approximations can be unfaithful[J]

    Hermans J, Delaunoy A, Rozet F, et al. A crisis in simulation-based inference? Beware, your posterior approximations can be unfaithful[J]. Transactions on Machine Learning Research, 2022

  15. [15]

    Variational methods for simulation-based inference[J]

    Glöckler M, Deistler M, Macke J H. Variational methods for simulation-based inference[J]. arXiv preprint arXiv:2203.04176, 2022

  16. [16]

    Simulation-based calibration checking for Bayesian computation: The choice of test quantities shapes sensitivity[J]

    Modrák M, Moon A H, Kim S, et al. Simulation-based calibration checking for Bayesian computation: The choice of test quantities shapes sensitivity[J]. Bayesian Analysis, 2023, 20(2): 461

  17. [17]

    Posterior SBC: simulation-based calibration checking conditional on data[J]

    Säilynoja T, Schmitt M, Bürkner P C, et al. Posterior SBC: simulation-based calibration checking conditional on data[J]. Statistics and Computing, 2026, 36(2): 78

  18. [18]

    E-valuating classifier two-sample tests[J]

    Pandeva T, Bakker T, Naesseth C A, et al. E-valuating classifier two-sample tests[J]. arXiv preprint arXiv:2210.13027, 2022

  19. [19]

    H2ST: Hierarchical two-sample tests for continual out-of-distribution detection[C]//Proceedings of the Computer Vision and Pattern Recognition Conference

    Liu Y, Zhao W, Guo Y. H2ST: Hierarchical two-sample tests for continual out-of-distribution detection[C]//Proceedings of the Computer Vision and Pattern Recognition Conference. 2025: 15413-15423

  20. [20]

    L-c2st: Local diagnostics for posterior approximations in simulation-based inference[J]

    Linhart J, Gramfort A, Rodrigues P. L-c2st: Local diagnostics for posterior approximations in simulation-based inference[J]. Advances in Neural Information Processing Systems, 2023, 36: 56384-56410

  21. [21]

    The frontier of simulation-based inference[J]

    Cranmer K, Brehmer J, Louppe G. The frontier of simulation-based inference[J]. Proceedings of the National Academy of Sciences, 2020, 117(48): 30055-30062

  22. [22]

    Parameter estimation and identifiability in a neural population model for electro-cortical activity[J]

    Hartoyo A, Cadusch P J, Liley D T J, et al. Parameter estimation and identifiability in a neural population model for electro-cortical activity[J]. PLoS computational biology, 2019, 15(5): e1006694

  23. [23]

    Universally sloppy parameter sensitivities in systems biology models[J]

    Gutenkunst R N, Waterfall J J, Casey F P, et al. Universally sloppy parameter sensitivities in systems biology models[J]. PLoS computational biology, 2007, 3(10): e189. 49

  24. [24]

    Perspective: Sloppiness and emergent theories in physics, biology, and beyond[J]

    Transtrum M K, Machta B B, Brown K S, et al. Perspective: Sloppiness and emergent theories in physics, biology, and beyond[J]. The Journal of chemical physics, 2015, 143(1)

  25. [25]

    Active subspaces: Emerging ideas for dimension reduction in parameter studies[M]

    Constantine P G. Active subspaces: Emerging ideas for dimension reduction in parameter studies[M]. Society for Industrial and Applied Mathematics, 2015

  26. [26]

    Mitigating parameter identifiability issues through model calibration on the data-informed active subspace: an example in tumor growth[J]

    Lewis A L, Everett R A. Mitigating parameter identifiability issues through model calibration on the data-informed active subspace: an example in tumor growth[J]. Available at SSRN 6055166

  27. [27]

    Dynamic causal modelling[J]

    Friston K J, Harrison L, Penny W. Dynamic causal modelling[J]. Neuroimage, 2003, 19(4): 1273-1302

  28. [28]

    Dynamic causal modelling revisited[J]

    Friston K J, Preller K H, Mathys C, et al. Dynamic causal modelling revisited[J]. Neuroimage, 2019, 199: 730-744

  29. [29]

    Universally sloppy parameter sensitivities in systems biology models[J]

    Gutenkunst R N, Waterfall J J, Casey F P, et al. Universally sloppy parameter sensitivities in systems biology models[J]. PLoS computational biology, 2007, 3(10): e189

  30. [30]

    Simulation-based Bayesian inference with ameliorative learned summary statistics–Part I[J]

    Befekadu G K. Simulation-based Bayesian inference with ameliorative learned summary statistics–Part I[J]. arXiv preprint arXiv:2601.22441, 2026

  31. [31]

    Using simulation-based inference for learning introductory statis- tics[J]

    Rossman A J, Chance B L. Using simulation-based inference for learning introductory statis- tics[J]. Wiley Interdisciplinary Reviews: Computational Statistics, 2014, 6(4): 211-221

  32. [32]

    All-in-one simulation-based inference[J]

    Gloeckler M, Deistler M, Weilbach C, et al. All-in-one simulation-based inference[J]. arXiv preprint arXiv:2404.09636, 2024

  33. [33]

    Detecting model misspecification in amortized Bayesian inference with neural networks[C]//Dagm german conference on pattern recognition

    Schmitt M, Bürkner P C, Köthe U, et al. Detecting model misspecification in amortized Bayesian inference with neural networks[C]//Dagm german conference on pattern recognition. Cham: Springer Nature Switzerland, 2023: 541-557

  34. [34]

    Simulation-based inference in agent-based models using spatio-temporal summary statistics[C]//International Conference on Computational Science

    Dignum E, Choudhary H, Lees M. Simulation-based inference in agent-based models using spatio-temporal summary statistics[C]//International Conference on Computational Science. Cham: Springer Nature Switzerland, 2025: 239-254

  35. [35]

    Learning robust statistics for simulation-based inference under model misspecification[J]

    Huang D, Bharti A, Souza A, et al. Learning robust statistics for simulation-based inference under model misspecification[J]. Advances in Neural Information Processing Systems, 2023, 36: 7289-7310

  36. [36]

    Active sequential posterior estimation for sample-efficient simulation-based inference[J]

    Griesemer S, Cao D, Cui Z, et al. Active sequential posterior estimation for sample-efficient simulation-based inference[J]. Advances in Neural Information Processing Systems, 2024, 37: 127907-127936.F

  37. [38]

    Eegpt: Pretrained transformer for universal and reliable rep- resentation of eeg signals[J]

    Wang G, Liu W, He Y, et al. Eegpt: Pretrained transformer for universal and reliable rep- resentation of eeg signals[J]. Advances in Neural Information Processing Systems, 2024, 37: 39249-39280

  38. [39]

    idecode: In-distribution equivariance for conformal out-of- distribution detection[C]//Proceedings of the AAAI conference on artificial intelligence

    Kaur R, Jha S, Roy A, et al. idecode: In-distribution equivariance for conformal out-of- distribution detection[C]//Proceedings of the AAAI conference on artificial intelligence. 2022, 36(7): 7104-7114. 50

  39. [40]

    Conformalk-NN Anomaly Detector for Univariate Data Streams[C]//Conformal and Probabilistic Prediction and Applications

    Ishimtsev V, Bernstein A, Burnaev E, et al. Conformalk-NN Anomaly Detector for Univariate Data Streams[C]//Conformal and Probabilistic Prediction and Applications. PMLR, 2017: 213-227

  40. [41]

    Testing for outliers with conformal p-values[J]

    Bates S, Candès E, Lei L, et al. Testing for outliers with conformal p-values[J]. The Annals of Statistics, 2023, 51(1): 149-178

  41. [42]

    On finding and using identifiable parameter combinations in nonlinear dynamic systems biology models and COMBOS: a novel web implementation[J]

    Meshkat N, Kuo C E, DiStefano III J. On finding and using identifiable parameter combinations in nonlinear dynamic systems biology models and COMBOS: a novel web implementation[J]. PloS one, 2014, 9(10): e110261

  42. [43]

    Active subspace methods in theory and practice: applications to kriging surfaces[J]

    Constantine P G, Dow E, Wang Q. Active subspace methods in theory and practice: applications to kriging surfaces[J]. SIAM Journal on Scientific Computing, 2014, 36(4): A1500-A1524

  43. [44]

    Strictly proper scoring rules, prediction, and estimation[J]

    Gneiting T, Raftery A E. Strictly proper scoring rules, prediction, and estimation[J]. Journal of the American statistical Association, 2007, 102(477): 359-378

  44. [45]

    Estimation of the continuous ranked probability score with limited information and applications to ensemble weather forecasts[J]

    Zamo M, Naveau P. Estimation of the continuous ranked probability score with limited information and applications to ensemble weather forecasts[J]. Mathematical Geosciences, 2018, 50(2): 209-234

  45. [46]

    Exceeding chance level by chance: The caveat of theoretical chance levels in brain signal classification and statistical assessment of decoding accuracy[J]

    Combrisson E, Jerbi K. Exceeding chance level by chance: The caveat of theoretical chance levels in brain signal classification and statistical assessment of decoding accuracy[J]. Journal of neuroscience methods, 2015, 250: 126-136

  46. [47]

    Feature selection with the Boruta package[J]

    Kursa M B, Rudnicki W R. Feature selection with the Boruta package[J]. Journal of statistical software, 2010, 36: 1-13

  47. [48]

    An introduction to the bootstrap New York[J]

    Efron B, Tibshirani R J. An introduction to the bootstrap New York[J]. NY: Chapman and Hall, 1993, 473

  48. [49]

    Statistical power analysis for the behavioral sciences[M]

    Cohen J. Statistical power analysis for the behavioral sciences[M]. routledge, 2013

  49. [50]

    Statistics on special manifolds[M]

    Chikuse Y. Statistics on special manifolds[M]. Springer Science & Business Media, 2003

  50. [51]

    On the nonconvexity of push-forward constraints and its consequences in machine learning[J]

    De Lara L, Deronzier M, González-Sanz A, et al. On the nonconvexity of push-forward constraints and its consequences in machine learning[J]. SIAM Journal on Mathematics of Data Science, 2025, 7(2): 597-620

  51. [52]

    Statistical aspects of Wasserstein distances[J]

    Panaretos V M, Zemel Y. Statistical aspects of Wasserstein distances[J]. Annual review of statistics and its application, 2019, 6(1): 405-431

  52. [53]

    The geometry of algorithms with orthogonality constraints[J]

    Edelman A, Arias T A, Smith S T. The geometry of algorithms with orthogonality constraints[J]. SIAM journal on Matrix Analysis and Applications, 1998, 20(2): 303-353

  53. [54]

    On the nature of seizure dynamics[J]

    Jirsa V K, Stacey W C, Quilichini P P, et al. On the nature of seizure dynamics[J]. Brain, 2014, 137(8): 2210-2230

  54. [55]

    Normative intracranial EEG maps epileptogenic tissues in focal epilepsy[J]

    Bernabei J M, Sinha N, Arnold T C, et al. Normative intracranial EEG maps epileptogenic tissues in focal epilepsy[J]. Brain, 2022, 145(6): 1949-1961

  55. [56]

    Topographic Variation in Human Neurotransmitter Receptor Densities Explains Differences in Intracranial EEG Spectra[J]

    Stoof U M, Friston K J, Tisdall M, et al. Topographic Variation in Human Neurotransmitter Receptor Densities Explains Differences in Intracranial EEG Spectra[J]. Human Brain Mapping, 2025, 46(16): e70393. 51

  56. [57]

    ERP CORE: An open resource for human event-related potential research[J]

    Kappenman E S, Farrens J L, Zhang W, et al. ERP CORE: An open resource for human event-related potential research[J]. NeuroImage, 2021, 225: 117465

  57. [58]

    Imaging human EEG dynamics using independent component analysis[J]

    Onton J, Westerfield M, Townsend J, et al. Imaging human EEG dynamics using independent component analysis[J]. Neuroscience & biobehavioral reviews, 2006, 30(6): 808-822

  58. [59]

    Enhancing EEG signals classification using LSTM-CNN architecture[J]

    Omar S M, Kimwele M, Olowolayemo A, et al. Enhancing EEG signals classification using LSTM-CNN architecture[J]. Engineering Reports, 2024, 6(9): e12827

  59. [60]

    Automatic posterior transformation for likelihood-free inference[C]//International conference on machine learning

    Greenberg D, Nonnenmacher M, Macke J. Automatic posterior transformation for likelihood-free inference[C]//International conference on machine learning. PMLR, 2019: 2404-2414

  60. [61]

    Methods and considerations for estimating parameters in biophysically detailed neural models with simulation based inference[J]

    Tolley N, Rodrigues P L C, Gramfort A, et al. Methods and considerations for estimating parameters in biophysically detailed neural models with simulation based inference[J]. PLOS Computational Biology, 2024, 20(2): e1011108

  61. [62]

    Is learning summary statistics necessary for likelihood-free inference?[C]//International Conference on Machine Learning

    Chen Y, Gutmann M U, Weller A. Is learning summary statistics necessary for likelihood-free inference?[C]//International Conference on Machine Learning. PMLR, 2023: 4529-4544

  62. [63]

    Individual brain structure and modelling predict seizure propagation[J]

    Proix T, Bartolomei F, Guye M, et al. Individual brain structure and modelling predict seizure propagation[J]. Brain, 2017, 140(3): 641-654

  63. [64]

    NeuroImage, 2020, 217: 116839

    HashemiM,VattikondaAN,SipV,etal.TheBayesianVirtualEpilepticPatient: Aprobabilistic framework designed to infer the spatial map of epileptogenicity in a personalized large-scale brain model of epilepsy spread[J]. NeuroImage, 2020, 217: 116839

  64. [65]

    Identifying spatio-temporal seizure propagation patterns in epilepsy using Bayesian inference[J]

    Vattikonda A N, Hashemi M, Sip V, et al. Identifying spatio-temporal seizure propagation patterns in epilepsy using Bayesian inference[J]. Communications biology, 2021, 4(1): 1244

  65. [66]

    A taxonomy of seizure dynamotypes[J]

    Saggio M L, Crisp D, Scott J M, et al. A taxonomy of seizure dynamotypes[J]. Elife, 2020, 9: e55632

  66. [67]

    Canonical microcircuits for predictive coding[J]

    Bastos A M, Usrey W M, Adams R A, et al. Canonical microcircuits for predictive coding[J]. Neuron, 2012, 76(4): 695-711

  67. [68]

    Dynamic causal modeling of the response to frequency deviants[J]

    Garrido M I, Kilner J M, Kiebel S J, et al. Dynamic causal modeling of the response to frequency deviants[J]. Journal of neurophysiology, 2009, 101(5): 2620-2631

  68. [69]

    Attentional enhancement of auditory mismatch responses: a DCM/MEG study[J]

    Auksztulewicz R, Friston K. Attentional enhancement of auditory mismatch responses: a DCM/MEG study[J]. Cerebral cortex, 2015, 25(11): 4273-4283. Appendix 1 Epileptor dynamics and observation mapping 1.1 Epileptor dynamical model The Epileptor is a phenomenological neural dynamical model designed to describe transitions between seizure and non-seizure sta...

  69. [70]

    Manual Feature Table Experiment Dim Feature group Features Epileptor/iEEG 1 Response difference response-baseline difference RMS 1 Peak timing peak time fraction 1 Difference area area under response-baseline difference 3 Temporal windows early difference mean; middle difference mean; late difference mean 2 Transition slopes onset slope; offset slope 4 Sp...