{"id":"325e2741-4eb6-4807-9d53-beb7a07c96b0","arxiv_id":"2607.24874","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":11,"one_line_summary":"A hierarchical audit framework separates model-coverage failure, summary-induced information loss, target non-identifiability, and joint parameter compensation in neural-mass simulation-based inference.","lead":"This paper builds a three-stage 'validity audit' for simulation-based inference with neural-mass brain models: check whether the model can generate the real recordings, check which of its parameters are actually recoverable, then check whether parameters stay consistent when inferred together. It shows that a model can pass a within-simulator recovery test yet fail on real seizure data, and that one reported 'good coverage' result depends on how coverage is defined.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Step-1 pass/fail verdicts use an unformalized threshold on conformal p-values: p≈0.176 is read as 'comparatively natural' and p≈0.004 as failure, but the method supplies no null model or cutoff, so the headline contrast is not yet established.","rationale":"The framework itself is coherent and the internal equations appear consistent; I find no load-bearing mathematical error. The most serious vulnerability is the interpretive step that converts two scalar conformal p-values into the central 'Epileptor fails / CMC conditionally passes' contrast. The reader's weakest assumption identified both the possible selection of the diagnostic space D and the missing pass/fail threshold. I focus on the threshold because it is the more formally addressable and directly determines the headline result. The proposed test would settle whether the threshold concern actually lands: if the CMC p-value is within the null range for an adequate model and the Epileptor p-value is not, the contrast is preserved under a prespecified rule. If not, the empirical demonstration is unsupported. Because this can be resolved by a calibration check, the appropriate verdict remains CONDITIONAL; my read does not move the reader's verdict, so I recommend UNCHANGED.","tokens_in":35916,"tokens_out":5500,"duration_ms":55718,"concrete_test":"Fix the diagnostic space D (freeze/version the 15 Epileptor and 39 ERP feature lists) and prespecify the decision rule before computing p-values. Then use each fitted simulator as a ground-truth adequate model: generate e.g. 100 pseudo-observed datasets from held-out parameter draws and fresh noise seeds, run the exact Step-1 observation-layer calibration, kNN support audit, and Eq. (4) p-value computation, and obtain the null distribution of the median diagnostic-space support p across known-adequate observations. Report the quantile of the real-data values p=0.004 and p=0.176 under this null. If p=0.176 sits above the prespecified lower-tail cutoff (e.g., 5th percentile) while p=0.004 sits below it, the claimed contrast is calibrated and stands; if p=0.176 also falls in the lower tail of the null, the CMC 'conditional pass' is not supported. Equivalently, compute the false-positive rat","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline contrast—Epileptor fails to cover SOZ-local iEEG, while CMC conditionally covers MMN—is decided in §3.3.1 by reading two diagnostic-space support p-values: 0.004 (Epileptor) and 0.176 (CMC). The method in §2.2 defines these as split-conformal upper-tail p-values (Eq. 4) and explicitly says they are 'not the p value of an independent hypothesis test.' It also says p≈0.5 is 'the most natural outcome.' Yet the classification of 0.176 as 'comparatively natural behavior' and 0.004 as a failure is made without any prespecified cutoff, null distribution, or error-rate control. The same §3.3.1 passage that calls CMC's p 'comparatively natural' would have to justify why 0.176 is functionally equivalent to 0.5 while 0.004 is not. This is load-bearing because both headline conclusions shift if the threshold changes: a cutoff at 0.01 keeps the contrast, a cutoff at 0.2 makes both configurations fail, and a symmetric 'typicality' rule around 0.5 would make CMC marginal at best. The paper's later 'graded, not binary' language does not fix this, because the graded values are converted into binary eligibility for Step 2–3 interpretation ('adequate coverage' vs 'only internal recoverability') at exactly this unstated boundary. The related worry that D was preselected knowing its outcome is secondary; even granting that D is fixed, the missing decision rule prevents the observed p-values from supporting the stated conclusions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces NMM-SBI Audit, a three-stage hierarchical validation framework for simulation-based inference with neural mass models. Stage 1 performs observation-layer calibration and then assesses coverage of real data in both the summary space used for SBI and a separate waveform diagnostic space, using split-conformal local-support and local-predictive p-values. Stage 2 trains target-specific posteriors for raw parameters, predefined mechanistic coordinates, and data-driven active directions, evaluating recoverability via proper-score gain, point-recovery metrics, and a zero-waveform-controlled summary-loss diagnostic. Stage 3 examines joint-posterior marginal stability, within-track parameter synchrony/coupling, and cross-track consistency (active–mechanism alignment, raw–mechanism distributional consistency, and local–global active consistency). The framework is applied to two real datasets: SOZ-local iEEG with a reduced single-source Epileptor model, and ERP CORE MMN with a fixed five-node CMC model. The authors conclude that the Epileptor configuration does not adequately cover core seizure dynamics (support p≈0.004), so its Step 2–3 results are only internal-recoverability evidence, whereas the CMC configuration conditionally covers the MMN data and supports limited interpretation of a few targets, while exposing summary information loss and instability in active-subspace alignment.","tokens_in":36309,"tokens_out":6330,"duration_ms":64016,"significance":"If the central claims hold, the framework makes a useful methodological contribution by separating four failure sources—configuration mismatch, summary-induced information loss, insufficient target information, and joint parameter compensation/cross-track coupling—and by explicitly warning against overinterpretation of within-simulator posterior recovery. The negative Epileptor result, the zero-waveform control, the structure-matched permutation reference for active–mechanism alignment, and the repeated caveat that Step 2–3 results are internal recoverability rather than real-data validation are all strengths. However, the headline contrast between the two applications rests on an unformalized interpretation of the conformal p-values in Stage 1, and the claim that the waveform diagnostic space was predefined is not substantiated by a preregistration artifact. These issues are fixable but currently prevent the strong binary conclusions from being fully supported.","major_comments":[{"comment":"The central pass/fail contrast—Epileptor fails while CMC conditionally passes—is read from two conformal p-values (0.004 vs 0.176) without a prespecified threshold, null model, or error-rate control. The text states that p≈0.5 is the most natural outcome but then treats 0.176 as 'comparatively natural behavior' and 0.004 as failure. A cutoff at 0.01 preserves the contrast, one at 0.2 makes both configurations fail, and a symmetric typicality rule around 0.5 makes CMC marginal at best. Because this classification determines which configurations are eligible for Step 2–3 interpretation, the headline conclusions are not yet established. Please add an explicit decision rule (e.g., a preregistered lower-tail cutoff with justification, or a calibrated reference distribution for the observed p-values), or present all downstream results purely as graded evidence without converting them into bina","section":"§3.3.1, Eq. (4)"},{"comment":"It is unclear whether the real samples used in the dual-space coverage audit are the held-out test partition. Section 2.1 partitions real data 7:3 and calibrates the observation layer on the training split; Section 2.2 says 'real observed data are strictly held out' but does not explicitly state that the coverage p-values in Fig. 2 and §3.3.1 are computed exclusively on the test real samples. If the same real samples used to fit the observation layer appear in the coverage audit, the p-values are optimistically biased. Please specify the split used for every reported coverage p-value and confirm that calibration of the observation layer used only training real data, with all coverage statistics computed on the test real data.","section":"§2.1, §2.2"},{"comment":"The claim that the waveform diagnostic space D is 'predefined before the experiment' is used to rule out outcome-dependent construction of the diagnostic space, but no preregistration or a priori feature-selection protocol is provided. For the Epileptor, the 15-feature diagnostic set is the same set on which the model fails, so the failure is partly a statement about that feature choice. Even granting that D is fixed, the missing decision rule from the first major comment prevents the observed p-values from supporting the stated conclusions. I therefore ask for either a time-stamped preregistration or an explicit outcome-independence argument, together with a sensitivity analysis using alternative diagnostic feature sets. This is secondary to the missing threshold, but it affects how strongly the Epileptor negative result can be interpreted.","section":"§2.2, §3.2"}],"minor_comments":[{"comment":"In the point-recovery text, 'effective fast-system drivets' should be 'teff' or 'effective fast drive'. The label 'Dynamics feature' in Tables 3 and 4 is inconsistent with 'dynamical-feature' used elsewhere.","section":"§3.3.2"},{"comment":"Several hyperparameters that affect the diagnostics are not specified: the kNN neighborhood size k, the number of reference centers Ncenter, the number of bootstrap resamples B, and the exact PCA dimensions for the waveform-complement branch are either omitted or only given in the experimental narrative. Please provide a complete parameter table for reproducibility.","section":"§2.2, §2.3"},{"comment":"For the three-parameter Epileptor, the structure-matched permutation reference sets are small, so p_AM values are highly discrete (e.g., minimum 1/6). The tables report medians and ranges but not the reference-set sizes. Please report B_h and note the resulting granularity when interpreting 'no stable alignment'.","section":"§2.4.2(a), Tables 5–6"},{"comment":"The manuscript does not state code or data availability. Given the reproducibility emphasis of the proposed audit framework, an availability statement for the analysis code, simulation pipelines, and processed data would strengthen the contribution.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern is valid and load-bearing. The missing threshold in Stage 1 is the main obstacle: the manuscript's own language converts continuous p-values into binary eligibility for interpretation, and the headline contrast shifts under reasonable alternative thresholds. The preregistration/D-selection issue is secondary but should also be addressed. The framework itself is sound in its internal estimators, and the authors are appropriately cautious in several places, so I see this as fixable with additional analysis and rewriting rather than a reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is worth your time. It builds a three-stage audit for simulation-based inference with neural mass models, and the core idea is right: before you report parameters from an SBI fit, check whether the model actually covers the real data, whether the summary carries information about the target you care about, and whether joint posteriors create compensatory structures that make single-target interpretation misleading. The framework implements that with proper-score gains (PSG), zero-waveform controls, a three-track decomposition into raw/mechanism/active targets, and active-subspace alignment tests. The internal equations are consistent, and the two case studies are genuinely informative. The Epileptor negative result is the strongest part: even after observation-layer calibration, the diagnostic-space support p-value stays near 0.004, so the authors correctly refuse to interpret parameters on real data. That's exactly the restraint the field needs.\n\nThe soft spot is real and load-bearing. The paper defines conformal p-values (Eq. 4) and explicitly says they are not hypothesis-test p-values, then uses them to dichotomize adequacy anyway. p≈0.176 in the CMC diagnostic space is read as 'comparatively natural behavior' while p≈0.004 is read as failure, but no cutoff, null model, or error-rate control is supplied. A symmetric typicality rule around 0.5 would make CMC marginal; a cutoff at 0.2 flips both conclusions. That's not a nitpick — both headline conclusions depend on it. The later 'graded, not binary' language doesn't fix the boundary, because the boundary is silently applied to decide which configurations proceed to Step 2–3 interpretation.\n\nSecondary concerns: the 'preregistered' waveform diagnostic features have no artifact, and for Epileptor they are the very features that produce the failure, so a skeptic could argue the mismatch is partly constructed. The paper also never benchmarks against SBC or L-C2ST, so the marginal value of the full hierarchy is asserted rather than measured. A posterior-faithfulness check (SBC/C2ST) for the amortized posteriors would strengthen the compensation claims.\n\nWho it's for: anyone doing SBI with neural mass models, especially EEG/iEEG inversion. It deserves a serious referee — the framework is novel enough and the failure modes it targets are real. But it needs a formalized threshold, a justification for D (or an actual preregistration), and a comparison against existing diagnostics before I'd publish it.","headline":"A valuable audit framework for NMM-SBI, but the headline pass/fail contrast rests on an unstated conformal p-value threshold that needs formalizing.","tokens_in":36914,"tokens_out":2779,"would_cite":true,"duration_ms":26573,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In neural-mass simulation-based inference, a model can pass simulated-data recovery tests and still fail to cover the real recordings it is supposed to explain; the paper's hierarchical audit makes these two verdicts independent and identif","keywords":["simulation-based inference","neural mass models","posterior validity audit","observational coverage","summary statistics","parameter identifiability","Epileptor","canonical microcircuit"],"falsifier":"Re-run the Epileptor coverage audit with a diagnostic feature set that is registered before any real data are viewed—for example, features the restricted three-parameter model can plausibly generate—and pre-specify one p-value threshold for both experiments. If the Epileptor waveform support p-value rises to the CMC's ~0.18 range, or if the same threshold applied to the CMC classifies it as a failure, the paper's central demonstration is not robust.","tokens_in":35673,"feed_emoji":"🧠","tokens_out":8170,"duration_ms":68310,"temperature":0.7,"pith_summary":"The paper argues that a neural mass model can look fully recoverable inside a simulator—tight posteriors, high calibration scores—and still fail when confronted with real recordings, because the model may not generate the observed dynamics, the summary features may have discarded the information needed for a specific parameter, or parameters may simply trade off against each other in the joint posterior. To keep these failure modes distinct, it introduces a three-stage audit: first test whether the model-prior-summary pipeline covers the real data in both the summary space and a separate fixed waveform-diagnostic space; then estimate target-specific posteriors for raw parameters, mechanistic ratios, and data-driven active directions, scoring them with a proper-score gain; finally examine joint-posterior coupling and cross-track consistency. Applied to real data, the audit finds the reduced single-source Epileptor does not cover the core seizure dynamics of seizure-onset-zone-local intracranial EEG (support p≈0.004), so none of its recovered coordinates can be read as patient-specific mechanisms, whereas the five-population canonical microcircuit conditionally covers mismatch-negativity data but exposes summary-induced information loss and a stable negative compensation between two gains. The payoff is a graded boundary on what 'successful SBI' may conclude.","feed_headline":"Simulated-data success does not prove model fits real data","feed_subtitle":"Three-stage audit: Epileptor fails seizure-dynamics coverage; CMC conditionally passes but hides information loss.","key_machinery":"The framework's carrying mechanism is a hierarchical audit with three stages and three target tracks. Stage one uses split-conformal local-support and local-predictive p-values in two parallel spaces—the SBI summary space and a fixed waveform-diagnostic space—so that compression cannot hide dynamical mismatch. Stage two defines targets as raw parameters, predefined mechanistic combinations (such as an excitation–inhibition gain ratio), and data-driven active directions, and scores each with a Proper Score Gain (the reduction in continuous ranked probability score relative to a condition-specific prior), adding a zero-waveform control and a waveform-complement branch to distinguish summary lo","core_discovery":"The central claim is that validity in neural-mass SBI must be established at three separate levels before interpretation: observational coverage, target-specific recoverability, and joint interpretability. The authors demonstrate that high posterior recovery under simulation (high proper-score gain, high R²) can coexist with failure to cover real observations in a fixed diagnostic space, and that a summary with excellent coverage (waveform PCA) can carry almost no parameter information while a poorer-coverage representation carries more. They isolate four diagnosable failure sources—model-configuration mismatch, summary-induced information loss, insufficient target information, and joint par","pith_inferences":["The same logic implies every SBI study that reports posterior means should also report a coverage diagnostic in a space not used for training; otherwise high calibration metrics can mask model misspecification.","The negative gain compensation detected in the CMC experiment is a testable signature: summaries that preserve it should be preferred, and simulations conditioned on real waveforms should reproduce the trade-off if it is genuine.","Applied to model comparison, the four-failure taxonomy could decide when added model complexity is justified: a candidate gains interpretability only if it passes coverage and target-invertibility audits, not merely if it fits simulated data better.","Requiring the diagnostic feature space to be fixed before real observations are examined, with a formal p-value threshold, would remove the main degree of freedom in the headline Epileptor failure."],"forward_implications":["If a model fails the waveform-diagnostic coverage check, all downstream posterior estimates are restricted to within-simulator recoverability and cannot support patient-specific physiological claims.","Coverage and recoverability must be reported as separate quantities: a summary can cover the observed distribution (high conformal p-value) while discarding exactly the information needed to invert a target.","The summary-loss branches can identify when a learned summary should be augmented rather than abandoned, using a zero-waveform control to rule out pure capacity effects.","Joint-posterior coupling can reveal negative compensation between individually recoverable parameters, so single-target recovery is insufficient to certify independent interpretation.","The framework outputs graded rather than binary evidence, so users can see which conclusion levels are supported for each target."],"fun_headline_variants":["Simulated success ≠ real-data fit in neural mass SBI","Three-level audit catches fake validity in neural mass models","Epileptor fails real-data coverage despite simulated recovery","SBI audit distinguishes four failure sources for neural mass models","Simulated posterior accuracy doesn't guarantee real-data coverage"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the waveform-diagnostic feature space was fixed before results were inspected and that the same p-value reading rules apply to both experiments: if the Epileptor's diagnostic set was chosen knowing it could not be satisfied, or if p≈0.176 is accepted as adequate while p≈0.004 is not, the headline Epileptor-fails/CMC-passes contrast collapses.","fun_headline_variants_meta":{"raw":{"variants":["Simulated success ≠ real-data fit in neural mass SBI","Three-level audit catches fake validity in neural mass models","Epileptor fails real-data coverage despite simulated recovery","SBI audit distinguishes four failure sources for neural mass models","Simulated posterior accuracy doesn't guarantee real-data coverage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000696,"raw_usage":{"total_tokens":3005,"prompt_tokens":789,"completion_tokens":2216,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":2146}},"tokens_in":533,"tokens_out":2216,"duration_ms":14440,"temperature":1.0,"reasoning_tokens":2146,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T23:08:14.988011+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the Epileptor coverage audit with a diagnostic feature set that is registered before any real data are viewed—for example, features the restricted three-parameter model can plausibly generate—and pre-specify one p-value threshold for both experiments. If the Epileptor waveform support p-value rises to the CMC's ~0.18 range, or if the same threshold applied to the CMC classifies it as a failure, the paper's central demonstration is not robust.","supporting_citations":[],"review_version":1}