{"id":"340bb770-64f8-4453-9dc4-350830be971c","arxiv_id":"2607.05282","paper_version":1,"verdict":"CONDITIONAL","confidence":"UNKNOWN","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":5,"one_line_summary":"A single sentence stating the discovery directly. ≤ 300 chars.","lead":"Two short sentences. ≤ 480 chars total. The first answers ","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"SHAP analysis contradicts the central biomarker claim: classifier decisions are driven by first-order (S1) features, while the paper's biomarker narrative rests on second-order (S2) dominance from univariate statistics.","rationale":"The reader's CONDITIONAL verdict is appropriate. The paper makes a genuine methodological contribution by applying WST with strict LOSO-CV to a well-known dataset, and the classification results (90.48% accuracy) are plausible given prior work on the same dataset. The preprocessing, LOSO protocol, and subject-level majority voting are methodologically sound. However, the central biomarker claim — that S2/cross-frequency coupling is the primary signature — rests on a tension that the paper does not adequately address. The SHAP analysis, which the paper claims provides independent validation, actually shows the opposite pattern (S1 dominance in classifier decisions). The non-significant Spearman correlation (p=0.0969) further undermines the claimed 'cross-methodological consistency.' This is not a fatal flaw in the classification pipeline itself, but it does weaken the biomarker discovery narrative that constitutes the paper's primary scientific contribution beyond pure classification. Additionally, the SVM results (79.76%) used FDR-significant features selected on the full dataset including test subjects — a feature selection leakage issue — though this is secondary since the RF is the headline classifier. The WST parameter choices (J=7, Q=(8,1)) are described as 'optimized' without specifying whether this optimization used LOSO performance, which could introduce additional optimistic bias if hyperparameters were selected by peeking at test folds. The proposed concrete test (S1-only vs S2-only RF comparison) would directly settle whether the S2 dominance claim holds under the paper's own classification framework.","tokens_in":17625,"tokens_out":3642,"duration_ms":70834,"concrete_test":"Train two separate Random Forest classifiers under identical LOSO-CV settings: one using only S1 features (46 paths × 16 channels × 2 time bins) and one using only S2 features (129 paths × 16 channels × 2 time bins). Compare subject-level LOSO accuracy, AUC, and sensitivity. If S1-only performance is comparable to or exceeds S2-only performance, the claim that second-order cross-frequency coupling is the 'primary electrophysiological signature' of schizophrenia is not supported by the classification evidence and should be revised.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central biomarker claim is that 'temporal amplitude modulation constitutes the primary electrophysiological signature of schizophrenia,' supported by the finding that 78.5% of FDR-significant features are second-order (S2) scattering coefficients. The paper explicitly states that SHAP was incorporated 'to independently validate the identified biomarkers through cross-methodological consistency.' However, the SHAP results contradict this: the top 5 SHAP-ranked features are ALL first-order (S1) coefficients (T5-S1[17], T6-S1[17], P3-S1[6], C4-S1[16], T6-S1[16]). The Spearman correlation between statistical F-score rankings and SHAP channel importance is ρs=0.429, p=0.0969 — non-significant. The paper frames this as 'moderate positive trend' and 'localized alignment,' but it is a null result. Furthermore, the statistical analysis identifies P3 as the single most discriminative electrode, while the top SHAP features come from T5, T6, and C4. The paper claims joint evidence from two independent methods, but the two methods disagree on the most fundamental structural property (scattering order) and on the top electrode sites. This means the biomarker discovery claim — that cross-frequency coupling (S2) is the primary signature — is supported only by univariate ANOVA, not by the classifier's actual decision process. If the classifier primarily relies on S1 (spectral energy) features, the claim that amplitude modulation dynamics are the 'primary' signature is overstated.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This manuscript proposes a Wavelet Scattering Transform (WST) framework for schizophrenia biomarker discovery and classification from resting-state EEG. The pipeline extracts multi-order scattering coefficients (S0, S1, S2), applies subject-level ANOVA with FDR correction for biomarker identification, and trains Random Forest and SVM classifiers under strict Leave-One-Subject-Out (LOSO) cross-validation with subject-level majority voting. SHAP explainability analysis is used to cross-validate statistical biomarker findings. The RF classifier achieves 90.48% accuracy (AUC = 0.9339). The central biomarker claim is that second-order (S2) scattering coefficients, encoding cross-frequency coupling, dominate the discriminative feature set (78.5% of FDR-significant features), and that temporal amplitude modulation constitutes the primary electrophysiological signature of schizophrenia.","tokens_in":18660,"tokens_out":1498,"duration_ms":144924,"significance":"The manuscript addresses a genuine methodological gap in EEG-based schizophrenia classification by combining the WST (which captures amplitude modulation structure inaccessible to standard PSD features) with strict LOSO cross-validation and SHAP explainability. The use of subject-level statistics to avoid pseudo-replication from temporally overlapping epochs is a sound methodological choice. The LOSO evaluation protocol is appropriately rigorous for the 84-subject dataset. The reproducible use of kymatio for WST implementation and the publicly available Kaggle dataset are strengths. However, the central biomarker claim faces an internal consistency problem between the statistical analysis and the SHAP validation, as detailed below.","major_comments":[{"comment":"§1 (Introduction) and §3.5: The paper states that SHAP was incorporated 'to independently validate the identified biomarkers through cross-methodological consistency.' However, the SHAP results (§3.5) contradict the central biomarker claim rather than validating it. The top 5 SHAP-ranked features are all first-order (S1) coefficients (T5-S1[17], T6-S1[17], P3-S1[6], C4-S1[16], T6-S1[16]), while the paper's biomarker narrative rests on S2 dominance (78.5% of FDR-significant features). The Spearman correlation between F-score rankings and SHAP channel importance is ρs=0.429, p=0.0969 — non-significant. The paper frames this as 'moderate positive trend' and 'localized alignment,' but with p>0.05 this is a null result. Furthermore, the statistical analysis identifies P3 as the most discriminative electrode, while top SHAP features come from T5, T6, and C4. The claim of 'cross-methodological'","section":null},{"comment":"§3.5 and §3.1: The spatial concordance claim is also inconsistent. The ANOVA identifies P3 as the single most discriminative electrode (12 of 27 Bonferroni-significant features), while SHAP channel-level importance (Fig. 11) shows P3, O1, and F4 as top regions. The paper states the model's decision architecture is 'primarily prioritized around the left-parietal (P3), left-occipital (O1), and right-frontal (F4) regions,' but the top individual SHAP features are from T5, T6, and C4. The authors should reconcile these discrepancies or temper the 'independent validation' framing. As written, the two methods disagree on both scattering order (S1 vs S2) and top electrode sites, which undermines the 'joint evidence' claim.","section":null},{"comment":"Table 4: The proposed method's validation is listed as 'Subject-Level Holdout,' but the text (§2.6) describes Leave-One-Subject-Out cross-validation with 84 folds. These are different protocols. The table should say 'LOSO CV' to match the methodology. Additionally, the comparison with Sravanthi et al. (2026) [77], which uses 'Subject-wise LOOCV' on the same dataset (MHRC, 45 SZ / 39 HC) and reports 96.7% accuracy, is not discussed in the text despite being the most directly comparable result. The authors should comment on why their method underperforms this benchmark on the same data.","section":null}],"minor_comments":[{"comment":"§2.2: The specific z-score threshold value is not stated ('an adaptive z-score threshold was used'). The threshold is listed as a free parameter in the axiom ledger but its value should be reported for reproducibility.","section":null},{"comment":"§2.6: The RF hyperparameters (200 trees, max depth 20) are stated but it is unclear whether these were selected via nested cross-validation or fixed a priori. If tuned on the same LOSO folds, this introduces optimistic bias. Please clarify the tuning protocol.","section":null},{"comment":"§2.3: The claim that J=7 is 'the optimal choice' is supported by qualitative arguments about deformation stability and temporal dynamics, but no quantitative comparison with alternative J values is provided. Consider softening to 'a principled choice' rather than 'optimal,' or provide empirical justification.","section":null},{"comment":"§3.1: The Bonferroni-significant subset (27 features) is described as comprising 'both first- and second-order scattering coefficients,' but the S1/S2 breakdown is not reported. Given that the FDR-significant set is 78.5% S2, the breakdown for the Bonferroni subset would be informative.","section":null},{"comment":"§3.2: The statement 'no delta or theta features survived correction' is notable given the schizophrenia EEG literature's emphasis on slow-wave abnormalities. The authors attribute this to 'high cohort variability for slower alterations' but do not provide evidence. Consider softening this interpretation.","section":null},{"comment":"Table 2: The 'Mean % Change' values appear to be computed from WST scattering coefficients, not raw spectral power. The table caption should clarify that these are percentage changes in scattering energy, not traditional band power.","section":null},{"comment":"§4 (Discussion): Several references in the discussion (e.g., [61, 62, 63] cited together for posterior predominance) make it difficult to trace which specific finding from which citation supports which claim. Consider separating multi-citation clusters where they support distinct sub-claims.","section":null},{"comment":"Fig. 10 caption: 'Fig' 10' has a stray apostrophe/quote mark. Should read 'Figure 10' or 'Fig. 10'.","section":null},{"comment":"Abstract: 'Hierarchical WST coefficients capturing multi-scale amplitude modulation structure was extracted' — subject-verb agreement; should be 'were extracted.'","section":null}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the S1/S2 discrepancy between SHAP and ANOVA is well-founded and is the primary reason for the major revision recommendation. The authors may be able to address this by reframing the biomarker claim (e.g., 'S2 features dominate the statistical biomarker set' rather than 'amplitude modulation is the primary signature'), but the current framing overstates what the evidence supports. The SHAP analysis was explicitly positioned as independent validation, so the contradiction is not a minor presentation issue."},"author_rebuttal":{"model":"glm-5.2","summary":"WST-based EEG framework achieves 90.48% LOSO accuracy for schizophrenia classification; S2 coefficients dominate statistical biomarkers but SHAP prioritizes S1 features, creating an internal consistency problem requiring revision.","responses":[{"response":"The referee is correct that the current framing overstates the degree of cross-methodological consistency. We acknowledge the following: (1) The top-5 SHAP-ranked features are indeed all first-order (S1) coefficients, which contrasts with the S2 dominance (78.5%) observed in the FDR-significant biomarker set. (2) The Spearman correlation between F-score rankings and SHAP channel importance (ρs = 0.429, p = 0.0969) does not reach conventional statistical significance, and describing this as evidence of 'independent validation' is not justified. We will revise the manuscript to accurately characterize this as a null result for the rank correlation and to explicitly acknowledge the discrepancy in scattering order between the statistical and SHAP analyses. We believe this discrepancy is itself scientifically informative: univariate ANOVA identifies S2 features as the most statistically discriminative between groups, while the multivariate Random Forest model relies more heavily on S1 features for its classification decisions. This suggests that group-level statistical separation and model-level discriminative utility are related but distinct properties, and we will discuss this distinction explicitly rather than claiming validation. The 'independent validation' framing in the Introduction and §3.5 will be removed and replaced with a more measured characterization: SHAP provides complementary, model-level interpretability that partially overlaps with but does not replicate the statistical biomarker findings. We agree this is a substantive revision to the paper's narrative.","revision_made":"yes","referee_comment":"The SHAP results contradict the central biomarker claim: top 5 SHAP features are all S1, while the paper's narrative rests on S2 dominance. The Spearman correlation (ρs=0.429, p=0.0969) is non-significant, yet framed as 'moderate positive trend.' The claim of 'cross-methodological consistency' is not supported."},{"response":"The referee correctly identifies a genuine inconsistency in our spatial concordance narrative. The discrepancy operates at two levels: (a) at the channel-aggregated level, P3 does appear prominently in both analyses (P3 is the most discriminative electrode in ANOVA and among the top three in SHAP channel-level importance), which represents partial spatial overlap; (b) at the individual feature level, the top SHAP features (T5-S1[17], T6-S1[17], C4-S1[16]) do not coincide with the top ANOVA features (dominated by P3). We will revise the manuscript to clearly separate these two levels of analysis and to state honestly that the spatial concordance is partial—limited to the channel-aggregated level for P3—rather than claiming broad 'joint evidence.' The discrepancy between channel-level SHAP importance (where P3, O1, F4 rank highly) and individual-feature-level SHAP importance (where T5, T6, C4 features top the list) likely reflects the fact that channel-level aggregation sums across many scattering paths, so a channel can rank highly without any single feature from it appearing in the top-5. We will add this explanation and remove the implication that the two methods converge on the same electrode sites at the feature level. The 'joint evidence' language will be removed.","revision_made":"yes","referee_comment":"Spatial concordance claim is inconsistent: ANOVA identifies P3 as most discriminative electrode, while SHAP channel-level importance shows P3, O1, and F4 as top regions, and top individual SHAP features come from T5, T6, and C4. The two methods disagree on both scattering order and top electrode sites, undermining the 'joint evidence' claim."},{"response":"The referee is correct on both points. First, the Table 4 label 'Subject-Level Holdout' is inaccurate; the methodology (§2.6) clearly describes 84-fold Leave-One-Subject-Out cross-validation. This is an error in the table and will be corrected to 'LOSO CV.' Second, we agree that the Sravanthi et al. (2026) result (96.7% accuracy on the same MHRC dataset with subject-wise LOOCV) is the most directly comparable benchmark and should be discussed explicitly. We will add a discussion paragraph addressing this comparison. We note the following relevant differences: Sravanthi et al. employ Variational Mode Decomposition with multi-domain features and evaluate 9 ML plus 7 optimized ML classifiers, selecting the best-performing configuration, whereas our approach uses a single Random Forest model on WST features. Their higher accuracy may reflect the broader feature space and classifier optimization strategy. However, we also note that our framework's primary contribution is not maximizing classification accuracy but rather the interpretable biomarker discovery pipeline (WST + ANOVA + SHAP), which provides neurophysiological insight that a VMD-based approach does not directly offer. We will state this comparison honestly, including the accuracy gap, rather than omitting it.","revision_made":"yes","referee_comment":"Table 4 lists validation as 'Subject-Level Holdout' but the text describes LOSO CV with 84 folds. The table should say 'LOSO CV.' Additionally, the comparison with Sravanthi et al. (2026), which uses Subject-wise LOOCV on the same dataset and reports 96.7% accuracy, is not discussed despite being the most directly comparable result."}],"tokens_in":18142,"tokens_out":1192,"duration_ms":157001,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The headline: the paper's central biomarker claim — that second-order scattering coefficients (cross-frequency coupling) are the primary electrophysiological signature of schizophrenia — is supported only by univariate ANOVA, not by the classifier's actual decision process. The SHAP analysis, which the paper explicitly frames as independent validation, shows the opposite: the top five SHAP-ranked features are all first-order (S1) coefficients. The Spearman correlation between statistical F-score rankings and SHAP channel importance is ρs = 0.429, p = 0.0969 — a null result the paper reframes as 'moderate positive trend' and 'localized alignment.' The two methods disagree on scattering order and on top electrode sites (ANOVA favors P3; SHAP favors T5, T6, C4). This is a real internal contradiction and the authors need to address it honestly rather than spinning a null result as concordance. The stress-test concern lands squarely on reading the paper; it is not an artifact of selective quotation. The paper does several things well. The WST framework is a legitimate and underused tool for EEG analysis, and applying it with strict LOSO cross-validation and subject-level majority voting is methodologically sound — a genuine improvement over the epoch-level k-fold leakage that pervades this literature. The BH-FDR correction at the subject level, the Bonferroni sensitivity check, and the Cohen's d effect size reporting are all done correctly. The 90.48% accuracy under LOSO on 84 subjects is a credible result for this dataset. The kymatio implementation is reproducible. The soft spots beyond the SHAP contradiction are moderate. The dataset is small (84 subjects, single site, adolescent cohort), and the WST hyperparameters (J=7, Q=(8,1)) are stated as 'optimized' without specifying the search procedure — these are free parameters chosen without documented justification. The z-score artifact rejection threshold is unspecified. The comparison table (Table 4) mixes datasets and validation schemes, making the favorable comparison somewhat misleading. The gamma-band finding (57.9% of significant features) is interesting but hard to interpret without knowing medication status of the cohort, which the paper itself acknowledges as a confound. The paper is for readers interested in interpretable EEG feature extraction for psychiatric diagnosis. The WST + LOSO + SHAP pipeline is a reasonable contribution to methodology. But the biomarker discovery claim is overstated relative to what the evidence supports: the paper has two lines of evidence that point in different directions, and it presents them as convergent. A serious referee should require the authors to either (a) downgrade the S2-dominance claim to 'univariate statistical finding not confirmed by the classifier' or (b) explain why the SHAP-classifier discrepancy does not undermine the biomarker narrative. The classification result itself is fine; the biomarker interpretation is where the work overreaches.","headline":"SHAP analysis contradicts the central biomarker claim: the classifier relies on first-order spectral features while the biomarker narrative rests on second-order scattering dominance from univariate statistics.","tokens_in":18341,"tokens_out":687,"would_cite":false,"duration_ms":160352,"reading_group":"yes","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Schizophrenia's EEG signature lives in cross-frequency amplitude modulation, not raw power","keywords":["Wavelet Scattering Transform","schizophrenia","EEG biomarker","cross-frequency coupling","amplitude modulation","LOSO cross-validation","SHAP explainability","resting-state EEG"],"falsifier":"If second-order WST coefficients no longer dominate the discriminative biomarker set when tested on an independent, multi-site adult dataset with medication-naive patients — or if the P3/gamma-band concentration fails to replicate — the central claim that amplitude modulation disruption is the primary electrophysiological signature of schizophrenia would not hold beyond this cohort.","tokens_in":17848,"feed_emoji":"🧠","tokens_out":1260,"duration_ms":34683,"temperature":0.7,"pith_summary":"This paper argues that the electrophysiological signature of schizophrenia is not primarily a change in how much power each brain rhythm has, but rather a disruption in how brain rhythms modulate each other over time — a property captured by second-order coefficients of the Wavelet Scattering Transform (WST). The WST decomposes an EEG signal into a hierarchy: zeroth-order coefficients capture slow baseline drift, first-order coefficients capture band-specific energy (like a robust spectrogram), and second-order coefficients capture how the amplitude envelope of a fast rhythm is itself modulated by slower rhythms — i.e., cross-frequency coupling. Applying this to resting-state EEG from 84 adolescent subjects (45 schizophrenia, 39 healthy controls), the authors find that 78.5% of the statistically significant biomarkers surviving false-discovery-rate correction are second-order coefficients, concentrated in the gamma band, with electrode P3 (left parietal) as the single most discriminative recording site. Under strict leave-one-subject-out cross-validation — which prevents the temporal data leakage the authors argue inflates prior work — a Random Forest classifier on the full WST feature space achieves 90.48% accuracy (AUC 0.934, sensitivity 95.56%). SHAP explainability analysis independently confirms that the model's decisions center on the same parietal-occipital and right-frontal regions identified by the statistical biomarker analysis. The paper's central claim is that schizophrenia is measurable as a disorder of multi-scale temporal coordination — disrupted amplitude modulation dynamics — rather than a disorder of altered mean spectral power, and that the WST provides a principled, interpretable feature space for capturing this.","feed_headline":"Schizophrenia EEG signature traced to cross-frequency modulation, not raw power","feed_subtitle":"Second-order wavelet scattering coefficients dominate schizophrenia biomarkers, pointing to disrupted amplitude modulation as the core elect","key_machinery":"Wavelet Scattering Transform (WST): a hierarchical signal decomposition that cascades wavelet convolutions with modulus and averaging operations. Zeroth-order coefficients (S0) capture local DC baseline; first-order (S1) capture band-limited spectral energy; second-order (S2) capture amplitude modulation of one frequency band by another (cross-frequency coupling). Configured here with invariance scale J=7 (1-second window) and quality factors Q=(8,1), yielding 176 scattering paths per epoch across 16 EEG channels.","core_discovery":"The dominant discriminative biomarkers for schizophrenia in resting-state EEG are second-order wavelet scattering coefficients — features that quantify cross-frequency amplitude modulation — rather than first-order spectral energy features. Of 1,255 features surviving Benjamini–Hochberg false-discovery-rate correction, 78.5% are second-order, 57.9% fall in the gamma band, and the left parietal electrode P3 produces the most top-ranked biomarkers. Schizophrenia patients show a systematic, near-universal reduction (99.8% of significant features with negative effect sizes) in scattering energy, with the largest deficit at P3 (Cohen's d = −1.19, a 28.7% drop). This pattern — cross-frequency-coug","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Cross-frequency amplitude modulation — not static power — marks schizophrenia EEG","Schizophrenia EEG deficits concentrate in gamma-band cross-frequency coupling at P3","Wavelet scattering reveals amplitude modulation as core schizophrenia EEG signature","Second-order wavelet features dominate schizophrenia EEG biomarkers under strict LOSO","Schizophrenia shows near-universal reduction in cross-frequency scattering energy at P3"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The entire framework is validated on a single dataset of 84 adolescent subjects with no medication controls, no multi-site replication, and no adult cohort. The claim that disrupted amplitude modulation is the primary electrophysiological signature of schizophrenia rests on this one sample. Additionally, the SHAP and ANOVA biomarker rankings show only moderate concordance (Spearman ρ = 0.429, p = 0.097), meaning the two methods agree on the general spatial pattern but not on ","fun_headline_variants_meta":{"raw":{"variants":["Cross-frequency amplitude modulation — not static power — marks schizophrenia EEG","Schizophrenia EEG deficits concentrate in gamma-band cross-frequency coupling at P3","Wavelet scattering reveals amplitude modulation as core schizophrenia EEG signature","Second-order wavelet features dominate schizophrenia EEG biomarkers under strict LOSO","Schizophrenia shows near-universal reduction in cross-frequency scattering energy at P3"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":809,"prompt_tokens":732,"completion_tokens":77,"prompt_tokens_details":null},"tokens_in":732,"tokens_out":77,"duration_ms":68048,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-07T19:46:35.609847+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If second-order WST coefficients no longer dominate the discriminative biomarker set when tested on an independent, multi-site adult dataset with medication-naive patients — or if the P3/gamma-band concentration fails to replicate — the central claim that amplitude modulation disruption is the primary electrophysiological signature of schizophrenia would not hold beyond this cohort.","supporting_citations":[],"review_version":1}