{"id":"e8465034-db63-41df-b6d3-ea1d7a041f43","arxiv_id":"2608.10887","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A null-distribution-aware framework, including a Fisher-normal semi-parametric model and the NNTS metric, makes neural tracking correlations comparable across features and models.","lead":"This paper shows that raw correlation scores used to measure how well brain signals track speech can mislead, because easy-to-reconstruct features score high by accident. The authors propose a statistically normalized score, NNTS, that corrects for each feature's own null distribution, and show it reverses conclusions in a 121-person EEG study.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Variance-rescaling in Eq. (3) is validated only up to 20 s windows yet Table 1 extrapolates to ~93 s; the assumed constancy of the autocorrelation inflation factor across window lengths is the load-bearing unvalidated step.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: Eq. (3) assumes that the ratio of the actual null variance of Fisher-transformed correlations to the theoretical independent-sample variance V(N) is independent of window length. This assumption is not explicitly stated as testable, is validated only on 1–20 s windows for a single story and EEG setup, and is then used in Table 1 to predict measurement times up to about 93 s. My independent reading of Sections 4.3, 5.1, and Table 1 confirms that the extrapolation is central to the '3–5 min across window lengths' claim and to the practical predictions of recording time. The concern is concrete and falsifiable, and it does not require rejecting the rest of the framework: the semi-parametric model and NNTS are largely independent of the extrapolation step and are supported by the provided code, data, and validation. A conditional verdict is appropriate because the core methodology is sound within the tested range, but the headline extrapolation claim should be either empirically supported at longer window lengths or explicitly qualified as unvalidated. I found no reason to move to ACCEPT or REJECT; the correct adjustment is to keep the conditional verdict and require the additional long-window validation.","tokens_in":30379,"tokens_out":1773,"duration_ms":19153,"concrete_test":"Using the same 876 s story and the 121-participant dataset, generate misalignment null correlations for target window lengths of 30, 45, 60, and 90 s (where the story length permits at least a few hundred permutations per participant), compute the empirical Fisher-domain variance sigma_hat^2 at each length, and compare it with the Eq. (3) prediction extrapolated from the 5 s baseline. If the ratio sigma_hat^2 / V(N) is approximately constant across 5–90 s for all eight features, the extrapolation holds; if it drifts by more than, say, 10% over that range, then Table 1's predicted measurement times—especially the 93 s smallband-envelope entry—are not supported and should be relabeled as unvalidated extrapolations or re-estimated at the actual target lengths.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central efficiency claim—that 3–5 min of data can reliably estimate significance levels across window lengths—depends on Eq. (3), which rescales a Fisher-transformed null variance estimated at one window length by the ratio V(N_target)/V(N_base), where V(N) is the theoretical independent-sample variance. The implicit assumption is that the empirical null variance equals c·V(N) with the same multiplicative constant c at every window length. This constant absorbs the autocorrelation inflation of the effective sample size. The paper validates the resulting extrapolation only for window lengths 1, 2, 5, 10, and 20 s (Figures 11–12), and even there the errors grow at the shortest window. Table 1 then uses the extrapolation to predict measurement times up to 93.4 s for the smallband envelope and 43 s for punctuation onset, far outside the validated range. If the autocorrelation inflation factor changes with window length—which is plausible because the decoded signal and the stimulus feature have different spectral properties and because the ridge-regularized decoder is retrained per window length—the extrapolated significance levels and the predicted recording times in Table 1 would be biased. The manuscript does not state this as a testable assumption nor provide evidence for it beyond the 1–20 s range. This is not an internal inconsistency, but it is the weakest support for the paper's headline practical recommendation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that raw Pearson correlations used to quantify neural tracking of natural stimuli are not comparable across stimulus features or models, because their null distributions differ. It compares four surrogate-data methods for constructing null distributions (random shuffling, circular shifting, phase scrambling, and stimulus-response misalignment), argues that each encodes a different null hypothesis, and adopts misalignment as the most appropriate for content-specific neural tracking. The paper then introduces a semi-parametric model in which Fisher-transformed null correlations are treated as zero-mean normal with an empirically estimated variance, claims that this yields accurate significance levels with far fewer permutations than the empirical null distribution, and extends this to predict significance levels at other window lengths by rescaling the estimated variance with the theoretical independent-sample variance V(N). On this basis it proposes the null-normalized tracking score (NNTS), shows a mathematical and empirical relationship between NNTS and match-mismatch accuracy, and applies the framework to EEG data from 121 participants listening to a continuous story, concluding that raw correlations would reverse the ranking of features relative to the NNTS-based ranking.","tokens_in":30649,"tokens_out":5084,"duration_ms":52687,"significance":"If the central claims hold, the paper provides a practically useful and statistically principled framework for a widespread analysis choice in EEG/MEG speech tracking: it offers a way to estimate significance levels efficiently, to compare features and models on a common scale, and to connect correlation-based tracking metrics to match-mismatch accuracy. The paper is unusually thorough in its empirical validation: it uses 100,000 permutations as a ground-truth null, reports both correlation-domain and percentile-domain errors, quantifies bias and variance through resampling, and verifies the NNTS-to-match-mismatch relationship with a reported 0.37 percentage point error. The theoretical V(N) rescaling is a parameter-free derivation, and the code, data, and toolbox availability statements are exemplary. The main limitation is that the cross-window extrapolation, which underpins the headline efficiency claim and the predicted measurement times, is validated only for window lengths up to 20 s and is then applied to far longer windows; this is a load-bearing issue rather than a cosmetic one, but it is addressable with additional validation or with appropriately restricted claims.","major_comments":[{"comment":"The extrapolation formula in Eq. (3) rests on the assumption that the empirical null variance of the Fisher-transformed correlations equals c*V(N) with the same multiplicative constant c at every window length, where c absorbs the autocorrelation-induced inflation of the effective sample size. The manuscript validates the resulting extrapolation only for window lengths of 1, 2, 5, 10, and 20 s (Figures 11-12), and Table 1 then extrapolates to 93.4 s for the smallband envelope and 43.0 s for punctuation onset, far beyond the validated range. Because the ridge-regularized decoder is retrained at each window length and because the spectral properties of the stimulus feature and the reconstruction differ, the inflation factor could plausibly change with N, which would bias the extrapolated significance levels and the predicted measurement times. The manuscript should either validate the constancy of the inflation factor at longer window lengths, provide a bound on its variation, or explicitly restrict the extrapolation claim to the validated range.","section":"4.3 (Eq. (3)); Table 1"},{"comment":"The predicted measurement times in Table 1 additionally assume that the mean correlation for each feature is approximately constant across window lengths. This assumption is stated in Section 5.1 with a citation to Lopez-Gordo et al. (2025), but that reference concerns unsupervised accuracy estimation in auditory attention decoding rather than a direct demonstration for the speech features and the 121-participant dataset used here. Since the predicted time to significance depends as much on this assumed constancy as on the null-variance extrapolation, the table should either provide supporting evidence that the mean correlations are stable for these features or present the assumption explicitly with a discussion of how its failure would change the predictions.","section":"5.1 (Table 1)"}],"minor_comments":[{"comment":"The caption contains a typo: 'readibility' should be 'readability'.","section":"Figure 3 caption"},{"comment":"The text contains a typo: 'non-trival problem' should be 'non-trivial problem'.","section":"Section 1"},{"comment":"The appendix appears to contain a duplicated and incomplete passage: the sentence beginning 'Given estimated variance of the per-window raw correlations (with mean assumed0):' is immediately followed by a second, nearly identical introduction of the same problem; this passage should be rewritten as a single clean statement.","section":"Appendix B"},{"comment":"The x-axis label 'mean predicted MM accuracy based on NNTS [%]' is ambiguous; the figure would be clearer if the axes were labeled 'Predicted match-mismatch accuracy [%]' and 'Observed match-mismatch accuracy [%]'.","section":"Figure 16(b)"},{"comment":"The key take-away box ends with an ellipsis and appears to omit the end of the sentence ('...across window lengths, features, ...'); the sentence should be completed or the ellipsis removed.","section":"Key take-away #4"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the manuscript is methodologically strong and the empirical evaluation is unusually thorough, including large-permutation ground truths, multiple error metrics, and resampling-based bias and variance analyses. The main concern is the unvalidated extrapolation of the variance-rescaling assumption beyond 20 s, which directly supports the headline recommendation that 3-5 minutes of data suffice across window lengths and the predicted measurement times in Table 1. This is fixable by additional validation or by narrowing the claim, so I recommend major revision rather than rejection. The paper is well within the scope of a methods-oriented neuroimaging or EEG journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. The genuine new content is the integration: it makes explicit that permutation procedures for neural-tracking nulls each encode a different null hypothesis, argues for misalignment as the most defensible default, and then builds a semi-parametric null model (normal after Fisher transform) that is validated against a 100,000-permutation ground truth. On top of that, it introduces the null-normalized tracking score and proves a clean mathematical link to match-mismatch accuracy, with the relationship verified empirically at 0.37 percentage point error. The code, data, and toolbox are public, and the empirical work is unusually thorough: resampling analyses, bias/variance checks, and multiple error metrics. This is a solid methodological contribution, not a repackaging of known parts.\n\nThe soft spot is exactly where the stress-test note points. Equation (3) rescales the estimated null variance by the theoretical ratio V(N_target)/V(N_base), which assumes the autocorrelation inflation factor is constant across window lengths. That is plausible but not tested; the validation only covers 1–20 s windows, and Table 1 extrapolates to 93.4 s. The authors do call the Table 1 numbers \"only illustrative\" and acknowledge the single-dataset limitation in the conclusion. So the overreach is contained: the core significance-level estimation at 1–20 s is well supported, and the extrapolation is a practical convenience that may well hold, but the headline \"across window lengths\" in Key take-away #4 goes beyond what is shown. I would not call this a fatal flaw; I would call it an unvalidated extrapolation that should either be tested on longer windows or explicitly labeled as such in the abstract and take-away boxes.\n\nThe missing check is not hard to design: record or simulate longer windows on the same story, or re-analyze existing longer recordings, and compare the extrapolated significance levels against the empirical nulls. The authors seem capable of doing this quickly.\n\nWho is this for? Anyone who uses neural tracking correlations to compare features, models, or recording setups, and methodologists working on surrogate nulls. It deserves a serious referee. My recommendation: accept for review, and in the review request that the extrapolation claim be either validated beyond 20 s or softened to the tested range.","headline":"A principled, well-validated framework for interpreting neural-tracking correlations, with one honest gap: the window-length extrapolation is used beyond its validated range.","tokens_in":31189,"tokens_out":2145,"would_cite":true,"duration_ms":68363,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Neural-tracking correlations only become comparable when scored against their null distribution, and 3–5 minutes of data suffice to build it.","keywords":["neural tracking","null distribution","Fisher transformation","permutation testing","significance level","EEG","match-mismatch accuracy"],"falsifier":"Measure the Fisher-transformed null variance from misaligned stimulus-response pairs at several window lengths (for example 1, 2, 5, 10, and 20 seconds), divide each by the theoretical independent-sample variance V(N), and check whether the resulting ratios are constant across N; if they drift, the extrapolation formula is systematically biased and the predicted significance levels and recording times are unreliable.","tokens_in":30119,"feed_emoji":"🧠","tokens_out":6627,"duration_ms":57566,"temperature":0.7,"pith_summary":"This paper argues that the Pearson correlation commonly used to quantify neural tracking of natural stimuli cannot be compared across stimulus features, models, or settings without first knowing the distribution of correlations expected by chance. It shows that the standard ways of building that null distribution are not interchangeable, each encoding a different null hypothesis, and adopts stimulus-response misalignment as the principled default. The central methodological contribution is a semi-parametric model, a normal distribution applied to Fisher-transformed correlations, that yields accurate significance levels from about 1000 permutations (3-5 minutes of data) and extrapolates them across analysis window lengths. The authors then introduce the null-normalized tracking score (NNTS), which places features and models on a common scale and connects directly to match-mismatch accuracy. Applied to EEG from 121 participants listening to continuous speech, the framework reverses conclusions drawn from raw correlations, demoting a narrowband envelope that had looked like one of the best features.","feed_headline":"3-5 minutes of data fix misleading neural-tracking correlations","feed_subtitle":"A Fisher-transform null model estimates significance at any window length and ranks EEG features on one scale.","key_machinery":"The carrying mechanism is the Fisher transform z = artanh(r) applied to null correlations, modeled as z ~ N(0, $sigma_z^{2}$) with the variance estimated empirically from misaligned stimulus-response pairs rather than set to 1/(N-3). The variance is then rescaled across window lengths using the ratio V(N_target)/V(N_base) of the theoretical independent-sample variances, and the significance level is recovered by applying the inverse transform tanh to a normal percentile. NNTS divides Fisher-transformed real correlations by the null standard deviation (window level) or by the pooled standard deviation of real and null correlations (participant level), making NNTS a d-prime-like sensitivity index.","core_discovery":"The paper's central claim is that the Pearson correlation between a decoded neural response and a stimulus feature is not a meaningful performance metric by itself, because its scale depends on the statistical properties of the feature and the decoder. The authors establish that the right way to interpret a tracking correlation is against the null distribution generated by stimulus-response misalignment, and that this null distribution is accurately and efficiently captured by a zero-mean normal distribution after the Fisher transform. From that model they derive significance levels at any window length and a null-normalized tracking score (NNTS) that puts different features and models on a common, unbounded scale, equals d-prime, and has a direct mathematical relation to match-mismatch accuracy. In EEG data from 121 listeners, the framework reverses the ranking suggested by raw correlations: a 1-1.1 Hz smallband envelope has among the highest raw correlations but the lowest NNTS.","pith_inferences":["If the variance-ratio assumption holds for other stimuli and recording modalities, the same 3-5 minute recipe could be used to design clinical or hearing-aid protocols that pre-specify recording duration from a target correlation.","Because NNTS depends only on the distributions of real and null correlations, it should also apply to non-linear or deep-learning decoders, whose raw correlation scales are even harder to interpret; this is a testable extension the paper does not run.","The equivalence between NNTS and match-mismatch accuracy suggests that existing match-mismatch pipelines could switch to NNTS to gain continuous, unbounded resolution without changing the underlying experiment.","A direct check of the variance-ratio assumption across window lengths on non-speech stimuli would tell whether the extrapolated significance levels and recording-time predictions transfer beyond the one dataset used here."],"forward_implications":["With only 3-5 minutes of data (about 1000 misaligned permutations), significance levels can be estimated reliably for any window length and feature, removing the need for tens of thousands of permutations.","Features and models can be compared on a common scale via NNTS; in the 121-participant EEG analysis, the smallband envelope drops from the top raw correlation to the worst NNTS, while the envelope, acoustic edge, and phoneme onset form a top tier.","NNTS is mathematically equivalent, up to a monotone transform, to match-mismatch accuracy, with predicted accuracy within 0.37 percentage points of measured accuracy, and it saturates less because it is unbounded.","Significance levels can be extrapolated across window lengths, allowing prediction of the measurement time needed for a target correlation to reach significance: about 17 seconds for the envelope, acoustic edge, and phoneme onset versus 93 seconds for the smallband envelope.","The framework is agnostic to the choice of permutation method and to stimulus modality, so the same pipeline should apply to music, video, MEG, or ECoG data, subject to empirical confirmation."],"supporting_citations":[{"why":"Supplies the variance-stabilizing normalizing transform z = artanh(r) that is the backbone of the null-distribution model.","marker":"Fisher (1921)"},{"why":"Supplies the cumulant-based expression V(N) for the Fisher-transform variance under the null, which underlies window-length extrapolation.","marker":"Fouladi and Steiger (2008)"},{"why":"Grounds the argument that each surrogate method encodes a different null hypothesis and must preserve the relevant signal statistics.","marker":"Lancaster et al. (2018)"},{"why":"Provides the Z-score idea that NNTS builds on and a misalignment-based alternative for generating null correlations.","marker":"MacIntyre et al. (2026)"},{"why":"Defines the match-mismatch task whose accuracy the paper proves equivalent to NNTS.","marker":"de Cheveigné et al. (2021)"},{"why":"Supplies the performance-curve modeling approach used to rescale variances across window lengths.","marker":"Geirnaert et al. (2025)"},{"why":"Supplies the linear backward-decoding model and validation framework used to generate the tracking correlations.","marker":"Crosse et al. (2021)"},{"why":"Provides the SparrKULee EEG dataset underlying the 121-participant analysis.","marker":"Accou et al. (2024)"}],"fun_headline_variants":["Raw neural tracking correlations mislead; null model fixes it","3-minute EEG data predicts neural tracking significance","Null-normalized score reorders EEG features, reverses conclusions","Fisher-transform null model gives significance from 3-minute EEG"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that autocorrelation inflates the spread of null correlations by the same factor at every window length, so a variance measured at one window length can be extrapolated to all others.","fun_headline_variants_meta":{"raw":{"variants":["Raw neural tracking correlations mislead; null model fixes it","3-minute EEG data predicts neural tracking significance","Null-normalized score reorders EEG features, reverses conclusions","Fisher-transform null model gives significance from 3-minute EEG"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000742,"raw_usage":{"total_tokens":3345,"prompt_tokens":1014,"completion_tokens":2331,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":630,"completion_tokens_details":{"reasoning_tokens":2267}},"tokens_in":630,"tokens_out":2331,"duration_ms":17180,"temperature":1.0,"reasoning_tokens":2267,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:59:49.942859+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the Fisher-transformed null variance from misaligned stimulus-response pairs at several window lengths (for example 1, 2, 5, 10, and 20 seconds), divide each by the theoretical independent-sample variance V(N), and check whether the resulting ratios are constant across N; if they drift, the extrapolation formula is systematically biased and the predicted significance levels and recording times are unreliable.","supporting_citations":[{"cited_title":"2008 , doi=","cited_arxiv_id":null,"evidence_quote":"Supplies the cumulant-based expression V(N) for the Fisher-transform variance under the null, which underlies window-length extrapolation."},{"cited_title":"2025 , volume=","cited_arxiv_id":null,"evidence_quote":"Supplies the performance-curve modeling approach used to rescale variances across window lengths."},{"cited_title":"and Zuk, Nathaniel J","cited_arxiv_id":null,"evidence_quote":"Supplies the linear backward-decoding model and validation framework used to generate the tracking correlations."}],"review_version":1}