{"id":"4ac44d72-ea8f-403a-869e-d31d1b6c7e31","arxiv_id":"2507.04442","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Entropy density and Lempel-Ziv complexity applied to binarized fMRI signals identify task-active cortical regions and cluster brain areas into functional groups, but the approach rests on a heuristic threshold and lacks error bars.","lead":"Using entropy scores computed from binarized brain-scan signals, the authors try to identify which cortical regions activate during motor, memory, emotion, and language tasks, and to group regions by shared patterns. The results mostly line up with known brain areas, but the method hinges on a heuristic threshold and lacks error bars or a direct ground-truth validation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The entropy-based activation classification rests on an uncalibrated, data-derived threshold; without independent validation, the central claim that entropy measures identify active regions is unsupported.","rationale":"The reader's weakest assumption identifies exactly the load-bearing flaw: the active/inactive partition is set by a heuristic threshold fit to the same data, with no independent calibration or statistical validation. My reading of the full text confirms that all subsequent claims—including the complexity-entropy interpretation, the task comparisons, and the dendrogram 'functional groupings'—inherit the validity of this threshold. The paper itself states that binarization is a drastic loss of information and that the results serve as a consistency check, while directed connectivity is deferred; this narrows the actual deliverable to activation identification, which is precisely the step that is not validated. I agree with the reader's rejection: as presented, the central claim is not supported by the evidence. A concrete test using standard GLM-based activation maps from the same HCP dataset could settle the question, and the verdict would be changed if such validation were provided, but with the current manuscript the appropriate outcome is unchanged rejection.","tokens_in":23725,"tokens_out":3329,"duration_ms":43713,"concrete_test":"On the same HCP task-fMRI data, define reference active ROIs from the official task contrast maps (e.g., the motor task's fixation-contrast z-statistics, thresholded at p<0.05 family-wise error and summarized per Glasser ROI). Then recompute h_LZ and E_LZ exactly as in the paper and evaluate the ROC AUC of the entropy-based active/inactive classification against this ground truth. Repeat with k = 5, 10, 15, 20 and with phase-randomized surrogate BOLD time series to build null distributions. If AUC is not significantly above 0.5, or if the active-region lists shift by more than 20% across k, the threshold-based method is not a reliable indicator of neural activation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's core result is the partition of the 360 ROIs into active and inactive sets from the complexity-entropy map. That partition is made by a threshold defined as the maximum of the mean residual of each point's ten nearest h_LZ neighbors with respect to a linear fit (Section III and Supplementary B). This threshold is computed from the very data it labels; changing k or the residual criterion changes every active-region list, and no error bars, cross-validation, or comparison to task activation maps from the same dataset are provided. The only external check is qualitative agreement with the literature, and the paper itself calls the results a 'consistency check' and defers directed connectivity analysis to a forthcoming paper. The abstract's claim of detecting dynamics 'without relying on parameters, models, or prior assumptions' is also contradicted by the method's own choices: mean-value binarization (explicitly admitted to be a drastic information-losing procedure), the 10-neighbor window, and the residual-maximum threshold. Because the dependent variable (activation) is defined by an unvalidated, data-dependent rule, the central claim that lower h_LZ and higher E_LZ mark task-active regions is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using Lempel-Ziv estimates of entropy density (h_LZ) and effective measure complexity (E_LZ), computed from mean-binarized fMRI time series of 360 Glasser regions, to identify task-active regions in HCP task-fMRI data (motor, working memory, emotion, and language, plus resting state). Active/inactive status is assigned by a threshold that is the maximum of the mean residual of each point's ten nearest h_LZ neighbors with respect to a linear fit to the (h_LZ, E_LZ) cloud. The paper reports that active regions have lower h_LZ and higher E_LZ, interprets these as less unpredictable and more patterned, and uses an undirected LZ-information-distance with dendrograms to claim functional groupings. The Conclusion states that the present results are a consistency check and that directed connectivity analysis is left to a forthcoming paper.","tokens_in":23935,"tokens_out":7624,"duration_ms":70637,"significance":"If the activation classification were validated, the approach could offer a relatively simple, entropy-based way to screen task-active regions from fMRI without an explicit hemodynamic or network model, and the use of a large public dataset (N=153, HCP) with multiple tasks is a genuine strength. The manuscript is also transparent about its limitations, including information loss from binarization and the undirected nature of the distance measure. However, the central scientific claim is not currently established because the activation threshold is fit to the very data it labels, with no independent calibration, cross-validation, or null-model testing. The paper therefore reads as a promising but unvalidated proposal rather than a demonstrated result.","major_comments":[{"comment":"The threshold separating active from inactive ROIs is defined as the maximum of the mean residual of each point's ten nearest h_LZ neighbors with respect to a linear least-squares fit to the same (h_LZ, E_LZ) data it partitions. Consequently, the statement that active regions have lower h_LZ and higher E_LZ is true by construction: the left-of-threshold set is selected by that property. The overlap with prior literature in Supplementary Table I is a qualitative sanity check, not an independent validation. To support the claims, the authors must calibrate or validate the threshold against an independent activation measure from the same data (e.g., HCP task-fMRI contrast maps) and provide stability analyses with respect to k, the fit model, and the binarization threshold, together with confidence intervals or null-model results.","section":"Section III, Figure 3, and Supplementary B"},{"comment":"The Abstract's claim that the tools detect dynamics 'without relying on pre-established parameters, models, or prior assumptions' is contradicted by the method's own choices: mean-value binarization per ROI (which the authors admit is 'a drastic procedure resulting in the loss of information'), the nearest-neighbor count k=10, the linear-fit model, and the residual-maximum threshold rule. These are ad hoc parameters that materially affect the classification; the parameter-free claim should be removed or substantially qualified.","section":"Section II (Data preprocessing) and Abstract"},{"comment":"The title promises 'connectivity paths,' but the paper explicitly defers directed connectivity analysis to a forthcoming study and states that 'it is possible to assign a time arrow to information distances, this was not done in the present analysis.' The LZ-distance is an undirected, time-symmetric similarity measure, and the dendrograms are descriptive hierarchical clusterings without statistical support. No evidence of information flow or directed paths is presented. The title should be revised, or the directed connectivity analysis must be included and validated.","section":"Title, Section III.B, and Conclusion"},{"comment":"The paper reports one (h_LZ, E_LZ) tuple per ROI but does not describe how the N=153 subjects' time series are combined (concatenation, averaging, or per-subject analysis followed by pooling), nor does it provide error bars or subject-level variability. Without this information, the reader cannot assess whether the separation left and right of the threshold is statistically meaningful. The authors should specify the subject-pooling scheme and provide per-subject or bootstrap confidence intervals for h_LZ and E_LZ.","section":"Section II (Methods) and Figure 5"},{"comment":"The reported percentages of active ROIs are inconsistent with the enumerated lists. For example, the motor task reports 8.61% of 360 regions (about 31 regions) but the listed entries correspond to 27 regions; the memory task reports 11.4% (about 41) but lists 23; the language task reports 13.61% (about 49) but lists 20. These discrepancies make the active-region results difficult to reproduce and must be resolved.","section":"Section III (active region lists)"}],"minor_comments":[{"comment":"A stray '1.' formatting artifact appears before 'Areas that are identified as active in this study but are not typically considered activated,' which should be removed.","section":"Section III"},{"comment":"The dendrogram labels are too small to read; a zoomed panel or a separate list of cluster members would improve clarity.","section":"Figure 6"},{"comment":"The resting-state 'active' set is selected manually from the low-h_LZ grouping ('22 regions can be considered in the far left end'), and this ad hoc criterion should be clearly distinguished from the threshold-based task analyses.","section":"Section III.2"},{"comment":"Reference 28 is incomplete: 'Wu-Minn and HCP' is not a proper author list for the HCP reference manual.","section":"References"},{"comment":"The in-text arXiv identifier and submission date ('arXiv:2507.04442v1 [q-bio.NC] 6 Jul 2025') should be removed from the body text.","section":"Page 2"},{"comment":"The asymmetry index is reported without a significance test or confidence interval, so the resting-state symmetry contrast is not statistically supported.","section":"Section III.2"}],"recommendation":"major_revision","confidential_remarks":"The core problem is the circular threshold: the classification rule is fit to the same data it labels, and no independent validation is provided. This is fixable within the manuscript's scope if the authors validate against HCP task-fMRI contrast maps, report stability and error analyses, and revise the title to match the content (removing the 'connectivity paths' claim or actually delivering it). If such validation cannot be provided, I would not support publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nThe short version: this is a genuinely exploratory application of LZ entropy and complexity to task fMRI, but the activation classification rests on a threshold drawn from the same data, and without independent validation the central claim is not yet supported.\n\nWhat the paper does well: it places each Glasser ROI in the complexity-entropy plane (h_LZ vs E_LZ) for four HCP tasks plus resting state, and the qualitative pattern is plausible—motor and somatosensory areas for the motor task, visual and fusiform areas for emotion, auditory and language areas for language, and a clear interhemispheric symmetry during tasks that breaks down in rest. The LZ-distance dendrograms show that active regions cluster together, and the authors are honest about calling the results a consistency check and deferring directed connectivity to a future paper.\n\nThe soft spot is the one your reader already flagged, and it lands. The vertical line separating active from inactive is the maximum of a mean-residual curve computed from ten-nearest-neighbor deviations from a linear fit to the same complexity-entropy cloud. There are no error bars, no null model, no cross-validation, and no comparison to a standard GLM activation map from the same HCP data. So the statement that active regions have lower h_LZ and higher E_LZ is true by construction: 'active' is defined as being on the low-entropy side of that data-derived threshold. Binarization with the mean is admitted to be drastic, and the abstract's 'no parameters, models, or prior assumptions' is overstated, because the method requires the binarization threshold, k, and the residual criterion.\n\nNone of this makes the paper nonsense. The qualitative agreement with previous studies is a genuine hint that entropy measures carry task-relevant information, but it is a hint, not a validated result. The missing piece is any statistical argument that the left-hand cluster is not an artifact of the fitting procedure.\n\nThis paper is for people who work on entropy-based fMRI or who want a concrete object lesson in why data-derived thresholds need external benchmarks. It deserves a serious referee—the combinatorial application is new and the HCP data are public—but a referee would need to require a comparison against existing activation maps and a stability check of the threshold with respect to k and the binarization rule. As it stands, I would not accept it; I'd reject with an invitation to revise.\n\nRegards,","headline":"Entropy-based fMRI activation mapping with a data-derived threshold that needs external validation before the central claim is credible.","tokens_in":24488,"tokens_out":3747,"would_cite":false,"duration_ms":35774,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Entropy measures alone can pick out task-active brain regions from binarized fMRI signals.","keywords":["entropy density","Lempel-Ziv complexity","effective measure complexity","complexity-entropy map","task-based fMRI","brain activation","information distance","resting-state asymmetry"],"falsifier":"Recompute the active-region lists after replacing each task time series with a phase-randomized surrogate that preserves the power spectrum but destroys nonlinear pattern structure; if the same regions still land left of the threshold, the entropy measures are responding to trivial autocorrelation rather than task-related dynamics.","tokens_in":23493,"feed_emoji":"🧠","tokens_out":11188,"duration_ms":111752,"temperature":0.7,"pith_summary":"The paper claims that two Lempel-Ziv-based entropy measures, the entropy density $h_{LZ}$ and the LZ-effective complexity $E_{LZ}$, computed from mean-binarized fMRI time series, separate active from inactive cortical regions during motor, working-memory, emotion, and language tasks. This matters because the separation is achieved without a hemodynamic model, a network model, or pre-set parameters, so the same pipeline could be used as an exploratory probe for tasks where the expected activations are unknown. The paper also claims that task conditions produce symmetric entropy-density maps across hemispheres, while the resting state is strongly asymmetric, and that LZ-distance dendrograms cluster functionally related regions. Directed connectivity between regions, which the title promises, is explicitly deferred to a forthcoming paper.","feed_headline":"Lower entropy marks active brain regions in fMRI scans","feed_subtitle":"Task-active cortex shows less random, more patterned signals across four tasks, with no model or preset parameters.","key_machinery":"The carrying object is the complexity-entropy map built from Lempel-Ziv factorization of each binarized fMRI time series: entropy density $h_{LZ}$ estimates randomness through the number of new patterns in the exhaustive history, and LZ-effective complexity $E_{LZ}$ estimates memory by comparing the original sequence with randomly shuffled versions. A third quantity, the LZ-distance $d_{LZ}$, estimates the normalized information distance between pairs of regions and is used to build distance matrices and Neighbor-Joining dendrograms. The underlying identity is the coding theorem that the Lempel-Ziv complexity divided by $N/\\log N$ converges to the entropy density for ergodic sources, which licenses the estimates on the short, discretized fMRI sequences.","core_discovery":"On the paper's own terms, active regions present lower entropy density and higher effective complexity than inactive regions: their binarized BOLD (blood-oxygenation-level-dependent) time series are less unpredictable and more patterned, while inactive regions keep a noisy background. Plotting every cortical region as a point in the $(h_{LZ}, E_{LZ})$ plane yields a linear trend with a data-derived residual threshold, and the regions to the left of that threshold match large parts of the known activation maps for each task, including visual areas for visually cued tasks, motor and sensory cortices for movement, face- and object-selective areas for emotion and memory, and auditory, language, and arithmetic areas for language. The LZ-distance dendrograms group active regions together at low hierarchy levels, and the resting state shows an asymmetry index of 0.41 compared with 0.078 to 0.11 across tasks.","pith_inferences":["Extension: the same two-number entropy summary could be tested on other coarse physiological recordings such as EEG or calcium imaging, where binarization is less drastic and ground-truth activation is available; the paper does not run that test.","Extension: the resting-state asymmetry index suggests a testable prediction that entropy-density asymmetry grows with mind-wandering or low arousal and shrinks with focused external attention, which the authors do not investigate.","Robustness check: rerunning the full pipeline with different window sizes for the residual threshold (for example $k=5$ or $k=15$ instead of $k=10$) and with median rather than mean binarization would show how much of the active-region list depends on those choices.","The authors defer directed connectivity, but the time-symmetric LZ-distance matrices could be extended to time-shifted or conditional versions that assign a direction, which would test whether the 'connectivity paths' of the title are actually recoverable."],"forward_implications":["If active regions really are the low-entropy, high-complexity points in the $(h_{LZ}, E_{LZ})$ plane, then task-activation maps can be produced directly from binarized fMRI time series without fitting a hemodynamic response or specifying a network model.","The consistency of the visual, motor, face, and language clusters across tasks means the same unparameterized measures can be applied to new tasks where the expected activation pattern is not known.","The interhemispheric symmetry of entropy density during tasks and its breakdown at rest suggests a simple scalar summary of hemispheric coordination that could be tracked across cognitive states.","Because the method also flags regions not usually reported as active and misses subcortical structures, it offers a complementary activation map rather than a replacement for conventional task contrasts.","The finding that active regions appear far from inactive ones in LZ-distance while inactive regions cluster tightly suggests that shared pattern redundancy, not just correlation, is a usable signal for functional grouping."],"supporting_citations":[{"why":"supplies the public task-fMRI dataset used for all analyses.","marker":"20"},{"why":"documents the scanning protocols and standard preprocessing pipeline for that dataset.","marker":"28"},{"why":"provides the 360-region cortical parcellation used to define regions of interest.","marker":"29"},{"why":"establishes that Lempel-Ziv-based entropy estimates are reliable on short symbolic sequences, justifying $h_{LZ}$.","marker":"30"},{"why":"gives the shuffling-based estimate used for the effective measure complexity $E_{LZ}$.","marker":"32"},{"why":"defines the normalized information distance that the LZ-distance $d_{LZ}$ estimates.","marker":"33"},{"why":"provides the Lempel-Ziv factorization procedure used to estimate Kolmogorov complexity in the distance calculation.","marker":"35"},{"why":"supports the interpretation that inactive regions do not suppress background noise and therefore appear more random.","marker":"36"},{"why":"supplies the task-contrast activation findings used as the reference for which regions are expected to be active.","marker":"37"},{"why":"states the theorem linking Lempel-Ziv complexity to entropy density for ergodic sources, the theoretical basis of the main estimate.","marker":"58"}],"fun_headline_variants":["Entropy dips reveal active brain networks","Lower entropy marks brain's working regions","Task-driven entropy shifts expose neural links","Entropy maps brain connectivity without models","Active cortex shows lower entropy, more pattern"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the data's own residual curve, with its ten-neighbor window, gives a true boundary between active and inactive regions, with no independent ground truth, and that mean-value binarization preserves the task-relevant dynamics even though it is admitted to discard information.","fun_headline_variants_meta":{"raw":{"variants":["Entropy dips reveal active brain networks","Lower entropy marks brain's working regions","Task-driven entropy shifts expose neural links","Entropy maps brain connectivity without models","Active cortex shows lower entropy, more pattern"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1441,"prompt_tokens":891,"completion_tokens":550,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":488}},"tokens_in":507,"tokens_out":550,"duration_ms":5925,"temperature":1.0,"reasoning_tokens":488,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:47:54.838195+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the active-region lists after replacing each task time series with a phase-randomized surrogate that preserves the power spectrum but destroys nonlinear pattern structure; if the same regions still land left of the threshold, the entropy measures are responding to trivial autocorrelation rather than task-related dynamics.","supporting_citations":[{"cited_title":"1200 subjects data release reference manual","cited_arxiv_id":null,"evidence_quote":"documents the scanning protocols and standard preprocessing pipeline for that dataset."},{"cited_title":"Lesne, J.L.Blanc, and L","cited_arxiv_id":null,"evidence_quote":"establishes that Lempel-Ziv-based entropy estimates are reliable on short symbolic sequences, justifying $h_{LZ}$."},{"cited_title":"Estevez-Rams, D","cited_arxiv_id":null,"evidence_quote":"gives the shuffling-based estimate used for the effective measure complexity $E_{LZ}$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the normalized information distance that the LZ-distance $d_{LZ}$ estimates."},{"cited_title":"Estevez-Rams, R","cited_arxiv_id":null,"evidence_quote":"provides the Lempel-Ziv factorization procedure used to estimate Kolmogorov complexity in the distance calculation."},{"cited_title":"Brincat, Markus Siegel, Ravi D","cited_arxiv_id":null,"evidence_quote":"supports the interpretation that inactive regions do not suppress background noise and therefore appear more random."}],"review_version":1}