{"id":"badf3e8a-4e04-4f12-bfae-ec44e6b847f9","arxiv_id":"2608.00655","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"Component-removing blind source separation can collapse causal graph recovery to zero in synthetic calcium traces and dramatically densify connectivity estimates from real v2a-RSN traces, so BSS denoising is not a neutral preprocessing step.","lead":"This paper tests whether blind source separation (BSS) cleaning of calcium-imaging traces preserves the signals used to decode behavior and infer brain connectivity. It finds that removing BSS components can destroy recoverable causal graph structure in simulations and sharply densify connectivity graphs from real zebrafish data, so denoising should be treated as an intervention, not a neutral cleanup.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central graph-collapse result is demonstrated only for cluster-retention component removal; the paper's own artifact-score selection rule is never tested, so the headline may overgeneralize.","rationale":"The abstract's central negative result is the strongest evidence that BSS component removal can distort causal structure learning. If that result is an artifact of the particular way 'component removal' was operationalized—removing whole k-means clusters rather than components flagged as artifacts by the paper's own scoring method—then the paper's main empirical support for the 'intervention' framing is much weaker. Readers would still accept that any non-identity transformation is an intervention, but that is a near-tautology; the paper's contribution is the surprising quantitative claim that even the best cluster-based removal collapses graph F1 to 0. The paper is honest about this limitation in the Discussion and Table S2, but the limitation is not peripheral: it defines the BSS intervention itself. I am not claiming the result is wrong; I am claiming it is under-tested at the exact point where the headline depends on a choice. A single controlled re-run with the artifact/protection scoring rule would settle whether the collapse is a property of component-removing BSS or of cluster-retention selection. This does not undermine the paper's internal consistency or the usefulness of its validation framework; it narrows the scope of the central claim until that check is run. Hence the reader's CONDITIONAL verdict should stand.","tokens_in":18936,"tokens_out":9065,"duration_ms":106997,"concrete_test":"Use the released repository to re-run the §3.1 synthetic benchmark (8 scenarios x 30 seeds; c-GC, c-GC*, PCMCI+, JPCMCI+) with a component-selection rule based on the artifact/protection scoring definition in Supplementary S1.6, instead of cluster retention. Sweep the artifact-score drop threshold and protection cap, fit all operations inside the same fold-local boundary, and compute median graph F1 per BSS method. Compare against Table 1 raw and cluster_keep_top_04 values. If median F1 remains at or near 0 for all thresholds, the concern is resolved. If F1 recovers toward the raw level (~0.5–0.65), the headline collapse is an artifact of the cluster-retention selection rule and the paper's central negative claim must be explicitly scoped to that rule.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing quantitative result is the synthetic graph-recovery collapse: 'component-removing BSS variants collapsed median graph F1 to 0' (Abstract, §3.1, Table 1). Every component-removal variant in the audited grid, including Table 4's cluster_keep_top_04, belongs to one decision family: k-means clustering of ICs into seven spectral clusters, then retaining/removing whole clusters (§2.6, Table S2). The paper's own Supplementary Methods define a different, more directly artifact-targeted rule: components are dropped only when an artifact score exceeds a threshold and a protection score (including behavior correlations) stays below its maximum; otherwise keep/review (S1.6). That rule is never run in the synthetic benchmark or in the v2a-RSN c-GC/c-GC* sensitivity analysis (Table 4). Thus the central claim that component-removing BSS erases recoverable graph structure is established only for cluster-retention removal. An expert who removes components by artifact/protection scores, or by explained variance, may not see F1 = 0. The high-level advice 'treat BSS as an intervention' is definitionally safe, but the quantitative headline and the empirical density shift are scoped to a narrow selection family. The Discussion already concedes 'one family of component-selection rules', but that family is exactly what the headline negative result depends on.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether blind source separation (BSS) applied to already-extracted calcium traces preserves the information needed for behavior decoding and causal structure learning. It introduces a synthetic benchmark with known lagged graphs and corrupted observations, and reports that raw corrupted traces retain recoverable graph information (median F1 ≈ 0.5–0.66), whereas component-removing BSS variants collapse median graph F1 to 0 across c-GC, c-GC*, PCMCI+, and JPCMCI+. On four larval zebrafish v2a-RSN recordings, the paper finds that BSS sometimes improves behavior-decoding scores, but matched PCA and low-pass controls often match or exceed BSS in fold-audited comparisons, and BSS-cleaned traces yield much denser c-GC/c-GC* graphs than raw traces. The authors conclude that BSS should be treated as an intervention on the measured process, not as a neutral cleanup step, and they provide a validation framework combining trace preservation, behavior scorecards, leakage audits, matched controls, causal-state diagnostics, and graph sensitivity.","tokens_in":19325,"tokens_out":5597,"duration_ms":62857,"significance":"If the central claim holds, the paper is a useful cautionary study for the calcium-imaging preprocessing community: component-removing BSS is not automatically safe for downstream causal analysis. The main strengths are the synthetic ground-truth benchmark with known lagged graphs, the use of four BSS methods and several causal estimators, the fold-local leakage audit, matched PCA/low-pass baselines, and the explicit, reproducible code repository. The authors are also transparent about many limitations, including post hoc best-of-grid selection and missing uncertainty intervals. However, the breadth of the headline negative result is constrained by the fact that only one component-selection family is tested, which limits the generality of the quantitative claims even though the qualitative advice is sound.","major_comments":[{"comment":"The central negative result — component-removing BSS collapses median graph F1 to 0 — is demonstrated only for the cluster-retention selection family (k-means spectral clustering into seven clusters, retaining/removing whole clusters; §2.6, Table S2). The artifact/protection scoring rule defined in S1.6, which drops a component only when artifact evidence exceeds a threshold and protection evidence stays below its maximum, is never run in the synthetic benchmark or in the v2a-RSN connectivity analysis (Table 4). Because that rule is explicitly designed to protect behavior-relevant components, the headline overgeneralizes. Either add experiments with the artifact/protection rule and at least one explained-variance-based selection rule, or qualify every occurrence of \"component-removing BSS\" as \"cluster-retention BSS\".","section":"Abstract, §3.1, Table 1"},{"comment":"The empirical connectivity-density result is also tied to a single selection rule, cluster_keep_top_04, applied with all four BSS methods. There is no ground-truth graph, so the only readout is density; using one operationalization of \"component removal\" makes the claim \"BSS cleaning makes v2a-RSN connectivity matrices denser\" narrower than stated. The paper's own artifact/protection rule could plausibly yield different density shifts. Please report sensitivity to at least one additional selection rule (e.g., artifact-score threshold or variance-retained) or explicitly state in the main text and abstract that this is a single-selection-rule sensitivity result.","section":"§3.4, Table 4"}],"minor_comments":[{"comment":"The sentence \"The lag τ=2 was an assumed value corresponding to approximately one-third of the acquisition frequency, and npasts=4 was chosen as it is a sufficient conditioning depth\" appears nearly verbatim twice. State it once and refer back.","section":"§2.10 and §3.4"},{"comment":"The caption uses abbreviations cgc, cgc_star, jpcmciplus, and pcmci without expanding them, and panel (a) \"Best BSS graph F1\" includes all-component and oracle-selection outputs. The main text notes this is a diagnostic envelope, but the caption alone should make that clear.","section":"Figure 2 caption"},{"comment":"Define abbreviations in captions/footnotes: \"Ctrl.\", \"dyn.\", \"BA\". Currently they are only clear from context after reading the Methods.","section":"Table 3 and Figure 5"},{"comment":"The phrase \"The best BSS variant improved several raw behavior readouts\" should explicitly say \"best over the audited grid\" to avoid the implication of a prespecified successful variant; the paper's own post hoc caveat is in the text but could be moved closer to the results.","section":"§3.3"},{"comment":"For readers not familiar with the author's earlier work, a one-sentence definition of c-GC versus c-GC* would be helpful at first use; currently the distinction is only clear from context and reference [52].","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a single-author preprint with a carefully bounded Discussion, but the Abstract and several results go beyond the tested component-selection family. The missing artifact/protection-rule experiment is the key gap: the paper defines a more targeted selection rule in the Supplementary Methods and then does not evaluate it, leaving the headline negative result dependent on one cluster-retention operationalization. I would support publication after the authors either add that experiment or consistently narrow the claims. The code availability, synthetic ground truth, and leakage-audited design are strengths that should be credited in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper shows something worth knowing: component-removing BSS on already extracted calcium traces can destroy recoverable causal graph structure in a synthetic benchmark, and it can massively densify real-data connectivity estimates. The authors also show that BSS behavior decoding gains are often matched by PCA or low-pass controls. That is a genuine, useful caution for the calcium-imaging community.\n\nWhat's new: they built a synthetic ground-truth benchmark with known lagged graphs, injected artifacts, and compared raw, clean, oracle, BSS-all, BSS-removed, PCA, and low-pass. Raw traces retained graph recovery (F1 around 0.5–0.65) while cluster-based component-removing BSS variants collapsed median F1 to 0 across c-GC, c-GC*, PCMCI+ and JPCMCI+. That's a stark quantitative result. They also did a leakage-audited empirical pipeline with fold-local fitting and matched controls, and they are transparent about missing error bars and post-hoc best-of-grid selection. The code is available.\n\nThe main soft spot is scope. The graph-collapse result is demonstrated for one component-selection family: k-means clustering of ICs in spectral feature space and retaining ordered clusters. The paper defines a different, more targeted removal rule (artifact score + protection score) in Supplementary S1.6, but never runs it in the synthetic benchmark or the v2a-RSN sensitivity analysis. So the abstract's blanket statement that 'component-removing BSS variants collapsed median graph F1 to 0' overgeneralizes. If practitioners remove components by artifact scores or explained variance, the distortion might be smaller. The Discussion does concede 'one family of component-selection rules,' but the abstract and title don't carry that caveat.\n\nOther softer spots: only four fish, three in the full audited grid; raw calcium data not clearly public (code yes, data not stated); some headline numbers are post hoc best-of-grid; no confidence intervals for BPI and graph density. The empirical connectivity densification is only shown for c-GC/c-GC* at fixed tau and npasts, but it's large (5% to 50–88%).\n\nThe central argument holds within its scope: BSS is an intervention, not neutral cleanup. The paper is honest about limitations and the framework is reusable.\n\nWho it's for: anyone preprocessing calcium traces before behavior decoding or causal structure learning. It deserves a serious referee and probably publication after the selection-rule gap is acknowledged in the abstract or partially addressed. Recommendation: engage with it, send to peer review, and ask the authors to test at least one non-cluster selection rule or trim the abstract's generalization.","headline":"Useful cautionary study with a genuine synthetic benchmark; the headline overreaches slightly because only one component-selection family is tested.","tokens_in":19750,"tokens_out":2213,"would_cite":true,"duration_ms":22538,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Component-removing blind source separation can destroy the causal signal in calcium traces.","keywords":["blind source separation","calcium imaging","causal structure learning","Granger causality","denoising","behavior decoding","component selection","connectivity"],"falsifier":"Run the same synthetic benchmark with a component-selection rule that removes only components whose time courses correlate with an injected artefact probe above a threshold (an oracle-like artefact removal). If median graph F1 stays near or above the raw-trace level (e.g., ≥0.5 for c-GC), then the collapse to zero is an artifact of the spectral-cluster selection rule, not of component-removing BSS in general.","tokens_in":18811,"feed_emoji":"🧠","tokens_out":4822,"duration_ms":47872,"temperature":0.7,"pith_summary":"The paper argues that denoising calcium traces by blind source separation (BSS)—decomposing the traces into components and discarding some—is not a neutral cleanup step. In a synthetic benchmark with known lagged causal graphs, raw corrupted traces still allowed graph recovery, but every component-removing BSS variant collapsed median graph recovery to zero across four estimator families. On four larval zebrafish recordings of v2a reticulospinal neurons, BSS sometimes improved behavior decoding, but the gains were inconsistent across fish, methods, and retained component clusters, and matched PCA or low-pass controls often matched or beat them. The paper's central conclusion is that BSS should be treated as an intervention on the measured process, and its effect on the specific downstream analysis—especially causal structure learning—should be validated before trusting the cleaned traces.","feed_headline":"Blind source separation can erase causal structure in calcium traces","feed_subtitle":"Cleaning calcium traces by removing components collapses graph recovery to zero and inflates connectivity graphs.","key_machinery":"The load-bearing mechanism is the component-removing reconstruction: BSS decomposes the trace matrix into source time courses and loadings, clusters components by spectral features (Welch power spectra, k-means), retains only a subset of clusters, and reconstructs traces from the retained components alone. This selection step deletes coordinates that may carry behavior-relevant and graph-relevant signal. The evaluation framework—fold-local leakage audits, matched PCA and low-pass controls, behavior scorecards, and causal-state diagnostics—makes the distortion visible, but the deletion itself is the intervention that erases causal evidence.","core_discovery":"The central discovery is that removing components after blind source separation can delete the temporal dependencies that causal structure learning needs, while leaving behavior decoding roughly intact or even improved. In the synthetic benchmark, median graph F1 for raw corrupted traces was 0.585 (c-GC), 0.500 (c-GC*), and 0.658 (PCMCI+/JPCMCI+), while all component-removing BSS variants returned median F1 of 0, with zero true and false positives and 17–18 false negatives. The paper further shows that on v2a-RSN recordings, BSS-cleaned traces produce much denser c-GC and c-GC* connectivity graphs (49–88% off-diagonal directed edges) than raw traces (4.5–6.7%), meaning preprocessing dominate","pith_inferences":["Editorial inference: The same audit logic applies to any preprocessing that transforms the trace matrix—PCA, low-pass filtering, deconvolution, or deep-learning denoisers—so the paper's framework could serve as a general validation protocol for denoising before causal analysis.","Editorial inference: The density inflation after BSS suggests that component removal may suppress variance that Granger causality uses to reject spurious edges, or introduce correlations by reconstruction; this could be tested by comparing BSS-cleaned graphs to graphs from rank-matched random subspaces in the synthetic benchmark.","Editorial inference: A natural extension is to test component-selection rules based on explained variance or artefact-probe correlation rather than spectral clustering; the paper's negative result may be sensitive to that choice, and such rules could preserve more graph-relevant information.","Editorial inference: Because behavior decoding gains from BSS were often matched by smoothing controls, a plausible hypothesis is that BSS improves decoding by smoothing, not by recovering neural signal; this could be tested by comparing BSS to a low-pass filter matched on dynamic state diagnostics."],"forward_implications":["Researchers who clean calcium traces with BSS before Granger causality or related causal analyses should report raw-trace connectivity as a sensitivity check, because cleaned graphs can be far denser.","Component-removing BSS should not be assumed safe for causal structure learning; its effect on graph recovery should be validated on synthetic data with known ground truth before use on real data.","BSS can improve behavior decoding, but because matched PCA and low-pass controls often achieve the same or better scores, such improvements do not by themselves justify BSS as a denoising step.","The separation between trace preservation and graph recovery (raw traces had median trace correlation 0.83 vs clean, BSS-removed 0.60) implies that trace similarity to raw data is a poor proxy for preserving causal evidence."],"fun_headline_variants":["BSS cleanup erases causal links in calcium traces","Cleaning calcium traces with BSS inflates connectivity","Blind source separation can kill calcium graph recovery","BSS cleanup: behavior gains, connectivity losses","Calcium trace cleaning distorts causal structure"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The negative result rests on how components are selected for removal: spectral clustering into seven clusters with retention of top clusters by peak mean power, a rule that may not match how practitioners choose components; if expert or variance-based selection keeps more graph-relevant components, the distortion could be smaller.","fun_headline_variants_meta":{"raw":{"variants":["BSS cleanup erases causal links in calcium traces","Cleaning calcium traces with BSS inflates connectivity","Blind source separation can kill calcium graph recovery","BSS cleanup: behavior gains, connectivity losses","Calcium trace cleaning distorts causal structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1253,"prompt_tokens":788,"completion_tokens":465,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":407}},"tokens_in":532,"tokens_out":465,"duration_ms":5337,"temperature":1.0,"reasoning_tokens":407,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T01:08:59.357608+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same synthetic benchmark with a component-selection rule that removes only components whose time courses correlate with an injected artefact probe above a threshold (an oracle-like artefact removal). If median graph F1 stays near or above the raw-trace level (e.g., ≥0.5 for c-GC), then the collapse to zero is an artifact of the spectral-cluster selection rule, not of component-removing BSS in general.","supporting_citations":[],"review_version":1}