{"id":"c646cbc5-5012-46ed-930a-8bb4c45b3c09","arxiv_id":"2411.18012","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A thresholded correlation Gaussian process prior enables Bayesian estimation and uncertainty quantification for sparse, piecewise smooth, spatially varying correlations between two imaging modalities.","lead":"This paper introduces a Bayesian model for spatially varying correlations between two brain imaging modalities, using a thresholded Gaussian process that forces weak correlations to exactly zero and keeps strong ones smooth. The method comes with consistency theorems and a fast Gibbs sampler, and it is demonstrated on resting-state and task fMRI data from the Human Connectome Project.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3's proof of exact-zero selection consistency is incomplete: it requires interchanging limits in n and k, and L1 posterior consistency alone does not force posterior mass of the exact-zero set to 1.","rationale":"The selected concern is load-bearing because Theorem 3 is the paper's headline selection-consistency result. Without a valid proof, the main advertised guarantee—that the method recovers the true sign pattern, including the exactly-zero region, as n and m grow—is unsupported. The reader's structural-assumption objection and the data-dependent-prior concern are legitimate but secondary: the former is an explicit modeling restriction, and the latter is a theory-implementation mismatch. The double-limit gap is internal to the proof of the central theorem. It may be repairable by proving a rate of L1 convergence or by exploiting the threshold gap (nonzero ρ is bounded away from zero), but that repair is nontrivial and is not present in the manuscript. The method and simulations are promising, and the identified gap does not clearly invalidate the empirical claims, so conditional acceptance remains appropriate. I therefore keep the reader's verdict unchanged.","tokens_in":117,"tokens_out":11880,"duration_ms":301443,"concrete_test":"Independently re-derive the final step of Theorem 3 in Appendix A.4 as a double limit: a_{n,k}=pr(∫_{R0}|ρ|dv < 1/k | Y_n). State precisely whether any uniformity in n is available for the convergence a_{n,k}→1, for example a rate in Theorem 2 or a uniform bound using the threshold gap |ρ|>δ>0 whenever ρ≠0. If no such uniformity exists, construct a posterior sequence that is consistent in every fixed L1 neighborhood but whose exact-zero mass stays bounded away from 1, and check whether such a sequence is compatible with the TCGP likelihood when the true ρ0=0 on R0. If it is compatible, the proof of Theorem 3 is invalid as written; if not, identify the property that rules it out.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Appendix A.4, the proof defines F_k(R0)= {∫_{R0}|ρ|dv < 1/k}, notes pr(F_k(R0)|Y)→1 for each fixed k by Theorem 2, and then writes {ρ=0 on R0}=∩_k F_k(R0), concluding by monotone continuity that pr({ρ=0 on R0}|Y)→1. This tacitly interchanges lim_{n→∞} and lim_{k→∞}. For a fixed data set, monotone continuity gives pr(∩_k F_k|Y)=lim_k pr(F_k|Y), but Theorem 2 only gives lim_n pr(F_k|Y)=1 for each k; the double limit need not commute. A sequence of posterior distributions can put probability 1−1/n on a point 1/n and 1/n on a point 0: every fixed ε-neighborhood of 0 has posterior mass →1, yet exact-zero mass →0. The thresholded prior can behave similarly, with posterior mass shifting to ρ values just above the threshold on tiny sets. Thus the central claim of sign consistency for the zero region is not established by the supplied proof, even under the model's own assumptions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Bayesian nonparametric model for spatially varying correlations between two imaging modalities, built on a thresholded correlation Gaussian process (TCGP) prior. The construction thresholds a latent Gaussian process so that at each spatial location the correlation is either zero or bounded away from zero, yielding piecewise smooth, sparse, and discontinuous correlation surfaces. The authors prove identifiability (Proposition 1), large prior support (Theorem 1), posterior consistency (Theorem 2), and sign/selection consistency (Theorem 3). They derive full conditional distributions and propose a Gibbs sampler and a hybrid mini-batch MCMC for computation. Simulations on 2D and 3D images and an analysis of Human Connectome Project fMRI data compare favorably with voxel-wise, region-wise, and integrated competitors. The main weaknesses are an apparently incomplete proof of exact-zero selection consistency in Theorem 3, a mismatch between the fixed prior analyzed in the theory and the data-adaptive prior used in the implementation, and the absence of an analysis of the Karhunen-Loève truncation error introduced in Section 4.1.","tokens_in":42053,"tokens_out":2561,"duration_ms":26184,"significance":"If the theoretical results hold, the paper is a substantial contribution to Bayesian nonparametric modeling of spatially varying correlations. The thresholded-correlation construction is novel, and the model has the desirable features of sparsity, piecewise smoothness, and jump discontinuities. The paper explicitly provides machine-checkable-style proofs, though they are not machine-checked, and the empirical results are strong: the proposed method substantially outperforms voxel-wise and region-wise baselines and the integrated method of Li et al. (2019) in the reported simulations, and the HCP analysis yields interpretable brain regions with plausible associations. The proposed Gibbs sampler with analytically derived full conditionals is an algorithmic contribution in its own right. However, the central theoretical claim of exact-zero selection consistency currently rests on a questionable double-limit interchange, and the implemented adaptive prior for the threshold parameter falls outside the scope of the theorems. If the proof gap is closed and the theory is aligned with the implementation, the paper would merit publication in a leading statistics journal.","major_comments":[{"comment":"The proof of exact-zero selection consistency is incomplete because it interchanges limits in n and k without justification. The argument defines F_k(R0) = {∫_{R0}|ρ|dv < 1/k}, obtains pr(F_k(R0)|Y) → 1 for each fixed k from Theorem 2, and then uses monotone continuity to conclude pr({ρ=0 on R0}|Y) = lim_k pr(F_k(R0)|Y) = 1. For a fixed data set, monotone continuity indeed gives pr(∩_k F_k|Y) = lim_k pr(F_k|Y), but Theorem 2 only yields lim_n pr(F_k|Y) = 1 for each k. The double limit need not commute; a sequence of posteriors can concentrate on points just above zero on R0 at a rate that outpaces the convergence in n. Thus the proof as written does not establish that the posterior mass on the exact-zero set tends to one, which is the central claim of Theorem 3.","section":"Appendix A.4, proof of Theorem 3"},{"comment":"The theoretical results assume a fixed prior specification for the threshold ω, but the implemented and recommended procedure uses an adaptive, data-dependent prior: Section 4.2 states that aω and bω are chosen as quantiles of {|ξ(v)|} and are 'adaptively changed based on the value of {|ξ(v)|}v∈B in each iteration,' and Section 6 sets aω to the 75% quantile and bω to the 100% quantile of the posterior draws. This empirical-Bayes-like prior is not covered by the proof of Theorem 2, which constructs sieves and tests under a fixed TCGP(ω0, κ, τ²) prior. Since ω controls which voxels are treated as exactly zero, the proof gap directly affects the sparsity and selection claims. The authors should either extend the theory to data-dependent priors, or modify the implementation to use a fixed prior whose range is justified independently of the data.","section":"Sections 4.2 and 6; Theorems 2 and 3"},{"comment":"The posterior consistency theorems are stated for the infinite-dimensional model (2.9)–(2.10), but all simulations and the HCP analysis use the finite-dimensional KL-truncated model (4.11) with a fixed truncation level L (L = 540 in the HCP analysis). No bound is given for the truncation error, and no rate condition linking L to n and m is stated. Consequently, the theoretical results do not apply to the model actually implemented, even if the proof in Appendix A.4 were correct. The authors should either extend the theory to the truncated model or demonstrate that the truncation error is asymptotically negligible under Assumption 3.","section":"Section 4.1, Eq. (4.11); Theorems 2 and 3"}],"minor_comments":[{"comment":"There is a typo in the last line of the proof of Lemma 5: 'ccompletes' should read 'completes'.","section":"Appendix A.4, Lemma 5"},{"comment":"In the proof of Lemma 7, 'at lease' should read 'at least'.","section":"Appendix A.4, Lemma 7"},{"comment":"The caption of Table 6 contains the misspelling 'False Discorvery Rate'; it should be 'False Discovery Rate'.","section":"Table 6"},{"comment":"The figure numbering appears inconsistent with the text: Section 2.1 refers to 'Figure 1' for the graphical model illustration, while Section 5 refers to 'Figure 1' for a simulation slice; later Section 6 refers to 'Figure 3' for the HCP activation maps. Please renumber the figures and correct all cross-references.","section":"Figures 1–3"},{"comment":"The 3D simulation generates data using posterior means obtained from the same HCP data set analyzed in Section 6; this is a model-based truth and constitutes self-validation. It would be useful to state this limitation explicitly in the simulation section and to include at least one simulation with a truth not derived from the model's own posterior, for example with smoothly varying but nonzero correlations, to assess robustness.","section":"Section 5, first paragraph"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for Statistica Sinica and addresses an important problem. The empirical results are strong, and the algorithmic contribution is solid. My main concern is the proof of Theorem 3: the double-limit issue in Appendix A.4 is not merely cosmetic and needs a corrected argument. In addition, the implemented prior for ω is data-dependent, so the theory and practice are not aligned. The KL truncation issue is also relevant but might be addressed by a supplementary result or by reframing the theorems for the truncated model. The paper is not ready for acceptance in its current form, but I believe the issues are addressable within the scope of a major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Off the record: this paper is worth engaging, but read the proofs before citing the headline results. The thresholded correlation Gaussian process is a genuinely new prior for sparse, piecewise-smooth correlation fields, and the equivalent data transformation that turns the correlation function into a mean parameter is clever. Identifiability, large support, and L1 posterior consistency are plausibly established, and the Gibbs sampler (mixtures of truncated normals and uniforms) is a real computational contribution. The HCP application is a nice bonus.\n\nThe main soft spot is Theorem 3, the selection consistency claim. The proof in Appendix A.4 defines events F_k(R0) = {∫_{R0}|ρ|dv < 1/k}, shows each has posterior mass going to 1 as n→∞ for fixed k, then uses monotone continuity to conclude the intersection {ρ=0 on R0} has posterior mass going to 1. That requires interchanging the limits in n and k. This does not commute: a posterior can put mass 1−1/n on functions with small positive blobs on R0 and mass 1/n on exactly zero; every fixed L1 neighborhood gets mass →1 but exact-zero mass goes to 0. The thresholded prior has enough discontinuity that this kind of behavior is not ruled out. So the paper's central claim of sign recovery for the zero region is not established by the supplied argument.\n\nOther issues are real but smaller. The data-dependent prior for ω (the quantile-based range in Section 4.2 and the empirical Bayes refinement) sits outside the fixed-prior theory, and the KL truncation error is never analyzed. The 3D simulation generates data from the fitted model itself, so it checks internal consistency rather than the model's assumptions. No code or data is provided, which slows independent verification. The structural assumption that the truth is exactly zero on a set with nonempty interior and bounded away from zero elsewhere is explicit and may be too strong for some applications, but that is a modeling choice, not a flaw in derivation.\n\nWho this is for: Bayesian methodologists in neuroimaging and researchers working on thresholded GP priors. It should go to peer review, but with the expectation of heavy revision. The L1 consistency story is plausible; the exact-zero selection theorem needs either a repaired proof with a union bound over the boundary or a weakened statement (e.g., posterior mass on small L1 neighborhoods of the zero set).","headline":"The TCGP prior is a genuine new construction for spatially varying correlations with real computational machinery, but Theorem 3's exact-zero selection consistency is not proven: the proof assumes a limit interchange that need not commute.","tokens_in":42619,"tokens_out":2611,"would_cite":true,"duration_ms":27183,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62G20","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a thresholded correlation Gaussian process prior that makes spatially varying correlations between two imaging modalities identifiable and consistently recoverable, and demonstrates it on brain imaging data.","keywords":["Bayesian nonparametrics","thresholded Gaussian process","spatially varying correlation","multimodal neuroimaging","posterior consistency","selection consistency","Gibbs sampling","functional MRI"],"falsifier":"Run the TCGP posterior on data generated from a true $\\rho_0$ that is everywhere nonzero with $0<|\\rho_0(v)|<\\gamma$, so the null set is empty and no jump exists; the Theorem 3 premise is violated, and the estimated sign map and posterior inclusion probabilities will either declare spurious zero regions or fail to recover the true sign pattern.","tokens_in":41585,"feed_emoji":"🧠","tokens_out":6924,"duration_ms":57689,"temperature":0.7,"pith_summary":"This paper proposes a Bayesian nonparametric model for estimating how the correlation between two imaging modalities varies across space, and for locating the brain regions where that correlation is nonzero. The prior, called the thresholded correlation Gaussian process (TCGP), is built by passing a latent Gaussian process through a hard threshold, so the resulting correlation map is smooth within regions, sparse (exactly zero where the latent process stays below a threshold), and can jump at region boundaries. The paper establishes that the model is identifiable, that the prior has large support, and that the posterior is consistent, with posterior probability tending to one on the correct sign of the correlation at every location. If correct, this gives researchers a principled, uncertainty-quantified way to identify activation regions from multimodal imaging data, with performance that is maintained when the number of subjects is small or the signal is weak.","feed_headline":"Thresholded Gaussian process maps brain regions where scans correlate","feed_subtitle":"Spare-data-friendly Bayesian model identifies positive and negative correlation regions in the brain.","key_machinery":"The load-bearing object is the thresholded correlation Gaussian process prior (Definition 1, equations (2.2)–(2.6)). A latent Gaussian process $\\xi(v)$ is thresholded at level $\\omega$ to define $\\sigma_+(v) = G_\\omega\\{\\xi(v)\\}$ and $\\sigma_-(v) = G_\\omega\\{-\\xi(v)\\}$, which in turn scale two independent Gaussian processes whose sum and difference form the two modality means; after integrating those out, the correlation is the closed-form function in (2.6). The threshold enforces sparsity and identifiability, since $\\sigma_+(v)\\sigma_-(v)=0$ everywhere, and the Gaussian process kernel imposes spatial smoothness on the nonzero part. The equivalent model (2.9) uses $Y_{\\pm,i}(v) = s\\{\\pm\\rho(v)\\}E_{\\pm,i}(v)+\\varepsilon_{\\pm,i}(v)$ for a monotone function $s$, which turns the correlation into a coefficient multiplying a Gaussian process; this representation supports the posterior-consistency proof and yields full conditional distributions that are mixtures of truncated normals or uniforms, leading to a Gibbs sampler without gradient approximations.","core_discovery":"The central claim is that the correlation function $\\rho(v)$ between two aligned image modalities can be represented by equation (2.6), driven by a single latent Gaussian process $\\xi(v)$ through the thresholding function $G_\\omega(x) = x I(x > \\omega)$. Because $\\sigma_+(v) = G_\\omega\\{\\xi(v)\\}$ and $\\sigma_-(v) = G_\\omega\\{-\\xi(v)\\}$, at most one of the positive and negative variance components is active at any voxel; this forces exact zeros where $|\\xi(v)| \\le \\omega$ and delivers piecewise smooth, jump-discontinuous correlation maps. The paper proves that this TCGP formulation is identifiable (Proposition 1), that the prior assigns positive probability to every small supremum-norm neighborhood of any true correlation function in the class (Theorem 1), and that the posterior concentrates and recovers the sign pattern $\\operatorname{sgn}\\{\\rho(v)\\}$ for all $v$ as the number of subjects and voxels grow (Theorems 2 and 3). A reparametrization in terms of average and contrast images $Y_+(v), Y_-(v)$ turns $\\rho(v)$ into a mean parameter, which is what makes both the theory and the Gibbs sampler tractable.","pith_inferences":["The exact-zero assumption is a modeling commitment; in real data where all correlations are small but nonzero, the method should be read as detecting correlations above the threshold, and the sign-consistency theorem should not be invoked.","The adaptive choice of the prior range for $\\omega$ from quantiles of $|\\xi(v)|$ (Sections 4.2 and 6) lies outside the fixed-prior theory; a natural diagnostic is to repeat the analysis across quantile levels and check that declared regions are stable.","The same construction could be transplanted to temporally varying or group-level correlations, replacing the spatial domain with time or subject index; the proof structure would need new assumptions on the domain and smoothness, but the thresholding mechanism is domain-agnostic.","Because the transformed model (2.9) is a mean regression with a Gaussian process, the framework could in principle be extended to non-Gaussian imaging data through link functions, although the current theory and Gibbs sampler are tailored to Gaussian likelihoods."],"forward_implications":["Neuroscientists can report whole-brain maps of positive, negative, and null correlations between two modalities, with posterior inclusion probabilities as uncertainty measures, rather than a single test statistic per voxel.","The method is usable in small-sample imaging studies: simulations with $n=500$ and weak signals still show sensitivity around 0.89–0.90 and low false discovery rates, while voxel-wise tests have near-zero power.","The hybrid mini-batch MCMC reduces per-iteration cost from $O(m^2)$ to $O(m_s^2)$, so the procedure scales to the roughly 117,000 voxels of a typical whole-brain analysis.","The prior has large support over a wide class of sparse, piecewise smooth, jump-discontinuous correlation functions, so the model is not restricted to a parametric family."],"supporting_citations":[{"why":"Supplies the posterior consistency framework for Gaussian process priors in nonparametric regression that the paper extends to the hierarchical TCGP model.","marker":"Choi (2005)"},{"why":"Provides the large-support and tail-probability tools, including Theorem 5, used to prove Theorem 1 and Lemma 4 for thresholded Gaussian process priors.","marker":"Ghosal and Roy (2006)"},{"why":"Gives posterior-consistency theory for logistic Gaussian process priors that the paper adapts to the thresholded correlation setting.","marker":"Tokdar and Ghosh (2007)"},{"why":"Introduces the thresholded Gaussian process prior for scalar-on-image regression that TCGP builds on.","marker":"Kang et al. (2018)"},{"why":"Shares the thresholded Gaussian process machinery and the Karhunen–Loève truncation practice used in the posterior computation section.","marker":"Wu et al. (2024)"},{"why":"The main comparator in simulations and the existing spatially adaptive varying correlation method for multimodal neuroimaging.","marker":"Li et al. (2019)"},{"why":"Provides Theorem A1, the posterior-consistency criterion that the paper verifies through the lemmas in the supplement.","marker":"Choudhuri et al. (2004)"},{"why":"Introduces the spatially varying coefficient model with jump discontinuities that motivates the function class in Definition 3.","marker":"Zhu et al. (2014)"}],"fun_headline_variants":["Thresholded GP maps where brain scans significantly correlate","Bayesian thresholding reveals sparse brain correlation regions","New Bayesian model maps significant brain scan correlations","Efficient Gibbs-sampled thresholded GP reveals brain correlation areas"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the true correlation is exactly zero throughout a whole region and bounded away from zero everywhere else; if real associations are merely small but never exactly zero, the thresholded prior is misspecified and the sign-consistency guarantee does not apply.","fun_headline_variants_meta":{"raw":{"variants":["Thresholded GP maps where brain scans significantly correlate","Bayesian thresholding reveals sparse brain correlation regions","New Bayesian model maps significant brain scan correlations","Efficient Gibbs-sampled thresholded GP reveals brain correlation areas"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00097,"raw_usage":{"total_tokens":4125,"prompt_tokens":948,"completion_tokens":3177,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":3115}},"tokens_in":564,"tokens_out":3177,"duration_ms":21772,"temperature":1.0,"reasoning_tokens":3115,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:36:25.899466+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the TCGP posterior on data generated from a true $\\rho_0$ that is everywhere nonzero with $0<|\\rho_0(v)|<\\gamma$, so the null set is empty and no jump exists; the Theorem 3 premise is violated, and the estimated sign map and posterior inclusion probabilities will either declare spurious zero regions or fail to recover the true sign pattern.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the posterior consistency framework for Gaussian process priors in nonparametric regression that the paper extends to the hierarchical TCGP model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the large-support and tail-probability tools, including Theorem 5, used to prove Theorem 1 and Lemma 4 for thresholded Gaussian process priors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives posterior-consistency theory for logistic Gaussian process priors that the paper adapts to the thresholded correlation setting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the thresholded Gaussian process prior for scalar-on-image regression that TCGP builds on."},{"cited_title":"Ghosal, and A","cited_arxiv_id":null,"evidence_quote":"Provides Theorem A1, the posterior-consistency criterion that the paper verifies through the lemmas in the supplement."}],"review_version":1}