{"id":"1755da63-33fd-4079-a936-f6fc9200ac18","arxiv_id":"1908.07176","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"GraphMM, a graph-based empirical Bayes mixture model, controls false discovery rates and improves power for detecting spatially coherent effects in large-scale brain-imaging tests.","lead":"GraphMM is a new empirical Bayes method that uses the known graph structure of brain imaging data to detect weak, spatially coherent changes in gray matter. It reports more significant differences between Alzheimer's patients and controls than standard voxel-wise testing, while controlling false discovery rate in simulations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The deployed l_v are local-subgraph posteriors, not draws from one joint model; the FDR-control argument in Section 2.2.1 therefore does not formally transfer to GraphMM as implemented, and calibration is supported only by simulation.","rationale":"The reader's conditional verdict is sound. The model is internally coherent, and the local-FDR formula is standard when computed under a single global posterior. The paper itself discloses the local-subgraph approximation, but gives no calibration theory for it; the simulations cover several favorable and mildly misspecified regimes, yet they cannot close the gap for arbitrary graphs. I chose this as the load-bearing concern rather than the block-constant assumption because the headline claim is about FDR control, and the local approximation directly affects whether l_v are valid posterior probabilities. The block-constant assumption mainly influences power, and the paper's robustness simulation suggests FDR remains controlled under that misspecification. An exact-versus-local comparison on a small lattice would settle whether the concern actually lands. Absent that, the appropriate verdict remains CONDITIONAL: the method is promising and reproducible, but the core FDR-control claim needs a sharper validation before being taken at face value.","tokens_in":13167,"tokens_out":7677,"duration_ms":81684,"concrete_test":"On a 3x3 lattice (N=9, only 1434 graph-respecting partitions), simulate 1000 datasets under null and block-constant alternatives; compute exact global l_v from Eq. 2.2 by enumerating all partitions, then run the GraphMM package's local-subgraph version with equivalent hyperparameters. Record empirical false-discovery proportion at thresholds 0.05, 0.10, and 0.20. If the local-approximation FDP exceeds the exact global FDP by more than about 0.02, the FDR-control claim for the scalable version is not reliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The FDR-control property cited in Section 2.2.1 (Efron 2007; Newton et al. 2004) is that, under a joint model, the sum of l_v over the discovery list bounds the conditional expected number of false positives. The implementation replaces the global posterior in Eq. 2.2 with per-vertex local-subgraph posteriors: 'for each vertex v in the original graph we consider a small local subgraph in which v is one of the central vertices, and we simply deploy GraphMM on this local graph.' Each l_v therefore conditions on a different data subset and a different induced partition prior; no single probability model generates the collection of l_v values. The sum of these local probabilities is not a conditional expectation of false positives under any one model, so the theoretical basis for thresholding l_v is absent for the deployed algorithm. The permutation null and the five simulation scenarios give empirical support, but they do not establish control for the ADNI graph, especially for true blocks that extend beyond the local window or cross its boundary. This is the main load-bearing gap in the headline FDR-control claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes GraphMM, an empirical Bayes mixture model for large-scale hypothesis testing on graph-associated data. The model constrains the expected values of two groups to be constant on the blocks of a graph-respecting partition, and scores each vertex by a local false-discovery rate computed from a mixture over partitions and block-specific null indicators. Because exact posterior computation over all partitions is intractable, the method deploys GraphMM on a small local subgraph around each vertex, with hyperparameters estimated globally from the data. The paper evaluates GraphMM in simulations that mimic the ADNI-2 brain MRI data, comparing FDR and power with standard methods (BH, locfdr, q-value, ASH), and reports substantial power gains when non-null cases form connected subgraphs. It also applies the method to detect gray-matter differences between cognitively normal controls and late MCI subjects in the ADNI-2 cohort.","tokens_in":13330,"tokens_out":5752,"duration_ms":56278,"significance":"If the claims hold, the paper offers a practical advance for graph-associated multiple testing in neuroimaging, with an open-source implementation and a thoughtful demonstration on a real dataset. The paper is genuinely careful in its simulation design: it varies block sizes, effect distributions, and violations of the model, and it includes permutation experiments. These constitute substantial evidence that the method has good operating characteristics in the studied regimes. The empirical power gains, if reproducible, would be of practical value for detecting subtle structural brain changes. Credit is also due for making the R package openly available, which aids reproducibility.","major_comments":[{"comment":"The FDR-control guarantee cited in Section 2.2.1 (Efron 2007; Newton et al. 2004) requires that the l_v values are posterior probabilities from a single joint model. However, the implementation computes each l_v from a different local subgraph: 'for each vertex v in the original graph we consider a small local subgraph in which v is one of the central vertices, and we simply deploy GraphMM on this local graph.' Consequently, no single probability model generates the collection of l_v, and the sum of these local probabilities is not a conditional expected number of false positives under any one model. This leaves the theoretical basis for thresholding l_v absent for the deployed algorithm. The permutation and simulation results provide empirical support, but they do not establish FDR control for the full ADNI application. Please either provide a rigorous analysis of the local-subgraph approximation (e.g., bound the discrepancy between local and full-graph posteriors under model (2.1)–(2.3)) or explicitly state that FDR control is an empirical property observed in simulations rather than a theoretical guarantee.","section":"Section 2.2.1, Eq. (2.2)"},{"comment":"The predictive density is computed via a Laplace approximation to integrate out means, and hyperparameters are estimated by empirical Bayes from the full graph. The paper does not provide error bounds or diagnostics for these approximations, and it does not assess how the plug-in estimation of hyperparameters affects the operating characteristics. Given that the local subgraph posteriors condition on these point estimates, a sensitivity analysis over hyperparameter values or a comparison against MCMC on a small graph would help confirm that the reported FDR control and power gains are not artifacts of the approximation.","section":"Section 2.2.2 and Supplementary Material Eq. (0.2)"},{"comment":"The local subgraph size is a free parameter that is central to the method's computational feasibility and statistical behavior, yet the paper never states what local subgraph sizes were used in the simulations, the permutation experiments, or the ADNI analysis. The sensitivity advantage in Scenario 2 is partially attributed to matching block-size priors, but the role of the local window size is not discussed. A sensitivity analysis varying the local subgraph size would clarify how this tuning choice affects FDR control and power, and would be necessary for a reader to reproduce or apply the method.","section":"Sections 2.2.1 and 3.1"}],"minor_comments":[{"comment":"The word 'probablity' appears twice; it should be 'probability'.","section":"Page 3, first paragraph"},{"comment":"'conjugage' should be 'conjugate'.","section":"Section 2.2.2, paragraph 1"},{"comment":"The citation 'Weiner and Zilles (6 03)' appears to be a typo; the intended year is likely 2016.","section":"Table 1, row 6"},{"comment":"The caption of Figure 3 and the text refer to 'empirical FDR' and 'controlled FDR' but the distinction between these two quantities could be stated more explicitly in the main text; the definitions in Section 2.3 are given in prose and would benefit from display equations.","section":"Section 3.1, first paragraph"},{"comment":"The term 'local subgraph' is used in several places but the construction (e.g., how many neighbors are included) is only described vaguely; a precise algorithmic description in the main text or a pointer to a supplementary algorithm box would improve reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for Biostatistics and presents a plausible and potentially useful methodology. The central concern is the mismatch between the theoretical FDR statement in Section 2.2.1 and the local-subgraph implementation; this is fixable by qualifying the claim or adding an approximation analysis. The authors should also report the local subgraph sizes used in all experiments and provide sensitivity results. The paper would be strengthened by explicit discussion of the empirical-Bayes estimation step and its impact on operating characteristics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe short version: GraphMM is a genuinely new empirical-Bayes method for large-scale testing on graphs that borrows strength across connected vertices via graph-respecting partitions. The simulations and real-data analysis support its main claims, but the headline FDR-control property is not backed by a theorem for the implemented algorithm.\n\nWhat's new: the graph-respecting partition prior, exact enumeration on local subgraphs, and a likelihood that avoids the product-partition simplification. This is distinct from FDR smoothing and spatial product partition models. The paper is careful about that. They also ship an R package, run a useful battery of simulations including robustness to misspecification, and include permutation experiments. The ADNI analysis shows substantially larger yield than standard methods, with sensible anatomical findings.\n\nThe soft spot is the one your stress-test flags, and I think it lands. Section 2.2.1 gives the usual joint-model FDR bound for the global l_v. The implementation, though, computes each l_v on a small local subgraph centered at v, so the collection of l_v values is not generated by a single probability model. The sum of these local probabilities is not a conditional expectation of false discoveries under any one model. The paper does not prove a bound for the deployed algorithm. That said, the simulations and permutation experiments give credible empirical evidence of FDR control in the settings tested. It is a real gap, but not a fatal one: the paper's own claims lean on those simulations, and they are fairly thorough. The authors should either develop a theory for the local algorithm or clearly present it as empirically calibrated.\n\nOther soft spots are proportionally minor. There is no comparison against graph-aware competitors like Tansey et al.'s FDR smoothing, which is a direct touchstone and conspicuously absent. The Laplace approximation and local-window size are not given any formal error analysis, and simulation curves lack error bars. I would not sink the paper on these, but they weaken the \"striking operating characteristics\" claim.\n\nThe paper is for statisticians and neuroimagers who work on graph-based multiple testing. It deserves a serious referee: the idea is original, the data work is real, and the gap I mention is addressable. I would recommend sending it to review with the expectation that the authors add a graph-aware benchmark and address the local-versus-global FDR issue.\n\nBest,\n[name]","headline":"GraphMM is a genuinely new graph-based testing method with strong empirical results, but the theoretical FDR guarantee in the paper does not formally cover the deployed local-subgraph implementation.","tokens_in":13913,"tokens_out":2521,"would_cite":true,"duration_ms":24045,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that GraphMM, an empirical Bayes mixture method using graph-respecting partitions, controls the local false-discovery rate and increases power over graph-ignoring tests whenever true effects form connected subgraphs.","keywords":["empirical Bayes","graph-respecting partition","GraphMM","image analysis","local false-discovery rate","mixture model","brain imaging","false discovery rate control"],"falsifier":"Simulate a two-group Gaussian dataset on a lattice where the non-null vertices are chosen independently with small probability so that no connected block structure exists, run GraphMM and a standard voxel-wise method at the same target FDR, and compare true positive rates; GraphMM's claimed power advantage would fail in this scenario, and any observed FDR inflation would contradict the paper's control claim.","tokens_in":12839,"feed_emoji":"🧠","tokens_out":9469,"duration_ms":87881,"temperature":0.7,"pith_summary":"The paper introduces GraphMM, an empirical Bayes mixture model that scores local false-discovery rates for large-scale testing when the measured variables sit on a known graph. Its central claim is that by constraining the group means to be constant within connected blocks of the graph, the method reduces the effective parameter dimension and thereby gains statistical power relative to procedures that test each vertex on its own. The paper demonstrates this on simulated data with block-structured signals and on structural MRI data comparing normal controls with people in late-stage mild cognitive impairment, where GraphMM reports many more significant voxels than conventional methods at the same nominal FDR. The broader message is that graph geometry can be a useful source of evidence, not just a correction for multiple testing.","feed_headline":"Graph-smart tests find more brain changes at equal FDR.","feed_subtitle":"A new empirical Bayes method uses graph structure to boost power for subtle, spatially clustered brain changes.","key_machinery":"The load-bearing object is the graph-respecting partition, a partition of the vertex set in which every block induces a connected subgraph of the known graph. It encodes the dimension constraint that all vertices in a block share a common mean in each group, so a change between groups is a block-level shift; the set of such partitions is much smaller than all partitions, which makes exact enumeration feasible after localizing to small subgraphs. GraphMM computes each vertex's local false-discovery rate as a sum over these discrete states, using a joint predictive density that allows general covariance (inverse Wishart) and a Laplace approximation for integrating out the block means. The method thereby regularizes contrast estimates between neighboring units without forcing a product-partition factorization, and it estimates all hyperparameters empirically on the whole graph.","core_discovery":"GraphMM treats the unknown arrangement of group differences as a latent graph-respecting partition: the vertex set is divided into connected blocks, every vertex within a block shares a common mean in each of the two groups, and a binary indicator per block records whether the means differ between groups. The local false-discovery rate for a vertex is then the posterior probability that its block's change indicator is zero, summed over all partitions consistent with local subgraphs. The method uses the marginal predictive density after integrating out covariance matrices and numerically integrating the block means, computing exact sums over all graph-respecting partitions of small local neighborhoods while estimating hyperparameters globally. In the simulation scenarios where true effects occupy connected blocks, GraphMM attains the target FDR and shows higher sensitivity than voxel-wise empirical Bayes procedures, adaptive shrinkage, and Benjamini-Hochberg; robustness checks show FDR control remains when the true signal is not block-structured. On brain MRI data, GraphMM detects additional gray-matter changes in regions consistent with the aging and Alzheimer's literature.","pith_inferences":["If the block-constant assumption is replaced by a soft penalty on between-vertex mean contrasts, the graph-respecting partition idea could be folded into a regularized regression or graph total-variation prior, yielding a general recipe for graph-structured multiple testing beyond Gaussian mixtures.","The local-neighborhood approximation suggests a natural scaling test: on graphs with low connectivity, enumeration is cheap and power gains should be largest, whereas on dense graphs the method should approach the behavior of graph-ignoring procedures because the partition space approaches all partitions.","A direct comparison with bandlimited graph-signal smoothing approaches might reveal that GraphMM's advantage is a prior on blockiness rather than on smoothness, which would help practitioners choose between methods based on expected signal geometry.","One could pre-screen data with a graph-aware homogeneity statistic to decide when the block assumption is credible; if blocks are absent, the user could fall back to univariate empirical Bayes without losing FDR control."],"forward_implications":["For 2D or 3D imaging data in which pathological changes affect coherent anatomical regions, GraphMM should generate longer discovery lists at a given FDR than voxel-by-voxel false-discovery-rate methods, as shown in the MRI analysis.","The method's graph-local computation keeps per-vertex cost manageable, so the same machinery can be applied to large brain volumes by processing slices or small neighborhoods.","FDR control appears robust to violations of the block assumption: the paper's simulations show target FDR is maintained even when partitions are not graph-respecting and effects are not uniform within blocks.","Applied to the mild cognitive impairment data, GraphMM flags regions such as precentral gyrus and middle frontal gyrus that conventional methods do not, suggesting additional candidate targets for Alzheimer's research."],"supporting_citations":[{"why":"Establishes the mixture-model local false-discovery-rate formalism and the conditional expected FDR bound that GraphMM's thresholding uses.","marker":"Newton and others (2004)"},{"why":"Supplies the size-power-lfdr theory that justifies constructing discovery lists by thresholding local false-discovery rates.","marker":"Efron (2007)"},{"why":"Defines the false-discovery-rate criterion and provides the baseline procedure GraphMM must outperform.","marker":"Benjamini and Hochberg (1995)"},{"why":"Provides the q-value approach used as a comparison method for empirical Bayes FDR control.","marker":"Storey (2003)"},{"why":"Introduces adaptive shrinkage, a modern voxel-wise empirical Bayes benchmark that GraphMM is compared against.","marker":"Stephens (2017)"},{"why":"Provides product partition models for change-point problems, the special case of graph-respecting partitions on a line graph that GraphMM generalizes.","marker":"Barry and Hartigan (1992)"},{"why":"Demonstrates multiple testing by clustering treatment effects, motivating the partition-based borrowing of strength in GraphMM.","marker":"Dahl and Newton (2007)"},{"why":"Supplies the Alzheimer's disease neuroimaging study data used in the empirical evaluation.","marker":"Weiner and Veitch (2015)"},{"why":"Describes the MRI preprocessing pipeline that produces the gray-matter tissue maps used in the applied analysis.","marker":"Ithapu and others (2015)"}],"fun_headline_variants":["GraphMM finds more Alzheimer's brain changes at fixed FDR","Graph-aware empirical Bayes ups power for connected brain signals","GraphMM leverages spatial structure to find more brain changes","GraphMM improves detection of clustered effects in brain imaging"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The power gain relies on the assumption that true group differences are blockwise constant on connected graph pieces; if the signal is scattered across vertices or fragmented across disconnected components, the dimension reduction no longer helps, even though the paper's robustness simulations indicate FDR control is not lost.","fun_headline_variants_meta":{"raw":{"variants":["GraphMM finds more Alzheimer's brain changes at fixed FDR","Graph-aware empirical Bayes ups power for connected brain signals","GraphMM leverages spatial structure to find more brain changes","GraphMM improves detection of clustered effects in brain imaging"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000565,"raw_usage":{"total_tokens":2621,"prompt_tokens":830,"completion_tokens":1791,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":446,"completion_tokens_details":{"reasoning_tokens":1726}},"tokens_in":446,"tokens_out":1791,"duration_ms":16101,"temperature":1.0,"reasoning_tokens":1726,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:23:40.281058+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a two-group Gaussian dataset on a lattice where the non-null vertices are chosen independently with small probability so that no connected block structure exists, run GraphMM and a standard voxel-wise method at the same target FDR, and compare true positive rates; GraphMM's claimed power advantage would fail in this scenario, and any observed FDR inflation would contradict the paper's control claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the mixture-model local false-discovery-rate formalism and the conditional expected FDR bound that GraphMM's thresholding uses."},{"cited_title":"(2007, 08)","cited_arxiv_id":null,"evidence_quote":"Supplies the size-power-lfdr theory that justifies constructing discovery lists by thresholding local false-discovery rates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the false-discovery-rate criterion and provides the baseline procedure GraphMM must outperform."},{"cited_title":"(2003, 12)","cited_arxiv_id":null,"evidence_quote":"Provides the q-value approach used as a comparison method for empirical Bayes FDR control."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces adaptive shrinkage, a modern voxel-wise empirical Bayes benchmark that GraphMM is compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides product partition models for change-point problems, the special case of graph-respecting partitions on a line graph that GraphMM generalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates multiple testing by clustering treatment effects, motivating the partition-based borrowing of strength in GraphMM."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Alzheimer's disease neuroimaging study data used in the empirical evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the MRI preprocessing pipeline that produces the gray-matter tissue maps used in the applied analysis."}],"review_version":1}