REVIEW 4 major objections 3 minor 40 references
Hypothesis Testing for Network Data with Power Enhancement
T0 review · 4 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Comparing two network populations reduces to a maximum of standardized link differences, and the resulting global test is asymptotically minimax optimal; a sum-based auxiliary statistic then enhances simultaneous-testing power while…
desk verdict A genuinely useful two-sample network inference paper with a clever power-enhancement step, but its motivating brain-connectivity application runs outside the theory's weak-dependence assumptions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing statistic is $T_{i,j} = (\bar{S}_{1,i,j} - \bar{S}_{2,i,j}) / \sqrt{V_{1,i,j}/n_1 + V_{2,i,j}/n_2}$, the standardized link-wise mean difference, with global statistic $M_n = \max_{i<j} T_{i,j}^2$ whose null limit is the type-I extreme-value distribution. The power-enhancement machinery is the auxiliary statistic $A_{i,j} = (\bar{S}_{1,i,j} + \hat{\kappa}_{i,j} \bar{S}_{2,i,j}) / \sqrt{V_{1,i,j}/n_1 + \hat{\kappa}_{i,j}^2 V_{2,i,j}/n_2}$, constructed from the sum rather than the difference, which is asymptotically independent of $T_{i,j}$ under the null; this independence lets the method group tests by $A_{i,j}$, estimate alternative proportions per group, reweight p-values, and apply an FDR-controlling threshold.
What would settle it
Simulate two groups of networks where one hub node connects to all other nodes, so many link-specific correlations approach 1, and check whether the null distribution of M_n still matches the claimed extreme-value law and whether the simultaneous procedure keeps FDR at the nominal level. A simpler check on real data would be to compute the maximum absolute correlation among the q link-specific variables in the brain dataset and compare it with the bound required by the paper's condition (A2).
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that equality of two network mean matrices can be tested by standardizing each link's mean difference, taking the largest standardized square, and comparing it to the extreme-value distribution; this test detects sparse alternatives with standardized differences of order (log q)^{1/2} and is minimax optimal under mild sparsity. For simultaneous inference, rejecting links whose standardized differences exceed a data-driven threshold controls FDR asymptotically. The paper's power-enhancement claim is that an auxiliary statistic built from the sum of the group means, which is asymptotically independent of the test statistic under the null, can be used to group hypotheses, reweight p-values, and yield asymptotic FDR control with power at least that of the unenhanced test.
Load-bearing premise
The main load-bearing premise is that no pair of link-specific variables is too strongly correlated: all correlations among the q entries of each network are bounded away from 1. If many links share a common brain region or hub, this can fail, and then the claimed extreme-value null distribution and FDR control no longer follow from the paper's conditions.
Editorial extensions
If this is right
- If the global minimax claim holds, practitioners can know that no other test can uniformly beat the maximum statistic for detecting sparse mean differences at the (log q)^{1/2} scale.
- If FDR control holds, the procedure gives valid simultaneous inference for individual network links without normality and with only moment conditions, broadening applicability.
- If power enhancement is valid, small-sample brain connectivity studies can find more differential links at the same FDR, as illustrated in the paper's application.
- The method handles binary, count, and transformed heavy-tailed network data, so it applies beyond Gaussian settings.
Reading between the lines
- A direct extension would apply the same sum-based auxiliary grouping to comparisons of covariance or precision matrices, where link sums may also carry sparsity information not visible in differences.
- Because the paper's correlation condition is not checked in the brain data, a quick diagnostic is to estimate the largest pairwise correlation among the link-specific variables; if it is near 1, the claimed extreme-value threshold would need correction.
- The data-adaptive choice of the number of groups (beyond the fixed K=3) is left open and could be selected by cross-validation or information criteria.
- The additional links found by the power-enhanced test but missed by the basic test could be treated as candidates for independent replication, since the theory says they are real small changes rather than noise.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes methods for comparing two populations of network data when the observations are symmetric matrices encoding each subject's network. It develops a global test based on the maximum of standardized link-wise differences and proves a Gumbel null limit, claims asymptotic minimax optimality under sparse alternatives, and proposes a simultaneous testing procedure with asymptotic FDR/FDP control. A power-enhanced procedure is then introduced by borrowing auxiliary information from sums of the two sample means, following the grouping-adjusting-pooling idea. The theoretical results are stated as Theorems 1-5, and the numerical section reports simulations and an application to the KKI-42 brain structural connectivity dataset. The main contribution is to handle directly observed network matrices under moment conditions rather than normality, and to combine global and simultaneous inference with power enhancement.
Significance. If the stated results hold, the paper addresses a real gap in network inference: comparing populations of directly observed network matrices with both global and link-wise inference, under moment conditions that cover binary and count data. The max-type global test and the FDR-controlling simultaneous procedure are natural and potentially useful tools, and the power-enhanced variant is well motivated for small-sample brain-connectivity studies. The paper is explicit about its conditions and provides simulation evidence that the procedures control FDR and improve power in the settings considered. However, the weak-dependence conditions are not linked to the brain-network application, and several asymptotic conditions fail for the sample size of the real-data analysis. These gaps currently limit the support for the paper's central practical claims.
major comments (4)
- [Section 3.2, Conditions (A2) and (B2)] The central asymptotic claims require that among the q = p(p-1)/2 link variables there are no large high-correlation neighborhoods: Condition (B2) asks that max_i |A_i(ξ)| = o(q^ν) for some ν < (1-r)/(1+r), and Condition (A2) requires all pairwise correlations to be bounded away from 1. In a structural brain network, links that share a region, especially a hub, are naturally dependent, and a hub of degree O(p) creates a high-correlation neighborhood of size O(p) = O(q^{1/2}). Since q^{1/2} is not o(q^ν) for the values of ν allowed when r is non-negligible, condition (B2) can fail unless the shared-node correlations are all below the (log q)^{-2-ξ} threshold. The simulations in Section 5.1 generate each edge independently, and the KKI-42 analysis reports no diagnostic for (A2) or (B2). Consequently, the paper does not demonstrate that Theorems 1, 3, and 4 apply to the motivating data type, even though those theorems drive the claimed Gumbel null limit and FDR control.
- [Section 5.2 and Theorem 4] The real-data application has n1 = 22, n2 = 20, and p = 68, giving q = 2278 and log q close to 7.7. Theorem 4 requires log q = o(n^{1/C}) for some C > 5; for n ≈ 20, n^{1/6} is about 1.65, so this condition is not met. Condition (C1) also requires log p = o(n^{1/5}), which fails at these sample sizes. Thus the asymptotic FDR control of Theorem 4 and the power comparison of Theorem 5 do not formally cover the KKI-42 analysis. The authors should either treat the real-data results as illustrative, provide finite-sample justification calibrated to the observed dependence structure, or weaken the claims made for this application.
- [Section 5.2, KKI-42 analysis] The paper treats the two resting-state diffusion tensor images from the same subject as independent samples, effectively doubling the sample size. This is a substantial and untested assumption: within-subject scan-rescan dependence can induce correlation between the two groups and among link statistics, and the theory assumes independent subjects. The manuscript mentions this choice but does not discuss its potential impact on the null distribution or on the reported discoveries. I would like to see a sensitivity analysis, or at least an explicit statement of this limitation and a justification that the effect is negligible.
- [Section 4.4, Theorem 5] Theorem 5 states Ψadj ≥ Ψ + o(1), which as written is a no-asymptotic-loss statement and does not exclude equality. The abstract and Section 4 claim that the procedure can 'substantially enhance the power' and is 'asymptotically more powerful', but the theorem does not establish strict power enhancement. The simulations show gains in the particular settings considered, but the formal claim should be adjusted, or the proof should give conditions under which the inequality is strict.
minor comments (3)
- [Section 4.2, Proposition 1] In the definition of a_{i,j}, the denominator appears to contain Var(S_{1,l,i,j}) twice; given the definition of κ_{i,j}, the second variance term should presumably be Var(S_{2,l,i,j}).
- [Section 2.1, Condition (C1)] Condition (C1) is stated as log p = o(n^{1/5}) using n = n1 n2/(n1+n2), but the notation n is defined earlier only in the notation section; please make the dependence on n1 and n2 explicit when the condition is introduced.
- [Algorithm 2 and Section 5.1] There are small typographical issues: Algorithm 2 uses inconsistent step labels such as 'Step 2:3', and Section 5.1 says 'identify matrix' instead of 'identity matrix'. A copy-editing pass is needed.
Circularity Check
No significant circularity: the global/simultaneous tests and the power-enhancement comparison are derived under explicit stated conditions, not fitted to the data; self-citations are not load-bearing.
full rationale
The central derivation is self-contained rather than circular. The test statistics T_{i,j} in equation (4) and the auxiliary statistics A_{i,j} are defined directly from the raw network matrices S_{d,l}, and the null limit in Theorem 1, the FDR limits in Theorems 3 and 4, and the power comparison in Theorem 5 are proved under the stated moment and weak-dependence conditions (C1)/(C2), (A1)-(A2), and (B1)-(B3). No parameter is fitted to a subset of the data and then renamed as a prediction. The minimax lower bound in Theorem 2 is a statement about all alpha-level tests, not a restatement of the proposed test's rejection rule. The auxiliary statistic is justified by Proposition 1, which proves asymptotic independence from T_{i,j} under the null, and the p-value weights in equation (5) are estimated from the data, but their effect on FDR and power is explicitly analyzed rather than assumed. The comparison Psi_adj >= Psi + o(1) is relative to Algorithm 1, so it is a theoretical efficiency statement for the proposed procedure relative to its own baseline, not an externally predicted quantity obtained by fitting. The self-citations to Xia et al. (2019a) for the GAP idea and for the practical choice K >= 4 do not carry the mathematical load; the theorems are proved for the new statistics and conditions. Concerns that conditions (A2)/(B2) might fail for hub-structured brain networks are assumption-validity and correctness concerns, not evidence that any claim reduces to its input by construction.
Assumptions & free parameters
free parameters (4)
- C1, C2 =
data-dependent; chosen as min and max of auxiliary statistics, truncated at +/-16 sqrt(log q)
- N =
chosen so that (log p)^{1/2}/N is about 0.1
- K =
3
- epsilon =
1e-5
assumptions (5)
- domain assumption Moment conditions (C1) or (C2): sub-Gaussian or polynomial tails for the network entries.
- domain assumption Conditions (A1)-(A2): covariance of vech(S_d,l) has bounded eigenvalues and correlations bounded away from 1.
- domain assumption Conditions (B1)-(B3): enough alternatives, sparse high-correlation neighborhoods, and a positive fraction of nulls.
- domain assumption Comparable sample sizes: n1 is asymptotically proportional to n2.
- ad hoc to paper In the real-data application, repeated scans from the same subject are independent samples.
Cite this review
Pith. "Pith review of Hypothesis Testing for Network Data with Power Enhancement." pith.science (2026). https://pith.science/paper/AAUP4TWL
@misc{pith2026190803836,
author = {Pith},
title = {Pith review of: Hypothesis Testing for Network Data with Power Enhancement},
year = {2026},
howpublished = {\url{https://pith.science/paper/AAUP4TWL}},
note = {Machine review of arXiv:1908.03836}
}
read the original abstract
Comparing two population means of network data is of paramount importance in a wide range of scientific applications. Many existing network inference solutions focus on global testing of entire networks, without comparing individual network links. Besides, the observed data often take the form of vectors or matrices, and the problem is formulated as comparing two covariance or precision matrices under a normal or matrix normal distribution. Moreover, many tests suffer from a limited power under a small sample size. In this article, we tackle the problem of network comparison, both global and simultaneous inferences, when the data come in a different format, i.e., in the form of a collection of symmetric matrices, each of which encodes the network structure of an individual subject. Such data format commonly arises in applications such as brain connectivity analysis and clinical genomics. We no longer require the underlying data to follow a normal distribution, but instead impose some moment conditions that are easily satisfied for numerous types of network data. Furthermore, we propose a power enhancement procedure, and show that it can control the false discovery, while it has the potential to substantially enhance the power of the test. We investigate the efficacy of our testing procedure through both an asymptotic analysis and a simulation study under a finite sample size. We further illustrate our method with an example of brain structural connectivity analysis.
Reference graph
Works this paper leans on
-
[1]
Ahn, M., Shen, H., Lin, W., and Zhu, H. (2015). A sparse reduced rank framework for group analysis of functional neuroimaging data. Statistica Sinica , 25:295--312
work page 2015
-
[2]
Benjamini, Y. and Hochberg, Y. (1995). Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, Series B. , 57:289--300
work page 1995
-
[3]
Bickel, P. J., Levina, E., et al. (2008). Regularized estimation of large covariance matrices. The Annals of Statistics , 36:199--227
work page 2008
-
[4]
Bullmore, E. and Sporns, O. (2009). Complex brain networks: graph theoretical analysis of structural and functional systems. Nature reviews. Neuroscience , 10:186--198
work page 2009
-
[5]
Cai, T. T., Liu, W., and Xia, Y. (2013). Two-sample covariance matrix testing and support recovery in high-dimensional and sparse settings. Journal of the American Statistical Association , 108:265--277
work page 2013
-
[6]
Cai, T. T., Liu, W., and Xia, Y. (2014). Two-sample test of high dimensional means under dependency. Journal of the Royal Statistical Society, Series B. , 76:349--372
work page 2014
-
[7]
Cai, T. T. and Zhang, A. (2016). Inference for high-dimensional differential correlation matrices. Journal of multivariate analysis , 143:107--126
work page 2016
-
[8]
Chen, S., Kang, J., Xing, Y., and Wang, G. (2015). A parsimonious statistical method to detect groupwise differentially expressed functional connectivity networks. Human Brain Mapping , 36:5196--5206
work page 2015
Show all 40 references
-
[9]
S., S \'e gonne, F., Fischl, B., Quinn, B
Desikan, R. S., S \'e gonne, F., Fischl, B., Quinn, B. T., Dickerson, B. C., Blacker, D., Buckner, R. L., Dale, A. M., Maguire, R. P., Hyman, B. T., et al. (2006). An automated labeling system for subdividing the human cerebral cortex on mri scans into gyral based regions of i...
2006
-
[10]
and Dunson, D
Durante, D. and Dunson, D. B. (2018). Bayesian inference and testing of group differences in brain networks. Bayesian Analysis , 13:29--58
2018
-
[11]
Fornito, A., Zalesky, A., and Breakspear, M. (2013). Graph analysis of the human connectome: Promise, progress, and pitfalls. NeuroImage , 80:426--444
2013
-
[12]
Fox, M. D. and Greicius, M. (2010). Clinical applications of resting state functional connectivity. Frontiers in Systems Neuroscience , 4
2010
-
[13]
E., Li, J., Balachandran, P., Rosenberg, S., Kolaczyk, E
Ginestet, C. E., Li, J., Balachandran, P., Rosenberg, S., Kolaczyk, E. D., et al. (2017). Hypothesis testing for network data in functional neuroimaging. The Annals of Applied Statistics , 11(2):725--750
2017
-
[14]
Han, F., Han, X., Liu, H., and Caffo, B. (2016). Sparse median graphs estimation in a high-dimensional semiparametric model. The Annals of Applied Statistics , 10:1397--1426
2016
-
[15]
R., Mueller, B
Kim, J., Wozniak, J. R., Mueller, B. A., Shen, X., and Pan, W. (2014). Comparison of statistical tests for group differences in brain functional networks. NeuroImage , 101:681--694
2014
-
[16]
Kolaczyk, E., Lin, L., Rosenberg, S., and Walters, J. (2019). Averages of unlabeled networks: Geometric characterization and asymptotic behavior. Annals of Statistics , to appear
2019
-
[17]
A., Huang, A
Landman, B. A., Huang, A. J., Gifford, A., Vikram, D. S., Lim, I. A. L., Farrell, J. A., Bogovic, J. A., Hua, J., Chen, M., Jarso, S., et al. (2011). Multi-parametric neuroimaging reproducibility: a 3-t resource study. Neuroimage , 54:2854--2866
2011
-
[18]
and Chen, S
Li, J. and Chen, S. X. (2012). Two sample tests for high-dimensional covariance matrices. The Annals of Statistics , 40:908--940
2012
-
[19]
Liu, W. (2013). Gaussian graphical model estimation with false discovery rate control. The Annals of Statistics , 41:2948--2978
2013
-
[20]
M., Madan Babu, M., Yu, H., Snyder, M., Teichmann, S
Luscombe, N. M., Madan Babu, M., Yu, H., Snyder, M., Teichmann, S. A., and Gerstein, M. (2004). Genomic analysis of regulatory network dynamics reveals large topological changes. Nature , 431:308--312
2004
-
[21]
Morrison, J. H. and Hof, P. R. (1997). Life and death of neurons in the aging brain. Science , 278:412--419
1997
-
[22]
Qiu, H., Han, F., Liu, H., and Caffo, B. (2016). Joint estimation of multiple graphical models from high dimensional time series. Journal of Royal Statistical Society, Series B. , 78:487--504
2016
-
[23]
Raichle, M. E. and Gusnard, D. A. (2002). Appraising the brain's energy budget. Proceedings of the National Academy of Sciences , 99:10237--10239
2002
-
[24]
J., Bickel, P
Rothman, A. J., Bickel, P. J., Levina, E., and Zhu, J. (2008). Sparse permutation invariant covariance estimation. Electronic Journal of Statistics , 2:494--515
2008
-
[25]
Schott, J. R. (2007). Some high-dimensional tests for a one-way MANOVA . Journal of Multivariate Analysis , 98:1825--1839
2007
-
[26]
and Spj tvoll, E
Schweder, T. and Spj tvoll, E. (1982). Plots of p-values to evaluate many tests simultaneously. Biometrika , 69:493--502
1982
-
[27]
Srivastava, M. S. and Yanagihara, H. (2010). Testing the equality of several covariance matrices with fewer observations than the dimension. Journal of Multivariate Analysis , 101:1319--1329
2010
-
[28]
Storey, J. D. (2002). A direct approach to false discovery rates. Journal of the Royal Statistical Society, Series B. , 64:479--498
2002
-
[29]
Van de Geer, S., B \"u hlmann, P., Ritov, Y., Dezeure, R., et al. (2014). On asymptotically optimal confidence regions and tests for high-dimensional models. The Annals of Statistics , 42:1166--1202
2014
-
[30]
and Craddock, R
Varoquaux, G. and Craddock, R. C. (2013). Learning and comparing functional connectomes across subjects. NeuroImage , 80:405--415. Mapping the Connectome
2013
-
[31]
Wang, L., Zhang, Z., and Dunson, D. (2017). Common and individual structure of multiple networks. arXiv preprint arXiv:1707.06360
2017 arXiv
-
[32]
B., and Guo, Y
Wang, Y., Kang, J., Kemmer, P. B., and Guo, Y. (2016). An efficient and reliable statistical method for estimating functional connectivity in large scale brain networks using partial correlation. Frontiers in Neuroscience , 10:1--17
2016
-
[33]
Xia, M., Wang, J., and He, Y. (2013). Brainnet viewer: A network visualization tool for human brain connectomics. PLOS ONE , 8:1--15
2013
-
[34]
Xia, Y., Cai, T., and Cai, T. T. (2015). Testing differential networks with applications to the detection of gene-gene interactions. Biometrika , 102:247--266
2015
-
[35]
T., and Sun, W
Xia, Y., Cai, T. T., and Sun, W. (2019a). GAP: a general framework for information pooling in two-sample sparse inference. Journal of the American Statistical Association , to appear
2019
-
[36]
and Li, L
Xia, Y. and Li, L. (2019). Matrix graph hypothesis testing and application in brain connectivity alternation detection. Statistica Sinica , 29:303--328
2019
-
[37]
N., and Jagust, W
Xia, Y., Li, L., Lockhart, S. N., and Jagust, W. J. (2019b). Simultaneous covariance inference for multimodal integrative analysis. Journal of the American Statistical Association , accepted:1--30
2019
-
[38]
Yuan, M. (2010). High dimensional inverse covariance matrix estimation via linear programming. Journal of Machine Learning Research , 11:2261--2286
2010
-
[39]
Zhang, Z., Descoteaux, M., Zhang, J., Girard, G., Chamberland, M., Dunson, D., Srivastava, A., and Zhu, H. (2018). Mapping population-based structural connectomes. NeuroImage , 172:130--145
2018
-
[40]
and Li, L
Zhu, Y. and Li, L. (2018). Multiple matrix gaussian graphs estimation. Journal of the Royal Statistical Society, Series B. , 80:927--950
2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.