{"id":"b29198ff-3ccf-4250-b3a5-9a6a99bcb2fd","arxiv_id":"1908.03836","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A two-sample testing framework for network matrices that tests global differences and individual links with FDR control, plus a power-enhancement step using auxiliary statistics.","lead":"This paper builds statistical tests to compare two groups of brain or gene networks, including a way to find which individual connections differ, not just whether the networks differ. The method works directly on the observed network matrices, without assuming the data is normally distributed, and it adds an auxiliary-statistic step that finds more true differences at the same false-discovery level.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"A2/B2 weak-dependence assumptions are unverified and are likely violated for hub-structured networks; the KKI-42 application and simulations do not exercise this regime.","rationale":"I read the paper as making two connected claims: (i) the max-type global test attains the (log q)^{1/2} minimax detection boundary, and (ii) the unenhanced and power-enhanced simultaneous tests control the false discovery rate at the target level while the enhanced test is asymptotically at least as powerful. Both claims depend on the extreme-value approximation of the maximum of q dependent t-like statistics. The weakest link in that chain is Theorem 1, whose conditions (A1)-(A2) are the least protected part of the argument. The paper calls (A2) 'mild' because |r| = 1 would make Gamma singular, but that observation only shows strict correlation less than 1 is necessary; it does not make the condition mild for network data. Structural brain networks have hubs, and edges sharing a hub are dependent through both physiology and the tractography estimation procedure. Condition (B2) restricts the number of such high-correlation neighbors to o(q^nu) with nu < (1-r)/(1+r), which is not a consequence of any stated property of network data and can fail at the paper's own dimensions. The simulations use independent edge generation, so they provide only partial, independence-regime support; the data analysis reports no diagnostic for the dependence conditions. A secondary but related gap is that Theorem 4 additionally assumes log q = o(n^{1/C}) with C > 5, which is not satisfied for n1 = n2 = 25 or for the real sample sizes n1 = 22, n2 = 20; this reinforces, but does not replace, the dependence concern. The theory itself appears coherent under the stated conditions, and the independent-edge simulations in Table 1 give some evidence of practical utility, so the appropriate response is to require a diagnostic or a hub-structured simulation before the practical claims are accepted, not to reject the theoretical framework.","tokens_in":17771,"tokens_out":10418,"duration_ms":120776,"concrete_test":"Simulate the hub regime with the paper's dimensions: p = 68 (q = 2278), n1 = n2 = 25, sparsity kq = 0.2q, and effect sizes as in Section 5.1, but generate each subject's network with a fixed hub h by setting S_{d,l,hj} = mu_{d,hj} + gamma_{d,l} + epsilon_{d,l,hj} for j != h and S_{d,l,ij} = mu_{d,ij} + epsilon_{d,l,ij} otherwise, with gamma and epsilon independent N(0,1) and mu chosen to match the Bernoulli mean differences (or transform the result to Bernoulli by thresholding). Run Algorithms 1 and 2 at alpha = 0.05 over 500 replications. Compare (a) the empirical FDR with the nominal 5% and (b) a Q-Q plot of M_n under the null against the Theorem 1 Gumbel limit. If the empirical FDR exceeds 5% or the Q-Q plot shows systematic departure, the unverified A2/B2 conditions are the likely cause; if neither occurs, the concern is not fatal.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central asymptotic results rest on a weak-dependence condition that is not tailored to the motivating application. Theorem 1's Gumbel limit for M_n requires (A1)-(A2): all q = p(p-1)/2 edge-wise variables have covariance eigenvalues bounded away from 0 and infinity, and all pairwise correlations are bounded away from 1. Theorems 3 and 4 additionally require (B2), that the number of variables whose correlation with a given variable exceeds (log q)^{-2-xi} is o(q^nu) with nu < (1-r)/(1+r). In a structural brain network, links that share a node, especially a hub region, are naturally dependent because they arise from the same tractography stream and the same subject-level physiology. A hub node of degree d makes the high-correlation neighborhood size at least O(d) = O(p) = O(q^{1/2}), which is not o(q^nu) for the small nu allowed when correlations are substantial. With p = 68, q = 2278, this is exactly the regime of Section 5. The paper's simulations generate each edge independently, so they cannot detect any violation of (A2) or (B2), and the KKI-42 analysis reports no diagnostic for these conditions. If they fail, the Gumbel extreme-value approximation and the FDR calculations in Theorems 1, 3, and 4 are not implied by the stated assumptions, which undermines both the claimed minimax optimality and the reported discoveries for the data type the paper targets.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes methods for comparing two populations of network data when the observations are symmetric matrices encoding each subject's network. It develops a global test based on the maximum of standardized link-wise differences and proves a Gumbel null limit, claims asymptotic minimax optimality under sparse alternatives, and proposes a simultaneous testing procedure with asymptotic FDR/FDP control. A power-enhanced procedure is then introduced by borrowing auxiliary information from sums of the two sample means, following the grouping-adjusting-pooling idea. The theoretical results are stated as Theorems 1-5, and the numerical section reports simulations and an application to the KKI-42 brain structural connectivity dataset. The main contribution is to handle directly observed network matrices under moment conditions rather than normality, and to combine global and simultaneous inference with power enhancement.","tokens_in":18013,"tokens_out":7139,"duration_ms":82944,"significance":"If the stated results hold, the paper addresses a real gap in network inference: comparing populations of directly observed network matrices with both global and link-wise inference, under moment conditions that cover binary and count data. The max-type global test and the FDR-controlling simultaneous procedure are natural and potentially useful tools, and the power-enhanced variant is well motivated for small-sample brain-connectivity studies. The paper is explicit about its conditions and provides simulation evidence that the procedures control FDR and improve power in the settings considered. However, the weak-dependence conditions are not linked to the brain-network application, and several asymptotic conditions fail for the sample size of the real-data analysis. These gaps currently limit the support for the paper's central practical claims.","major_comments":[{"comment":"The central asymptotic claims require that among the q = p(p-1)/2 link variables there are no large high-correlation neighborhoods: Condition (B2) asks that max_i |A_i(ξ)| = o(q^ν) for some ν < (1-r)/(1+r), and Condition (A2) requires all pairwise correlations to be bounded away from 1. In a structural brain network, links that share a region, especially a hub, are naturally dependent, and a hub of degree O(p) creates a high-correlation neighborhood of size O(p) = O(q^{1/2}). Since q^{1/2} is not o(q^ν) for the values of ν allowed when r is non-negligible, condition (B2) can fail unless the shared-node correlations are all below the (log q)^{-2-ξ} threshold. The simulations in Section 5.1 generate each edge independently, and the KKI-42 analysis reports no diagnostic for (A2) or (B2). Consequently, the paper does not demonstrate that Theorems 1, 3, and 4 apply to the motivating data type, even though those theorems drive the claimed Gumbel null limit and FDR control.","section":"Section 3.2, Conditions (A2) and (B2)"},{"comment":"The real-data application has n1 = 22, n2 = 20, and p = 68, giving q = 2278 and log q close to 7.7. Theorem 4 requires log q = o(n^{1/C}) for some C > 5; for n ≈ 20, n^{1/6} is about 1.65, so this condition is not met. Condition (C1) also requires log p = o(n^{1/5}), which fails at these sample sizes. Thus the asymptotic FDR control of Theorem 4 and the power comparison of Theorem 5 do not formally cover the KKI-42 analysis. The authors should either treat the real-data results as illustrative, provide finite-sample justification calibrated to the observed dependence structure, or weaken the claims made for this application.","section":"Section 5.2 and Theorem 4"},{"comment":"The paper treats the two resting-state diffusion tensor images from the same subject as independent samples, effectively doubling the sample size. This is a substantial and untested assumption: within-subject scan-rescan dependence can induce correlation between the two groups and among link statistics, and the theory assumes independent subjects. The manuscript mentions this choice but does not discuss its potential impact on the null distribution or on the reported discoveries. I would like to see a sensitivity analysis, or at least an explicit statement of this limitation and a justification that the effect is negligible.","section":"Section 5.2, KKI-42 analysis"},{"comment":"Theorem 5 states Ψadj ≥ Ψ + o(1), which as written is a no-asymptotic-loss statement and does not exclude equality. The abstract and Section 4 claim that the procedure can 'substantially enhance the power' and is 'asymptotically more powerful', but the theorem does not establish strict power enhancement. The simulations show gains in the particular settings considered, but the formal claim should be adjusted, or the proof should give conditions under which the inequality is strict.","section":"Section 4.4, Theorem 5"}],"minor_comments":[{"comment":"In the definition of a_{i,j}, the denominator appears to contain Var(S_{1,l,i,j}) twice; given the definition of κ_{i,j}, the second variance term should presumably be Var(S_{2,l,i,j}).","section":"Section 4.2, Proposition 1"},{"comment":"Condition (C1) is stated as log p = o(n^{1/5}) using n = n1 n2/(n1+n2), but the notation n is defined earlier only in the notation section; please make the dependence on n1 and n2 explicit when the condition is introduced.","section":"Section 2.1, Condition (C1)"},{"comment":"There are small typographical issues: Algorithm 2 uses inconsistent step labels such as 'Step 2:3', and Section 5.1 says 'identify matrix' instead of 'identity matrix'. A copy-editing pass is needed.","section":"Algorithm 2 and Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely acceptable after revision if the authors either verify the weak-dependence conditions for the real-data setting or appropriately reframe the practical claims. The mismatch between the asymptotic assumptions and the brain-network application is the main substantive concern, rather than a technical error in the derivations. The novelty relative to the GAP framework of Xia et al. (2019a) is moderate but sufficient for a statistics journal if the conditions are made credible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper fills a real gap—comparing two populations of directly observed network matrices, with simultaneous link-level tests and a power-enhancement step. The problem is well motivated, the moment conditions are a genuine improvement over normality, and the numerical work is honest about its claims. The global test achieves the minimax rate for sparse alternatives, and the FDR theory is standard but adapted cleanly to this setting.\n\nThe soft spots are not fatal to the core idea, but they matter. The weakest point is the gap between the theory's weak-dependence conditions (A2, B2) and the brain-connectivity application that motivates the paper. In a structural network, edges sharing a node—especially a hub—are naturally dependent, and B2 requires the high-correlation neighborhood of each edge to be o(q^nu) with nu < (1-r)/(1+r). For p=68, q=2278, a hub of degree 20 already gives a neighborhood of order p, which swamps o(q^nu) for the nu values allowed when r is moderate. The paper never checks these conditions on the KKI-42 data, and its simulations generate edges nearly independently, so the violation is never exercised. That doesn't make the method useless; it makes the asymptotic guarantees that are claimed for the application not supported.\n\nSecond, the 42 images from 21 subjects are treated as independent. The paper says this is common for this dataset, but a sensitivity analysis or at least a note on cluster correlation would be needed to trust the discovery counts.\n\nThird, Algorithm 2 selects the grid that maximizes rejections. The theorems treat the final adjusted p-values as if the grid were fixed; the search over JK creates an additional multiple-testing layer that isn't accounted for in the FDR proof. This is a genuine gap, though it may be fixable by a uniform argument.\n\nOn credit: the auxiliary statistic A_{i,j} is cleverly constructed to be asymptotically independent of the test statistic under the null, and the simulations show clear power gains, especially at n=25. The comparison with Xia et al.'s GAP is careful and correct.\n\nBottom line: for a statistician working on network comparisons, this is a useful paper with a promising method, and it deserves a serious referee. My recommendation: send it to review, but the revision must address the weak-dependence gap—either by checking the conditions on the real data, adding robustness checks when they fail, or weakening the assumptions. As it stands, the application is outside the theory's proven regime.","headline":"A genuinely useful two-sample network inference paper with a clever power-enhancement step, but its motivating brain-connectivity application runs outside the theory's weak-dependence assumptions.","tokens_in":18579,"tokens_out":2669,"would_cite":true,"duration_ms":26950,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H15","62F03","62F05","62G20","62H12"],"pacs":[],"model":"deepseek-v4-flash","headline":"Comparing two network populations reduces to a maximum of standardized link differences, and the resulting global test is asymptotically minimax optimal; a sum-based auxiliary statistic then enhances simultaneous-testing power while…","keywords":["network data comparison","two-sample hypothesis testing","power enhancement","false discovery rate","multiple testing","minimax optimality","brain connectivity","moment conditions"],"falsifier":"Simulate two groups of networks where one hub node connects to all other nodes, so many link-specific correlations approach 1, and check whether the null distribution of M_n still matches the claimed extreme-value law and whether the simultaneous procedure keeps FDR at the nominal level. A simpler check on real data would be to compute the maximum absolute correlation among the q link-specific variables in the brain dataset and compare it with the bound required by the paper's condition (A2).","tokens_in":17490,"feed_emoji":"🧠","tokens_out":5860,"duration_ms":57298,"temperature":0.7,"pith_summary":"This paper develops a two-sample hypothesis test for comparing the mean matrices of two collections of network data, where each observation is itself a symmetric matrix of network links. It claims that a global test based on the maximum of standardized link-wise differences is asymptotically minimax optimal, meaning no uniformly better test exists at the detection threshold it achieves. It further claims that a simultaneous testing procedure controls the false discovery rate, and that a power-enhancement step using the sum of the two sample means as auxiliary information is asymptotically more powerful than the unenhanced procedure while still controlling FDR. The method is designed for small-sample settings such as brain connectivity studies and does not require normality, only moment conditions.","feed_headline":"Maximum-link network test reaches minimax power","feed_subtitle":"A new procedure compares mean networks link by link, controls false discoveries, and finds more differences using sum-based auxiliary…","key_machinery":"The load-bearing statistic is $T_{i,j} = (\\bar{S}_{1,i,j} - \\bar{S}_{2,i,j}) / \\sqrt{V_{1,i,j}/n_1 + V_{2,i,j}/n_2}$, the standardized link-wise mean difference, with global statistic $M_n = \\max_{i<j} T_{i,j}^2$ whose null limit is the type-I extreme-value distribution. The power-enhancement machinery is the auxiliary statistic $A_{i,j} = (\\bar{S}_{1,i,j} + \\hat{\\kappa}_{i,j} \\bar{S}_{2,i,j}) / \\sqrt{V_{1,i,j}/n_1 + \\hat{\\kappa}_{i,j}^2 V_{2,i,j}/n_2}$, constructed from the sum rather than the difference, which is asymptotically independent of $T_{i,j}$ under the null; this independence lets the method group tests by $A_{i,j}$, estimate alternative proportions per group, reweight p-values, and apply an FDR-controlling threshold.","core_discovery":"On the paper's own terms, the central discovery is that equality of two network mean matrices can be tested by standardizing each link's mean difference, taking the largest standardized square, and comparing it to the extreme-value distribution; this test detects sparse alternatives with standardized differences of order (log q)^{1/2} and is minimax optimal under mild sparsity. For simultaneous inference, rejecting links whose standardized differences exceed a data-driven threshold controls FDR asymptotically. The paper's power-enhancement claim is that an auxiliary statistic built from the sum of the group means, which is asymptotically independent of the test statistic under the null, can be used to group hypotheses, reweight p-values, and yield asymptotic FDR control with power at least that of the unenhanced test.","pith_inferences":["A direct extension would apply the same sum-based auxiliary grouping to comparisons of covariance or precision matrices, where link sums may also carry sparsity information not visible in differences.","Because the paper's correlation condition is not checked in the brain data, a quick diagnostic is to estimate the largest pairwise correlation among the link-specific variables; if it is near 1, the claimed extreme-value threshold would need correction.","The data-adaptive choice of the number of groups (beyond the fixed K=3) is left open and could be selected by cross-validation or information criteria.","The additional links found by the power-enhanced test but missed by the basic test could be treated as candidates for independent replication, since the theory says they are real small changes rather than noise."],"forward_implications":["If the global minimax claim holds, practitioners can know that no other test can uniformly beat the maximum statistic for detecting sparse mean differences at the (log q)^{1/2} scale.","If FDR control holds, the procedure gives valid simultaneous inference for individual network links without normality and with only moment conditions, broadening applicability.","If power enhancement is valid, small-sample brain connectivity studies can find more differential links at the same FDR, as illustrated in the paper's application.","The method handles binary, count, and transformed heavy-tailed network data, so it applies beyond Gaussian settings."],"supporting_citations":[{"why":"Supplies high-dimensional two-sample mean testing under dependence and the extreme-value limit that motivates the global test statistic.","marker":"Cai et al. (2014)"},{"why":"Earlier method for testing differential networks based on partial correlations that the current data format generalizes.","marker":"Xia et al. (2015)"},{"why":"Matrix-normal based network comparison that this paper extends to directly observed network matrices without normality.","marker":"Xia and Li (2019)"},{"why":"The GAP grouping-adjusting-pooling framework whose ideas the power-enhancement procedure adapts.","marker":"Xia et al. (2019a)"},{"why":"The false-discovery-rate controlling procedure applied to the reweighted p-values.","marker":"Benjamini and Hochberg (1995)"},{"why":"Method used to estimate the proportion of alternatives within each auxiliary-statistic group.","marker":"Storey (2002)"},{"why":"Plotting p-values to estimate the null proportion, used in the group-level adjustment.","marker":"Schweder and Spjøtvoll (1982)"},{"why":"Two-sample covariance testing and support recovery, connected to the transformed Wishart network example.","marker":"Cai et al. (2013)"}],"fun_headline_variants":["Link-wise network comparison reaches minimax power","Power-enhanced network testing with FDR control","Test network links simultaneously with minimax optimality","Network mean test: power enhancement without false discoveries","Extreme-value test for sparse network differences"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main load-bearing premise is that no pair of link-specific variables is too strongly correlated: all correlations among the q entries of each network are bounded away from 1. If many links share a common brain region or hub, this can fail, and then the claimed extreme-value null distribution and FDR control no longer follow from the paper's conditions.","fun_headline_variants_meta":{"raw":{"variants":["Link-wise network comparison reaches minimax power","Power-enhanced network testing with FDR control","Test network links simultaneously with minimax optimality","Network mean test: power enhancement without false discoveries","Extreme-value test for sparse network differences"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000244,"raw_usage":{"total_tokens":1520,"prompt_tokens":918,"completion_tokens":602,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":534}},"tokens_in":534,"tokens_out":602,"duration_ms":6533,"temperature":1.0,"reasoning_tokens":534,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:00:44.623031+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate two groups of networks where one hub node connects to all other nodes, so many link-specific correlations approach 1, and check whether the null distribution of M_n still matches the claimed extreme-value law and whether the simultaneous procedure keeps FDR at the nominal level. A simpler check on real data would be to compute the maximum absolute correlation among the q link-specific variables in the brain dataset and compare it with the bound required by the paper's condition (A2).","supporting_citations":[{"cited_title":"T., Liu, W., and Xia, Y","cited_arxiv_id":null,"evidence_quote":"Supplies high-dimensional two-sample mean testing under dependence and the extreme-value limit that motivates the global test statistic."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier method for testing differential networks based on partial correlations that the current data format generalizes."},{"cited_title":"and Hochberg, Y","cited_arxiv_id":null,"evidence_quote":"The false-discovery-rate controlling procedure applied to the reweighted p-values."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Method used to estimate the proportion of alternatives within each auxiliary-statistic group."},{"cited_title":"and Spj tvoll, E","cited_arxiv_id":null,"evidence_quote":"Plotting p-values to estimate the null proportion, used in the group-level adjustment."},{"cited_title":"T., Liu, W., and Xia, Y","cited_arxiv_id":null,"evidence_quote":"Two-sample covariance testing and support recovery, connected to the transformed Wishart network example."}],"review_version":1}