REVIEW 3 major objections 4 minor 1 cited by
Statistical inference for core-periphery structures
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A single population parameter makes core-periphery structure a statistically testable property, and the maximizing sample metric recovers the true labels exactly under a sparsity condition.
desk verdict A genuine step forward for CP inference—label recovery and the ER-null test are real advances, but the CL-null test (Theorem 3.2) is not proven as stated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the population parameter $\rho(P,c)$, a correlation-like measure between centered edge probabilities and the core-periphery indicator $\Delta_{ij}=c_i+c_j-c_ic_j$, with the sample version $T(A,c)$ replacing $P$ by the adjacency matrix $A$. This sample metric coincides with the Borgatti–Everett template-matching metric, so the new inference attaches statistical meaning to an established descriptive quantity. The proof machinery combines Bernstein-type concentration bounds, union bounds over labelings, and counting lemmas that relate the population gap $\rho(P,c^*)-\rho(P,c)$ to the misclassification fraction $\xi_n(c)$; the hypothesis tests use analytic cutoffs $C_1$ and $C_2$ rather than bootstrap thresholds.
What would settle it
Simulate a Chung-Lu network with no endogenous CP structure, choosing the core size $k$ so that $\alpha_n/\sqrt{\varrho_n}$ is not $o(1/\sqrt{n\log n})$, and run the intersection test of Theorem 3.2; if the rejection rate under this null does not converge to zero, the claimed size control fails in that regime.
Extended reading notes
Core claim
The paper's central claim is that the strength of a core-periphery structure can be quantified at the level of the data-generating mechanism by the parameter $\rho(P,c)$, the normalized centered expected edge count falling on core-core and core-periphery pairs, and that the sample analogue $T(A,c)$ obtained by maximizing over labelings is a statistically valid estimator of the labels. Under the CP-SBM, Chung-Lu, and degree-corrected SBM with the stated separation conditions, Theorem 2.1 gives a misclassification bound of order $o(\alpha_n)$, and under the additional condition $\alpha_n/\sqrt{\varrho_n}=o(1/\sqrt{n\log n})$ the labels are recovered exactly with probability tending to one. The paper further claims that the intersection test based on $T_1(A)=\max_c T(A,c)$ and $T_2(A)=\hat p_{11}-\hat p_{12}$ drives type I error to zero under the Erdős–Rényi and Chung-Lu nulls and power to one under the specified alternatives, thereby distinguishing endogenous CP structure from exogenous structure induced by degree heterogeneity.
Load-bearing premise
The main premise is that the true core is separated from the periphery by a large enough edge-probability gap and is small enough relative to network density; additionally, the Chung-Lu test's error-rate claim depends on a label-perfect-recovery condition that the theorem statement does not list among its assumptions.
Editorial extensions
If this is right
- Maximizing the sample metric becomes a consistent way to find core-periphery labels in stochastic block, Chung-Lu, and degree-corrected block models, with an explicit error rate controlled by the signal-to-noise ratio.
- The two-stage testing recipe gives a practical decision rule: reject the Erdős–Rényi null to detect any CP-like structure, then reject the Chung-Lu null to conclude that the structure is stronger than degree heterogeneity alone can explain.
- Analytic rejection thresholds make the significance tests scalable to large networks, removing the computational bottleneck of bootstrap-based p-values.
- Exact label recovery is guaranteed only when the core is small enough relative to network density, so the theorem identifies the regime in which estimated core labels can be treated as trustworthy.
Reading between the lines
- A natural extension, not developed in the paper, is to adapt $\rho(P,c)$ to weighted or directed networks by replacing the adjacency indicator with a suitable edge-weight summary; the population parameter itself does not depend on the graph being binary.
- The rarity of significant CP structure in the thirteen networks suggests that many reported cores in applied work may be assortative or disassortative community structure rather than true core-periphery structure, though this is an interpretation of the data analysis rather than a proven general claim.
- A configuration-model null that preserves exact degrees would sit between the Erdős–Rényi and Chung-Lu nulls; building the same intersection test around it could clarify whether the authors' Chung-Lu threshold is too conservative for practitioners.
- The size control of the Chung-Lu test in the proof replaces the estimated labels with the true labels using a strong-consistency condition that is not listed among the theorem's stated assumptions; whether this condition is theoretically necessary remains open.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a model-agnostic population parameter rho(P,c) (Eq. 1) and its sample analogue T(A,c) (Eq. 3) to quantify core-periphery strength. It studies these under the ER, CP-SBM, Chung-Lu, and CP-DCBM models, proving label recovery guarantees (Theorem 2.1) and constructing intersection tests against ER and Chung-Lu nulls (Theorems 3.1 and 3.2) with analytic cutoffs. Simulations and thirteen real-world networks are presented. The central claim is that core-periphery structure becomes a testable statistical property, with a formal distinction between endogenous and exogenous CP structure.
Significance. If fully established, this is a valuable contribution: it provides the first recovery guarantees for a CP metric, analytic cutoffs for CP hypothesis tests, and a principled taxonomy of endogenous versus exogenous CP structure. The paper is ambitious and contains substantial proof effort in the supplemental materials, and the empirical evaluation is extensive and reproducible. The main theorems, however, have several load-bearing gaps that prevent the claims from being accepted as stated, particularly in the size control of the Chung-Lu null test and in the power proof of the ER test. These issues are fixable but require added assumptions or modified proof arguments.
major comments (3)
- [Section 3.2, Theorem 3.2; Supplemental Section 12, proof under H0] The type I error proof for the Chung-Lu null test invokes the condition alpha_n/sqrt(rho_n) = o(1/sqrt(n log n)) to obtain strong consistency and then replaces T(A, hat c) with T(A, c*) and rho(hat P, hat c) with rho(hat P, c*). This condition is not stated anywhere in Theorem 3.2. Likewise, the CL strong-consistency result in Theorem 2.1 also requires the ordering condition theta_(k)theta_(n) > theta_(k+1)theta_(k+2), which is also not stated for the null hypothesis in Theorem 3.2. Without these assumptions, the inequality T(A, hat c) <= C1 used to control the size is not established. The theorem as stated is therefore unproven and needs either the missing conditions added or a proof that avoids the strong-consistency step.
- [Theorem 3.1 proof, Supplemental Section 12, around Eq. (20)] The power proof for T2 asserts that 'Since alpha_n >= (log(n log n))^2/n, we have k^{1.5} rho_n log n = o(n^2 alpha_n^2 rho_n)'. This implication is false in general: if alpha_n = c (log n)^2 / n, then k = c log^2 n and the ratio equals 1/sqrt(c), a positive constant, not o(1). The dominance argument requires k >> log^2 n, i.e., alpha_n log^2 n / n -> infinity. Without this stronger condition, the proof that the third term dominates the fluctuation term in Eq. (20) fails, and the power-one claim for T2 under the stated alternatives is not justified. The condition in the theorem should be strengthened accordingly, or the proof revised.
- [Section 2.4-2.5, Eq. (4) and Algorithm 1] There is a mismatch between the theoretical optimization problem and the implemented estimator. Eq. (4) defines hat c as the argmax of T(A,c) over all labelings c, and Algorithm 1 swaps a single node's label at each step, which changes the core size. The theoretical guarantees, however, are proved only for labelings with the same core size as c*: Lemma 2 and the proof of Theorem 2.1 explicitly restrict to 'any c != c* such that the size of the core is the same for c and c*'. The paper states that k is assumed known but does not impose the constraint |c| = k in (4) or in Algorithm 1. Consequently, Theorem 2.1 as proved does not apply to the estimator actually used in simulations and data analysis. Either the optimization and algorithm should be modified to preserve core size k, or the theory must be extended to cover varying k.
minor comments (4)
- [Supplemental Section 12, Eqs. (22) and (24)] The probability statements in (22) and (24) appear reversed: the proofs establish concentration, i.e., P[|...| <= ...] -> 1, but the displayed results state P[|...| <= ...] -> 0. Please correct these directions.
- [Supplemental Section 12, proof of Theorem 3.2 under H1] The text 'strong consistency is achieved as per Theorem 2.3' should refer to Theorem 2.1, not Theorem 2.3.
- [Theorem 3.2, power part] The power proof for T2 also uses the condition alpha_n >= (log(n log n))^2/n (to claim C2 = o(rho_n)), but this condition is not listed among the three assumptions in the theorem statement. Furthermore, as noted in the second major comment, this condition is insufficient for the claimed order, so the same fix is needed here.
- [Remark 1] There is a duplicated 'are' in 'the sharper error bounds in Theorem 2.1 are are crucial'.
Circularity Check
No significant circularity: the recovery and testing theorems are proved from stated model assumptions; the Theorem 3.2 proof gap is a correctness issue, not a circular dependency.
full rationale
The derivation chain is self-contained rather than circular. The population parameter ρ(P,c) in (1) is defined independently of the four models, and the sample estimator T(A,c) in (3) is explicitly identified with the Borgatti–Everett metric rather than being relabeled as a new result. The consistency results in Theorem 2.1 are proved from Assumption A1 plus model-specific separation conditions in Lemmas 2–4, and the proof shows—rather than assumes—that the true model labels maximize ρ(P,c) (Supplemental Section 11). For the Chung–Lu model, the true labels are a definitional choice (the k vertices with largest θ values), but recovering those labels from adjacency data is still a nontrivial estimation theorem rather than a restatement of the definition. Theorem 3.1's ER-null type I error follows from ρ(P,c)=0 for all c under ER and the concentration bound in Lemma 1; Theorem 3.2's CL-null control uses plug-in estimates of the null model, which is a standard testing construction, and the claimed size and power are not obtained by fitting the target result. The self-citations (Yanchenko 2022 for the greedy algorithm; Yanchenko and Sengupta 2023 and 2024 for background) are not load-bearing: the theory analyzes the argmax of T(A,c) rather than the greedy implementation, and the cited works supply an algorithm or context, not the central premises. One caveat should be flagged as a correctness risk rather than circularity: in the Supplemental proof of Theorem 3.2, the line stating that strong consistency is achieved as per Theorem 2.1 under α_n/√ϱ_n = o(1/√(n log n)) invokes a condition not stated in Theorem 3.2, and the Chung–Lu strong-consistency result also requires θ(k)θ(n)>θ(k+1)θ(k+2). If those conditions fail, the type I error control in Theorem 3.2 is unproven; this is a missing-assumption issue in the proof, not a circular step.
Assumptions & free parameters
assumptions (5)
- domain assumption Edges are independent Bernoulli draws given P_ij (A_ij | P_ij ~ Bernoulli(P_ij)).
- domain assumption The core size k is known, alpha_n = k/n -> 0, and n*rho_n*alpha_n -> infinity (Assumption A1).
- domain assumption Separation conditions for identifiable CP structure: p11 > p12 > p22 for CP-SBM; theta(k)*theta(n) > theta(k+1)*theta(k+2) for CL; p12*theta_c*theta_p,min > p22*theta_p,max^2 for DCBM.
- ad hoc to paper Strong consistency condition alpha_n/sqrt(rho_n) = o(1/sqrt(n log n)) is needed in the proof of Theorem 3.2 but is omitted from the theorem statement.
- standard math Standard large-deviation and Taylor expansion machinery, including Bernstein's inequality, first-order expansions, and union bounds.
Cite this review
Pith. "Pith review of Statistical inference for core-periphery structures." pith.science (2026). https://pith.science/paper/EJ64YZPI
@misc{pith2026250804730,
author = {Pith},
title = {Pith review of: Statistical inference for core-periphery structures},
year = {2026},
howpublished = {\url{https://pith.science/paper/EJ64YZPI}},
note = {Machine review of arXiv:2508.04730}
}
read the original abstract
Core-periphery (CP) structure is an important meso-scale network property where nodes group into a small, densely interconnected {core} and a sparse {periphery} whose members primarily connect to the core rather than to each other. While this structure has been observed in numerous real-world networks, there has been minimal statistical formalization of it. In this work, we develop a statistical framework for CP structures by introducing a model-agnostic and generalizable population parameter which quantifies the strength of a CP structure at the level of the data-generating mechanism. We study this parameter under four canonical random graph models and establish theoretical guarantees for label recovery, including exact label recovery. Next, we construct intersection tests for validating the presence and strength of a CP structure under multiple null models, and prove theoretical guarantees for type I error and power. These tests provide a formal distinction between exogenous (or induced) and endogenous (or intrinsic) CP structure in heterogeneous networks, enabling a level of structural resolution that goes beyond merely detecting the presence of CP structure. The proposed methods show excellent performance on synthetic data, and our applications demonstrate that statistically significant CP structure is somewhat rare in real-world networks.
Figures
Figures from the paper (6 more)
Forward citations
Cited by 1 Pith paper
-
Community Detection on a Randomly Growing Network
Global community recovery is impossible under planted preferential-attachment forests with ER noise, but central nodes can be recovered consistently via degree-based pruning and anchor propagation.
Reference graph
Works this paper leans on
-
[1]
Adamic, L. A. and Glance, N. (2005). The political blogosphere and the 2004 us election: divided they blog. In Proceedings of the 3rd international workshop on Link discovery , pages 36--43
work page 2005
-
[2]
Bhadra, S., Pensky, M., and Sengupta, S. (2025). Scalable community detection in massive networks via predictive assignment. arXiv preprint arXiv:2503.16730
work page Pith review arXiv 2025
-
[3]
Bickel, P. J. and Chen, A. (2009). A nonparametric view of network models and newman--girvan and other modularities. Proceedings of the National Academy of Sciences , 106(50):21068--21073
2009
-
[4]
Borgatti, S. P. and Everett, M. G. (2000). Models of core/periphery structures. Social Networks , 21(4):375--395
work page 2000
-
[5]
Boyd, J. P., Fitzgerald, W. J., and Beck, R. J. (2006). Computing core/periphery structures and permutation tests for social relations data. Social Networks , 28(2):165--178
work page 2006
-
[6]
Boyd, J. P., Fitzgerald, W. J., Mahutga, M. C., and Smith, D. A. (2010). Computing continuous core/periphery structures for social relations data with minres/svd. Social Networks , 32(2):125--137
work page 2010
-
[7]
Chakrabarty, S., Sengupta, S., and Chen, Y. (2025). Subsampling based community detection for large networks. Statistica Sinica , 35(3):1--42
work page 2025
-
[8]
Cho, A., Shin, J., Hwang, S., Kim, C., Shim, H., Kim, H., Kim, H., and Lee, I. (2014). Wormnet v3: a network-assisted hypothesis-generating server for caenorhabditis elegans. Nucleic acids research , 42(W1):W76--W82
work page 2014
Show all 51 references
-
[9]
and Lu, L
Chung, F. and Lu, L. (2002). The average distances in random graphs with given expected degrees. Proceedings of the National Academy of Sciences , 99(25):15879--15882
2002
-
[10]
Csermely, P., London, A., Wu, L.-Y., and Uzzi, B. (2013). Structure and dynamics of core/periphery networks . Journal of Complex Networks , 1(2):93--123
2013
-
[11]
Csárdi, G., Nepusz, T., Traag, V., Horvát, S., Zanini, F., Noom, D., and Müller, K. (2025). igraph : Network Analysis and Visualization in R . R package version 2.1.4
2025
-
[12]
H., and Porter, M
Cucuringu, M., Rombach, P., Lee, S. H., and Porter, M. A. (2016). Detection of core--periphery structure in networks using spectral methods and geodesic paths. European Journal of Applied Mathematics , 27(6):846--887
2016
-
[13]
and Sengupta, S
Dasgupta, A. and Sengupta, S. (2022). Scalable estimation of epidemic thresholds via node sampling. Sankhya A , 84:321--344
2022
-
[14]
Elliott, A., Chiu, A., Bazzi, M., Reinert, G., and Cucuringu, M. (2020). Core--periphery structure in directed networks. Proceedings of the Royal Society A , 476(2241):20190783
2020
-
[15]
and R \'e nyi, A
Erd \"o s, P. and R \'e nyi, A. (1959). On random graphs. Publicationes Mathematicae Debrecen , 6:290--297
1959
-
[16]
G., Fullin, K., Gutierrez, G., Omodt, N., Zinnecker, S., Sprint, G., and McCulloch, S
Fink, C. G., Fullin, K., Gutierrez, G., Omodt, N., Zinnecker, S., Sprint, G., and McCulloch, S. (2023). A centrality measure for quantifying spread on weighted, directed networks. Physica A
2023
-
[17]
J., Young, J.-G., and Welles, B
Gallagher, R. J., Young, J.-G., and Welles, B. F. (2021). A clarified typology of core-periphery structure in networks. Science advances , 7(12):eabc9800
2021
-
[18]
Gao, J., Liang, F., Fan, W., Sun, Y., and Han, J. (2009). Graph-based consensus maximization among multiple supervised and unsupervised models. Advances in neural information processing systems , 22
2009
-
[19]
and Cunningham, P
Greene, D. and Cunningham, P. (2013). Producing a unified graph representation from multiple social network views. In Proceedings of the 5th annual ACM web science conference , pages 118--121
2013
-
[20]
Holland, P., Laskey, K., and Leinhardt, S. (1983). Stochastic blockmodels: first steps. Social Networks , 5:109--137
1983
-
[21]
Holme, P. (2005). Core-periphery organization of complex networks. Phys. Rev. E , 72:046111
2005
-
[22]
Ji, M., Sun, Y., Danilevsky, M., Han, J., and Gao, J. (2010). Graph regularized transductive classification on heterogeneous information networks. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases , pages 570--586. Springer
2010
-
[23]
and Newman, M
Karrer, B. and Newman, M. E. J. (2011). Stochastic blockmodels and community structure in networks. Physical Review E , 83:016107
2011
-
[24]
and Masuda, N
Kojaku, S. and Masuda, N. (2017). Finding multiple core-periphery pairs in networks. Physical Review E , 96(5):052313
2017
-
[25]
and Masuda, N
Kojaku, S. and Masuda, N. (2018). Core-periphery structure requires something else in the network. New Journal of Physics , 20(4):043012
2018
-
[26]
Krugman, P. (1996). The self organizing economy . John Wiley & Sons
1996
-
[27]
and Krevl, A
Leskovec, J. and Krevl, A. (2014). SNAP Datasets : Stanford large network dataset collection. http://snap.stanford.edu/data
2014
-
[28]
and Mcauley, J
Leskovec, J. and Mcauley, J. (2012). Learning to discover social circles in ego networks. Advances in neural information processing systems , 25
2012
-
[29]
and Sallan, J
Lordan, O. and Sallan, J. M. (2017). Analyzing the multilevel structure of the european airport network. Chinese Journal of Aeronautics , 30(2):554--560
2017
-
[30]
and Sallan, J
Lordan, O. and Sallan, J. M. (2019). Core and critical cities of global region airport networks. Physica A: Statistical mechanics and its applications , 513:724--733
2019
-
[31]
Lovekar, K., Sengupta, S., and Paul, S. (2025). Testing for the network small-world property. Electronic Journal of Statistics , 19(1):1453--1506
2025
-
[32]
M., Laffan, B., and Schweiger, C
Magone, J. M., Laffan, B., and Schweiger, C. (2016). Core-periphery relations in the European Union: Power and conflict in a dualist political economy . Routledge
2016
-
[33]
Michalski, R., Palus, S., and Kazienko, P. (2011). Matching organizational structure and social network extracted from email communication. In Business Information Systems: 14th International Conference, BIS 2011, Pozna \'n , Poland, June 15-17, 2011. Proceedings 14 , pages 19...
2011
-
[34]
Naik, C., Caron, F., and Rousseau, J. (2021). Sparse networks with core-periphery structure . Electronic Journal of Statistics , 15(1):1814 -- 1868
2021
-
[35]
Nepusz, T., Petr \'o czi, A., N \'e gyessy, L., and Bazs \'o , F. (2008). Fuzzy communities and the concept of bridgeness in complex networks. Physical Review E , 77(1):016107
2008
-
[36]
Newman, M. (2018). Networks . Oxford university press
2018
-
[37]
Newman, M. E. J. (2010). Networks: An Introduction . Oxford University Press
2010
-
[38]
A., Fowler, J
Rombach, P., Porter, M. A., Fowler, J. H., and Mucha, P. J. (2017). Core-periphery structure in networks (revisited). SIAM Review , 59(3):619--646
2017
-
[39]
Rossi, R. A. and Ahmed, N. K. (2015). The network data repository with interactive graph analytics and visualization. In AAAI
2015
-
[40]
D., and Lehmann, S
Sapiezynski, P., Stopczynski, A., Lassen, D. D., and Lehmann, S. (2019). Interaction data from the copenhagen networks study. Scientific Data , 6(1):315
2019
-
[41]
Sengupta, S. (2023). Statistical network analysis: Past, present, and future. arXiv preprint arXiv:2311.00122
2023 arXiv
-
[42]
and Chen, Y
Sengupta, S. and Chen, Y. (2015). Spectral clustering in heterogeneous networks. Statistica Sinica , pages 1081--1106
2015
-
[43]
and Chen, Y
Sengupta, S. and Chen, Y. (2018). A block model for node popularity in networks with community structure. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , 80(2):365--386
2018
-
[44]
Stehl \'e , J., Voirin, N., Barrat, A., Cattuto, C., Isella, L., Pinton, J.-F., Quaggiotto, M., Van den Broeck, W., R \'e gis, C., Lina, B., et al. (2011). High-resolution measurements of face-to-face contact patterns in a primary school. PloS one , 6(8):e23176
2011
-
[45]
Vanhems, P., Barrat, A., Cattuto, C., Pinton, J.-F., Khanafer, N., Regis, C., Kim, B.-A., Comte, B., and Voirin, N. (2013). Estimating potential infection transmission routes in hospital wards using wearable proximity sensors. PloS one , 8:e73970
2013
-
[46]
Yanchenko, E. (2022). A divide-and-conquer algorithm for core-periphery identification in large networks. Stat , page e475
2022
-
[47]
and Sengupta, S
Yanchenko, E. and Sengupta, S. (2023). Core-periphery structure in networks: a statistical exposition. Statistic Surveys , 17:42--74
2023
-
[48]
and Sengupta, S
Yanchenko, E. and Sengupta, S. (2024). A generalized hypothesis test for community structure in networks. Network Science , 12(2):122--138
2024
-
[49]
and Moore, C
Zhang, P. and Moore, C. (2014). Scalable detection of statistically significant communities and hierarchies, using message passing for modularity. Proceedings of the National Academy of Sciences , 111(51):18144--18149
2014
-
[50]
Zhang, X., Martin, T., and Newman, M. E. J. (2015). Identification of core-periphery structure in networks. Physical Review E , 91:032803
2015
-
[51]
Zhao, Y., Levina, E., and Zhu, J. (2012). Consistency of community detection in networks under degree-corrected stochastic block models. The Annals of Statistics , 40(4):2266--2292
2012
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.