REVIEW 3 major objections 4 minor 1 cited by
Reproducing the first and second moments of empirical degree distributions
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A two-parameter canonical model reproduces the sample variance of empirical degree distributions where standard linear exponential random graph models cannot.
desk verdict A useful two-parameter canonical ERG that reproduces degree variance on sparse financial networks; the main approximation is unquantified but actually tiny on their data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing identity is $\mathrm{Var}[k] = \frac{2S}{N} + \frac{2L}{N}(1-\frac{2L}{N})$, which translates the variance of the degree distribution into the total number of links $L$ and the total number of two-stars $S$, where a two-star is a pair of links that share a common node. The mechanism that enforces it is the softened link probability $p_{ij} = \frac{z s_i s_j y^{\kappa_i+\kappa_j}}{1 + z s_i s_j y^{\kappa_i+\kappa_j}}$: replacing observed degrees by expected degrees $\kappa_i$ keeps the model canonical and node-heterogeneous while avoiding the degeneracy that makes the mean-field degree-corrected two-star model collapse to deterministic edges. The two global parameters $z$ and $y$ are then fixed by requiring the expected link count and expected two-star count to equal their empirical values.
What would settle it
On a sparse eMID snapshot, sample many configurations from the fitted fit2SM, measure the variance of the total number of links directly, and compare $4\,\mathrm{Var}[L]/N^2$ with the observed degree variance; if the ratio is not negligible, the central variance-reproduction claim fails. Repeating the comparison with fully self-consistent $\kappa$ values from Eq. (17) would check whether the single-iteration implementation masks the discrepancy.
Extended reading notes
Core claim
The central claim is that a deliberately minimal, canonical model -- the fitness-induced two-star model (fit2SM) -- reproduces both the first and second moments of empirical degree distributions while keeping the explanatory power of the density-corrected Gravity Model. Its link probability is $p_{ij} = \frac{z s_i s_j y^{\kappa_i+\kappa_j}}{1 + z s_i s_j y^{\kappa_i+\kappa_j}}$, where $s_i$ are node strengths used as exogenous fitnesses, $\kappa_i$ are expected degrees, and the two parameters $z,y$ solve the coupled moment equations $\langle L\rangle = L$ and $\langle S\rangle = S$. The variance identity $\mathrm{Var}[k] = \frac{2S}{N} + \frac{2L}{N}(1-\frac{2L}{N})$ shows why matching the two-star count $S$ in expectation is sufficient for matching the variance, up to the neglected $\frac{4\mathrm{Var}[L]}{N^2}$ fluctuation term. In tests on the eMID interbank market, the model reproduces the empirical degree variance in snapshots where the UBCM systematically overestimates it and the dcGM over- or underestimates it, and it yields the smallest errors in spectral-radius reconstruction on daily and weekly aggregations.
Load-bearing premise
The paper assumes that the fluctuation in the total number of links, divided by $N^2$, is negligible when computing the degree variance; if that fluctuation is not tiny, the fit2SM's predicted variance misses the observed one by exactly that amount.
Editorial extensions
If this is right
- On sparse snapshots the UBCM overestimates the degree variance and the dcGM can over- or underestimate it, while the fit2SM reproduces it, so variance-based diagnostics become reliable under a soft-constrained model.
- At daily and weekly time scales the fit2SM reproduces the spectral radius more accurately than both benchmarks, with average absolute errors of 0.79 ± 0.51 and 0.97 ± 0.66.
- Model selection favors the fit2SM over the UBCM and the dcGM on sparse daily and weekly snapshots, with the advantage shrinking on denser quarterly and yearly snapshots.
- As a generative model, the fit2SM can replace the true network topology for computing spectral early-warning z-scores, including in the pre-crisis period.
- Because the two-star count is empirically related to the link count by $S \approx 0.51 L^{1.58}$, the variance constraint can be imposed even when the two-star count is not directly observable.
Reading between the lines
- The same construction should extend to higher-order nonlinear constraints, such as a Strauss model constraining triangle counts, giving a three-parameter canonical model for clustering or higher moments; the paper only hints at this direction.
- Because the implementation freezes $\kappa$ at dcGM expected degrees, the reported probabilities solve a linearized system; testing the fully self-consistent equations would show whether variance reproduction survives exact iteration.
- The fit2SM needs only strengths and two global parameters, so it is a candidate reconstruction tool for other sparse economic networks, such as trade, input-output, or payment systems, where counterparty links are unobserved but total activity per node is known.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a new exponential random graph model, the fitness-induced two-star model (fit2SM), which is designed to reproduce the first and second moments of an empirical degree distribution while remaining in the canonical (soft-constrained) framework. After deriving an exact relationship between the sample variance of the degree distribution and the number of links L and two-stars S, the authors argue that the problematic term involving Var[L] in the expected variance may be neglected, arriving at a two-parameter model whose probabilities are z s_i s_j y^{kappa_i+kappa_j}/(1+z s_i s_j y^{kappa_i+kappa_j}), with z and y fixed by the constraints <L>=L* and <S>=S*. The model is tested on the eMID interbank network at five temporal aggregations, with claims that it reproduces the degree variance, improves spectral-radius estimation at daily and weekly scales, and outperforms the UBCM and dcGM under BIC on sparse snapshots. The manuscript also contains a negative result on the mean-field degree-corrected two-star model, which degenerates when both degree and two-star constraints are imposed.
Significance. If the central claim is established, the fit2SM is a genuinely useful minimal model: it reproduces the first two moments of the degree distribution with only two global parameters and arbitrary node fitnesses, it is fast to solve, and it appears to improve on both the UBCM and the dcGM for sparse snapshots. The algebraic derivation connecting the degree variance to L and S is clean, the numerical solver reproduces the target constraints to high precision (Table I), and the code is publicly available. The spectral-radius and BIC comparisons are also welcome because they test consequences of the second moment rather than merely re-fitting it. However, the headline variance-reproduction claim is only approximate as stated, and the approximation is never quantified; this limits the significance of the paper unless the authors either measure the omitted term or explicitly reframe the claim as an approximation with a bounded error.
major comments (3)
- [Section V, Eqs. (6)-(12)] The central claim that the fit2SM reproduces the sample variance is exact only if the term -4Var[L]/N^2 is negligible, but this is asserted rather than demonstrated. For any dyad-independent model, Eq. (6) gives <Var[k]> = 2<S>/N - 4Var[L]/N^2 + 2<L>/N(1 - 2<L>/N). Since the constraints in Eq. (15) fix only <L> and <S>, the model's expected variance differs from the empirical sample variance by -4Var[L]/N^2. The sentence 'such an addendum may, thus, be expected not to play a relevant role' is not a quantitative argument. A simple bound is Var[L] = sum_{i<j} p_ij(1-p_ij) <= <L> = L, so the omitted term is at most 4L/N^2; for the tested sparse eMID snapshots this may indeed be small, but the paper never reports its value. I request that the authors compute Var[L] from the fitted probabilities for the reported snapshots, and either state the resulting relative error explicitly or rewrite the claim as holding up to an O(4L/N^2) correction.
- [Fig. 1 and Section V] The headline empirical support for variance reproduction consists of only two weekly snapshots in Fig. 1. Given that <Var[k]> is determined, up to the omitted Var[L] term, by the two constraints in Eq. (15), agreement in two cases is more a consistency check of the approximation than an independent model test. The manuscript should systematically report, for all snapshots and temporal aggregations, the empirical Var[k] versus the model's <Var[k]>, together with the estimated omitted term 4Var[L]/N^2 or its relative contribution. Without this, the statement that the fit2SM 'correctly reproduces' the sample variance is not supported by the presented evidence.
- [Eqs. (14), (17)-(18) and Appendix D] The model as implemented uses kappa_i fixed at the dcGM values from the single-iteration initialization, as Eq. (26) and Appendix D make clear, so the reported probabilities are not the self-consistent solutions of the non-linear system advertised in Eqs. (14)-(17). The authors do compare single-iteration and self-consistent solutions for BIC in Appendix D, but not for the variance-reproduction or spectral-radius claims. Please clarify which version of the model is used for each reported result, and show that the single-iteration choice does not materially affect the variance-reproduction and spectral-radius conclusions, not only the BIC values.
minor comments (4)
- [Section VI C] The bullet on the monthly time-scale contains an apparent contradiction: it says BIC_fit2SM < BIC_UBCM on 33.94% of monthly snapshots but then says BIC_fit2SM < BIC_UBCM on 97.78% of monthly snapshots since 2009. Please rephrase to distinguish the full-sample result from the post-2009 subsample.
- [Eqs. (16)-(17)] The iteration indices are confusing: Eq. (16) uses t for the z and y updates, while Eq. (17) uses e for epochs; the relationship between these two iteration levels should be defined explicitly.
- [Appendix D] The statement that numerical errors 'never exceed O(10^-1)' is imprecise given that Table I lists maximum relative errors around 10^-9 to 10^-11; please state the actual error ranges.
- [Section VII and Appendix D] The main text reports the S=aL^b fit with a~0.51, b~1.58, while Appendix D reports aggregation-specific values such as a_daily~0.36, a_yearly~0.69; please reconcile the two presentations so that the reader understands whether the main-text values are averages of the aggregation-specific fits or a single pooled fit.
Circularity Check
The variance-reproduction claim reduces to the model's own L and S constraints; only the spectral-radius and BIC comparisons provide independent evidence.
-
self definitional
[Section II Eq. (5); Section V Eqs. (6), (12), (15), and the sentence after Eq. (17)]
"Var[k] = k2 − k̄2 = 2S/N + 2L/N(1 − 2L/N). ... ⟨Var[k]⟩ = 2⟨S⟩/N − 4Var[L]/N2 + 2⟨L⟩/N(1−2⟨L⟩/N). ... The system above ensures that the fit2SM correctly reproduces the total number of links, the total number of two-stars and the sample variance."
Eq. (5) makes the empirical variance an exact function of the two aggregate observables L and S. The fit2SM parameters z and y are then solved from Eq. (15) so that ⟨L⟩=L* and ⟨S⟩=S*. Consequently, the model's expected variance, Eq. (6) under the paper's approximation Eq. (12), is algebraically forced to equal Var[k] up to the dropped −4Var[L]/N2 term. Reproducing the second moment is therefore not an independent prediction of the fitted model; it is a restatement of the constraints used to fit the two parameters. The paper itself states that reproducing the sample variance requires the total number of links to be reproduced exactly, so the claimed variance reproduction is by construction.
full rationale
The headline claim that the fit2SM reproduces the sample variance is circular in the precise sense that the model is fitted to the two aggregates that algebraically determine that variance. Because Eq. (5) expresses Var[k] in terms of L and S and Eq. (15) fixes ⟨L⟩ and ⟨S⟩ to their empirical values, the variance match is automatic up to the explicitly dropped Var[L] term; it is a constraint, not a prediction. The paper is transparent about deriving the variance from the constraints, but the claim is still definitional rather than independent. The empirical spectral-radius and BIC comparisons are not forced by the constraints and provide genuine evidence of performance, so the circularity is partial. The unquantified dismissal of −4Var[L]/N2 as 'not expected to play a relevant role' is an approximation gap rather than circularity, but it means the exact variance-reproduction claim is not established on the reported evidence. No load-bearing self-citation chain appears: citations to the authors' earlier work are used as benchmarks and methods, not to justify the fit2SM construction. The Appendix D single-iteration implementation is a self-consistency issue, not a circularity issue. Overall score 6: one central prediction reduces by construction, while other tests and the proposed model remain substantive.
Assumptions & free parameters
free parameters (3)
- z =
varies by snapshot; solved from Eq. (15)
- y =
varies by snapshot; roughly 0.85-1.05 in Fig. 7
- a, b (S=aL^b) =
a ~ 0.36-0.69, b ~ 1.56-1.61 by aggregation
assumptions (6)
- domain assumption Maximum-entropy canonical ensemble with soft constraints is the appropriate framework for network reconstruction.
- domain assumption The mean-field approximation replaces the graph A by its expectation P in separable Hamiltonians.
- ad hoc to paper The term 4Var[L]/N^2 in Eq. (6) is negligible.
- domain assumption Node strengths s_i are observable and can serve as exogenous fitnesses.
- ad hoc to paper One iteration of the fixed-point scheme, starting from κ_i = κ_i^dcGM, is sufficient.
- domain assumption The eMID symmetrization and time-aggregation procedure yields the correct binary network representation.
Cite this review
Pith. "Pith review of Reproducing the first and second moments of empirical degree distributions." pith.science (2026). https://pith.science/paper/6YFX2WWZ
@misc{pith2026250510373,
author = {Pith},
title = {Pith review of: Reproducing the first and second moments of empirical degree distributions},
year = {2026},
howpublished = {\url{https://pith.science/paper/6YFX2WWZ}},
note = {Machine review of arXiv:2505.10373}
}
read the original abstract
The study of probabilistic models for the analysis of complex networks represents a flourishing research field. Among the former, Exponential Random Graphs (ERGs) have gained increasing attention over the years. So far, only linear ERGs have been extensively employed to gain insight into the structural organisation of real-world complex networks. None, however, is capable of accounting for the variance of the empirical degree distribution. To this aim, non-linear ERGs must be considered. After showing that the usual mean-field approximation forces the degree-corrected version of the two-star model to degenerate, we define a fitness-induced variant of it. Such a `softened' model is capable of reproducing the sample variance, while retaining the explanatory power of its linear counterpart, within a purely canonical framework.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
Missing links prediction: comparing machine learning with physics-rooted approaches
On two economic networks, maximum-entropy network models predict missing links about as accurately as gradient-boosting machine learning, and adding geographic distance makes the physics-style model the better performer.
Reference graph
Works this paper leans on
-
[1]
Colizza, V., Barrat, A., Barth´ elemy, M. & Vespignani, A. The role of the airline transportation network in the prediction and predictability of global epidemics.Proc. Natl. Acad. Sci. U.S.A.103, 2015–2020 (2006)
work page 2006
-
[2]
& Vespignani, A.Dynamical Processes on Complex Networks(Cambridge University Press, Cambridge, 2008)
Barrat, A., Barth´ elemy, M. & Vespignani, A.Dynamical Processes on Complex Networks(Cambridge University Press, Cambridge, 2008)
work page 2008
-
[3]
Newman, M.Networks: An Introduction(Oxford Uni- versity Press, Oxford, 2010)
work page 2010
-
[4]
Pastor-Satorras, R., Castellano, C., Van Mieghem, P. & Vespignani, A. Epidemic processes in complex networks. Rev. Mod. Phys.87, 925 (2015)
work page 2015
-
[5]
Squartini, T., van Lelyveld, I. & Garlaschelli, D. Early- warning signals of topological collapse in interbank net- works.Sci. Rep.3, 3357 (2013)
work page 2013
-
[6]
Battiston, S.et al.Complexity theory and financial reg- ulation.Science351, 818–819 (2016)
work page 2016
-
[7]
Bardoscia, M., Battiston, S., Caccioli, F. & Caldarelli, G. Pathways towards instability in financial networks. Nat. Commun.8, 14416 (2017)
work page 2017
-
[8]
Macchiati, V., Marchese, E., Mazzarisi, P., Garlaschelli, D. & Squartini, T. Spectral signatures of structural change in financial networks.Chaos, Solitons & Fractals 193, 116065 (2025)
work page 2025
Show all 42 references
-
[9]
Jaynes, E. T. Information theory and statistical me- chanics.Phys. Rev.106, 620–630 (1957)
1957
-
[10]
& Newman, M
Park, J. & Newman, M. E. J. Statistical mechanics of networks.Phys. Rev. E70, 066117 (2004)
2004
-
[11]
& Loffredo, M
Garlaschelli, D. & Loffredo, M. I. Maximum likeli- hood: Extracting unbiased information from complex networks.Phys. Rev. E78, 015101 (2008)
2008
-
[12]
& Garlaschelli, D
Squartini, T. & Garlaschelli, D. Analytical maximum- likelihood method to detect patterns in real networks. New J. Phys.13, 083001 (2011)
2011
-
[13]
& Squartini, T
Saracco, F., Di Clemente, R., Gabrielli, A. & Squartini, T. Randomizing bipartite networks: The case of the world trade web.Sci. Rep.5, 10595 (2015)
2015
-
[14]
& Garlaschelli, D
Squartini, T. & Garlaschelli, D. Maximum-entropy net- works.Phys. Rev. E96, 032317 (2017)
2017
-
[15]
Cimini, G.et al.The statistical physics of real-world networks.Nat. Rev. Phys.1, 58–71 (2019)
2019
-
[16]
& Ho lyst, J
Fronczak, A., Fronczak, P. & Ho lyst, J. A. Statistical mechanics of the international trade network.Phys. Rev. E73, 016108 (2006)
2006
-
[17]
Entropy of network ensembles.Phys
Bianconi, G. Entropy of network ensembles.Phys. Rev. E79, 036114 (2009)
2009
-
[18]
& Fronczak, P
Fronczak, A. & Fronczak, P. Statistical mechanics of the international trade network: Structural correlations and modeling.Phys. Rev. E85, 056113 (2012)
2012
-
[19]
& Garlaschelli, D
Squartini, T., Mastrandrea, R. & Garlaschelli, D. Unbi- ased sampling of network ensembles.New J. Phys.17, 023052 (2015). 11
2015
-
[20]
& Reed, B
Molloy, M. & Reed, B. A critical point for random graphs with a given degree sequence.Random Struct. Algorithms6, 161–180 (1995)
1995
-
[21]
& Stone, L
Artzy-Randrup, Y. & Stone, L. Generating uniformly distributed random networks.Physica A364, 513–536 (2006)
2006
-
[22]
I., Kim, H., Toroczkai, Z
Del Genio, C. I., Kim, H., Toroczkai, Z. & Bassler, K. E. Efficient and exact sampling of simple graphs with given arbitrary degree sequence.PLoS One5, e10012 (2010)
2010
-
[23]
& Diaconis, P
Blitzstein, J. & Diaconis, P. A sequential importance sampling algorithm for generating random graphs with prescribed degrees.Internet Math.6, 489–522 (2011)
2011
-
[24]
L., Mikl´ os, I
Kim, J., Toroczkai, Z., Erd˝ os, P. L., Mikl´ os, I. & Kon- dor, I. Network-based prediction of unobserved edges in complex networks.New J. Phys.14, 023012 (2012)
2012
-
[25]
Roberts, E. S. & Coolen, A. C. C. Unbiased degree- preserving randomization of directed binary networks. Phys. Rev. E85, 046103 (2012)
2012
-
[26]
& Garlaschelli, D
Squartini, T., Caldarelli, G., Cimini, G., Gabrielli, A. & Garlaschelli, D. Reconstruction methods for networks: The case of economic and financial systems.Phys. Rep. 757, 1–47 (2018)
2018
-
[27]
& Gabrielli, A
Cimini, G., Squartini, T., Garlaschelli, D. & Gabrielli, A. Systemic risk analysis on reconstructed economic and financial networks.Sci. Rep.5, 15758 (2015)
2015
-
[28]
& Garlaschelli, D
Almog, A., Bird, R. & Garlaschelli, D. Enhanced grav- ity model of trade: Reconciling macroeconomic and net- work models.Front. Phys.7, 55 (2019)
2019
-
[29]
& Buiten, G
Hooijmaaijers, S. & Buiten, G. Methodology for esti- mating the dutch interfirm trade network. Tech. Rep., Statistics Netherlands (2019). Technical Report
2019
-
[30]
Mattsson, C. E. S.et al.Functional structure in pro- duction networks.Front. Big Data4, 666712 (2021)
2021
-
[31]
N.et al.Reconstructing firm-level interac- tions in the dutch input–output network.Sci
Ialongo, L. N.et al.Reconstructing firm-level interac- tions in the dutch input–output network.Sci. Rep.12 (2022)
2022
-
[32]
& Squartini, T.Recon- structing Networks
Cimini, G., Mastrandrea, R. & Squartini, T.Recon- structing Networks. Elements in the Structure and Dy- namics of Complex Networks (Cambridge University Press, 2021)
2021
-
[33]
& Cal- darelli, G
Battiston, S., Puliga, M., Kaushik, R., Tasca, P. & Cal- darelli, G. Debtrank: Too central to fail? financial net- works, the fed and systemic risk.Sci. Rep.2, 541 (2012)
2012
-
[34]
& Vespignani, A
Pastor-Satorras, R. & Vespignani, A. Epidemic spread- ing in scale-free networks.Phys. Rev. Lett.86, 3200 (2001)
2001
-
[35]
& Redner, S
Sood, V. & Redner, S. Voter model on heterogeneous graphs.Phys. Rev. Lett.94, 178701 (2005)
2005
-
[36]
& Redner, S
Sood, V., Antal, T. & Redner, S. Voter models on het- erogeneous networks.Phys. Rev. E77, 041121 (2008)
2008
-
[37]
D., Precup, O., Gabbi, G
Iori, G., Masi, G. D., Precup, O., Gabbi, G. & Cal- darelli, G. A network analysis of the italian overnight money market.J. Econ. Dyn. Control32, 259–278 (2008)
2008
-
[38]
& Lux, T
Finger, K., Fricke, D. & Lux, T. Network analysis of the e-mid overnight money market: The informational value of different aggregation levels for intrinsic dynamic processes. Tech. Rep. 1782, Kiel Institute for the World Economy (2012). Kiel Working Paper
2012
-
[39]
& Newman, M
Park, J. & Newman, M. E. J. Solution of the two-star model of a network.Phys. Rev. E70, 066146 (2004)
2004
-
[40]
J., Cimini, G
Vallarano, N., Straka, M. J., Cimini, G. & Squartini, T. Fast and scalable likelihood maximization for exponen- tial random graph models via gpu parallelization.Sci. Rep.11, 5725 (2021)
2021
-
[41]
Chung, F. & Lu, L. Connected components in ran- dom graphs with given expected degree sequences.Ann. Comb.6, 125–145 (2002)
2002
-
[42]
Horn, R. A. & Johnson, C. R.Matrix Analysis(Cam- bridge University Press, Cambridge, 1985). 12 APPENDIX A. DATA DESCRIPTION AND TEMPORAL AGGREGATIONS We employ transaction-level records from the Electronic Market for Interbank Deposits (eMID), a screen-based market for unsecur...
1985
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.