{"id":"14055937-4636-4686-9912-be0f42e7001a","arxiv_id":"2608.10845","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Under the random dot product graph model, the rows of degree-α spectral embeddings are asymptotically Gaussian with explicit covariance, and the preferred normalization degree depends on network density and community balance.","lead":"This paper proves a central limit theorem for a continuous family of degree-normalized spectral embeddings of random graphs, covering the adjacency matrix and symmetric Laplacian as special cases. It then uses the limiting distributions to show that the best degree normalization depends on network density, community imbalance, and block structure, not a universal default.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the row-wise CLT proof is internally consistent and the stated assumptions are used as declared.","rationale":"The reader identified Condition 1(c) as the weakest assumption. This is a real limitation of scope, especially for degree-corrected models with degree-correction parameters arbitrarily close to zero, but it is an explicit assumption of Theorem 1 rather than a hidden or circular step. My own check of the proof found no internal inconsistency: the spectral remainder is handled correctly through the exchangeability argument, the Taylor expansion of the degree matrix has the right leading terms, the diagonal self-loop correction is negligible at the stated scaling, and the Lindeberg condition holds under Condition 1. The comparison in Section 5 is honestly labeled as an idealized diagnostic, and the paper openly discusses why alpha* need not match the empirical k-means minimizer. Therefore the central claim, as qualified in the paper, is not undermined. I recommend keeping the reader's ACCEPT verdict unchanged.","tokens_in":36556,"tokens_out":36897,"duration_ms":291152,"concrete_test":"Independently recompute the leading-term expansion in the proof of Theorem 1 by expanding D^{-alpha} A D^{-alpha} - T^{-alpha} P T^{-alpha} in powers of D-T and verifying that, after multiplying by n^{alpha+1/2} rho^alpha, all remainder terms in Lemmas 8-10 are o_P(1). As a second check, simulate a symmetric two-community SBM (p=0.3, q=0.1, n=2000) and compare the empirical covariance of n^{alpha+1/2} rho^alpha (\\hat X_alpha Q - X_alpha) rows with the formula Sigma_alpha(xi_i) for alpha in {0, 0.5, 1}; if the entrywise relative error exceeds 10%, the covariance formula or scaling is wrong.","verdict_should_be":"UNCHANGED","load_bearing_attack":"I read the proof of Theorem 1 in detail, including Lemmas 1-10 and Proposition 1. The expansion of E_alpha, the treatment of the spectral remainder via the exchangeability argument, and the Lindeberg-Feller step all check out. Condition 1(c) is the only structural restriction that could fail in practice (e.g., DCSBM with degree-correction parameters approaching zero), but it is explicitly assumed and is not used beyond what is stated; moreover, under Condition 1(a) with no zero latent positions, the uniform lower bound <x,mu> >= c is essentially automatic since the support lies in a closed cone with pairwise nonnegative dot products. The finite-sample discrepancy between alpha* and the empirical minimizer is acknowledged and does not affect the central theorem. I could not identify a load-bearing objection to the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the family of degree-normalized graph matrices D^{-α}AD^{-α}, α∈[0,1], in the random dot product graph model. The main result is a conditional row-wise central limit theorem for the rank-r spectral embedding: after a suitable orthogonal alignment, each row is asymptotically Gaussian with mean equal to the corresponding row of the population embedding and covariance given explicitly in terms of the latent-position distribution. The theorem covers both dense graphs and sparse graphs with expected degree growing faster than (log n)^2. The authors specialize the result to degree-corrected stochastic block models, describe how degree normalization changes the population geometry and the covariance shape, and introduce a projected-Gaussian Bayes-error diagnostic to compare values of α for two-community SBMs. The predicted ordering is checked against simulations of row-normalized spectral clustering.","tokens_in":36640,"tokens_out":31021,"duration_ms":285345,"significance":"The result is significant because it unifies and extends the known CLTs for adjacency and symmetric Laplacian embeddings and makes the dependence on α explicit at the second order, which is exactly the order needed to compare normalizations after row normalization. The appendix contains a complete proof with the key algebra spelled out; the covariance formulas are internally consistent, and the finite-sample discrepancy between α* and the empirical minimizer is reported honestly rather than explained away. The paper also makes falsifiable predictions about when stronger normalization helps, and it confronts those predictions with simulations. No parameters are fitted to force agreement between theory and simulations.","major_comments":[],"minor_comments":[{"comment":"The claim that an intermediate limit ρ_n→ρ*∈(0,1] can be absorbed into the latent distribution is not literally covered by Theorem 1 as stated: the dense covariance formula has the factor 1−⟨x,ξ'⟩, the sparse formula has ⟨x,ξ'⟩, and an intermediate limit would produce 1−ρ*⟨x,ξ'⟩. Please either extend the theorem to this case or restrict the claim to the two regimes actually used in the applications.","section":"Section 3, paragraph after Model 1"},{"comment":"The uniform lower bound ⟨x,μ⟩≥c is a genuine structural restriction; it excludes latent positions with ⟨x,μ⟩=0, such as a support containing 0, even when the support spans R^r. A remark connecting this condition to the DCSBM requirements (Θ bounded away from zero and γ_k>0) would help readers assess when the theorem applies.","section":"Condition 1(c)"},{"comment":"The statement that Cauchy–Schwarz gives 0≤r_α≤1 is correct but not immediate; including the log-convexity argument (EΘ^{2−2α})^2≤EΘ·EΘ^{3−4α} would improve verifiability.","section":"Section 4.2"},{"comment":"The captions describe shading, but the figures also need explicit legends or marker descriptions for grayscale/print reproduction, particularly to identify which curve corresponds to which α in the finite-sample panels.","section":"Figures 2 and 3"},{"comment":"The Ali and Couillet (2018) entry is listed as JMLR 18:1–49, 2018; please verify the volume and year, since JMLR volume 18 appeared in 2017.","section":"References"}],"recommendation":"minor_revision","confidential_remarks":"I found no citation or novelty concerns. The paper is within the journal's scope; the local issues listed in the report can be addressed without re-review of the central argument."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper gives a row-wise CLT for the full degree-alpha family D^{-alpha} A D^{-alpha} under the RDPG model, with explicit population centers and covariances, specializations to the DCSBM, and a projected-Gaussian Bayes-error comparison across alpha. The proof is detailed and internally consistent; I checked the key algebra in the final CLT step and the block-model specialization, and the exchangeability argument controlling the spectral remainder is clean. The authors also deserve credit for being upfront about the finite-sample discrepancy between the asymptotic alpha* and the empirical minimizer, and for not overclaiming the practical reach of the diagnostic.\n\nThe genuinely new piece is the explicit covariance structure under RDPG for the whole continuum, which lets them decompose the effect of alpha into geometry reweighting and uncertainty shape. The DCSBM corollaries, including the flattening of the covariance ellipses toward the population hyperplane, are informative and match simulations. The Chernoff-information benchmark in the appendix is a nice addition that connects to Cape et al.\n\nSoft spots are real but proportionate. Condition 1(c) -- a uniform lower bound on <x, mu> -- is the most fragile structural assumption. As the stress-test note says, it is essentially automatic under Condition 1(a) when the support lies in a closed cone with nonnegative pairwise dot products, but it would fail for degree-correction parameters approaching zero. That is a genuine limitation for some DCSBM regimes, though the paper acknowledges the model restriction. The comparison criterion itself is idealized twice over: Gaussian approximation treated as exact and Bayes rule replacing k-means. The authors are honest about this, and the broad ordering still shows up in finite samples, so it does not sink the paper.\n\nThe RDPG/PSD restriction excludes disassortative and indefinite block structures, but that is a known boundary and the paper says so. I did not find a load-bearing flaw. The citation pattern is fair: the debt to Ali and Couillet, Fan et al., and Tang and Priebe is clearly stated, and the self-citations are to work that this paper builds on.\n\nWho is this for? Spectral clustering researchers who want to know whether to use the adjacency matrix, the symmetric Laplacian, or something in between, and theorists who want explicit covariance formulas for this family. It deserves serious refereeing; I would send it out with minor-revision expectations, mainly on sharpening the discussion of Condition 1(c) and the limits of the Bayes-error diagnostic.","headline":"A solid, careful CLT for the degree-alpha Laplacian family under RDPG, with explicit covariances and an honest practical comparison; the central claim holds up.","tokens_in":37204,"tokens_out":1212,"would_cite":true,"duration_ms":858442,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H30"],"pacs":[],"model":"deepseek-v4-flash","headline":"The degree-α Laplacian family of spectral embeddings has explicit Gaussian limits whose covariance explains when degree normalization helps community detection.","keywords":["central limit theorem","degree-corrected stochastic block model","spectral clustering","random dot product graph","degree normalization","community detection","row normalization","Bayes error"],"falsifier":"Simulate an RDPG with latent positions drawn from a distribution satisfying Condition 1, fix $\\alpha=0.75$, and for one node with known $\\xi_i$ compute the scaled residual $n^{\\alpha+1/2}\\rho_n^\\alpha[(\\hat X_\\alpha Q_n)_{i*}-(X_\\alpha)_{i*}]^T$ across many replicates, then test its empirical distribution against $N(0,\\Sigma_\\alpha(\\xi_i))$; a systematic mismatch would show the covariance formula or the CLT is wrong.","tokens_in":36310,"feed_emoji":"📊","tokens_out":6938,"duration_ms":64322,"temperature":0.7,"pith_summary":"This paper proves a central limit theorem for spectral embeddings built from the degree-α Laplacian family $D^{-\\alpha}AD^{-\\alpha}$, which interpolates between the adjacency matrix ($\\alpha=0$) and the symmetric Laplacian ($\\alpha=1/2$). Under a random dot product graph model, the rows of the embedding, after orthogonal alignment, are asymptotically Gaussian with an explicit covariance that depends on the latent position of the row and on $\\alpha$. The result lets the authors compare degree normalizations for community detection: after row-normalization, the population community directions are identical across $\\alpha$, but the surrounding Gaussian fluctuations are not. In two-community stochastic block models, a Bayes-error diagnostic shows that no single $\\alpha$ wins uniformly; stronger normalization is favored in sparser or more imbalanced networks. The paper thereby offers a distributional explanation for when and why degree normalization improves spectral clustering.","feed_headline":"Spectral embeddings converge to Gaussian limits at every degree power","feed_subtitle":"The covariance formula shows why stronger normalization helps in sparser, more imbalanced networks.","key_machinery":"The object is the degree-α Laplacian family $\\hat L_\\alpha=D^{-\\alpha}AD^{-\\alpha}$ for $\\alpha\\in[0,1]$, with population analogue $L_\\alpha=T^{-\\alpha}PT^{-\\alpha}$; the embedding is $\\hat X_\\alpha=\\hat U_\\alpha\\hat\\Lambda_\\alpha^{1/2}$ from the rank-$R$ eigendecomposition. The proof expands the perturbation $\\hat L_\\alpha-L_\\alpha$ via a Taylor expansion of $D^{-\\alpha}$ around $T^{-\\alpha}$, identifies the leading row-wise terms as a sum of independent mean-zero vectors, and applies the Lindeberg–Feller central limit theorem; the main technical difficulty is that $\\Lambda_\\alpha^{-1/2}$ has norm $O(\\delta_n^{\\alpha-1/2})$, which diverges when $\\alpha>1/2$ and requires sharper remainder control. The covariance formula $\\Sigma_\\alpha(x)$ is the load-bearing output: it encodes how degree normalization reweights both the population latent positions (through $T^{-\\alpha}$) and the local noise (through $\\Gamma_{\\rho,\\alpha}$), and the projected-Gaussian Bayes-error diagnostic built on it becomes the tool for comparing normalizations in community detection.","core_discovery":"The central claim is Theorem 1: fix $\\alpha\\in[0,1]$ and assume Condition 1; under the random dot product graph model, there exist orthogonal matrices $Q_n$ such that for each fixed row $i$, conditional on the latent position $\\xi_i$, the scaled embedding residual $n^{\\alpha+1/2}\\rho_n^\\alpha[(\\hat X_\\alpha Q_n)_{i*}-(X_\\alpha)_{i*}]^T$ converges in distribution to a mean-zero Gaussian with covariance $\\Sigma_\\alpha(\\xi_i)=\\Upsilon_\\alpha^{-1}\\Gamma_{\\rho,\\alpha}(\\xi_i)\\Upsilon_\\alpha^{-1}/\\langle\\xi_i,\\mu\\rangle^{2\\alpha}$, where $\\mu$ is the mean latent position and $\\Upsilon_\\alpha=\\mathbb{E}[\\xi\\xi^T/\\langle\\xi,\\mu\\rangle^{2\\alpha}]$. The covariance is explicit and separates into a population-geometric part (depending on the distribution of latent positions) and a fluctuation part $\\Gamma_{\\rho,\\alpha}$ that changes between the dense and sparse regimes. Specializing to the degree-corrected stochastic block model, the paper shows that spherical row normalization makes the limiting population centers identical across $\\alpha$, while the limiting covariance remains $\\alpha$-dependent and can flatten toward the hyperplane containing those centers. The authors then use the projected Gaussian approximations to compute Bayes errors and find that the preferred $\\alpha$ shifts with network density, community imbalance, and block-probability structure.","pith_inferences":["The explicit covariance formula suggests a principled way to estimate an optimal $\\alpha$ from a single network by plugging in estimated latent positions or block parameters, although the paper itself only tunes $\\alpha$ through a subsampling procedure in a companion work.","The Bayes-error diagnostic treats the asymptotic Gaussian approximation as exact at finite $n$; one testable extension is to replace the limiting covariance by the finite-sample covariance from the proof, which would reveal whether the recommended $\\alpha$ shifts when higher-order terms are included.","The restriction to positive-semidefinite block matrices, noted by the authors, means disassortative structures are excluded; extending the covariance formula to the generalized RDPG would likely introduce additional terms from an indefinite population matrix, and it is plausible that the qualitative ordering across $\\alpha$ would change."],"forward_implications":["For any fixed $\\alpha$, the degree-α spectral embedding is comparable to the population target $X_\\alpha=T^{-\\alpha}X$ up to rotation, with Gaussian row-wise errors of scale $n^{-\\alpha-1/2}\\rho_n^{-\\alpha}$.","After spherical row normalization in a degree-corrected stochastic block model, the population community directions are identical for all $\\alpha$, so any performance difference across $\\alpha$ comes from second-order covariance differences, not first-order geometry.","In the balanced symmetric two-community SBM, the variance ratio $F_\\alpha=\\sigma^2_{N,\\alpha}/\\sigma^2_{T,\\alpha}$ is strictly decreasing in $\\alpha$ and reaches zero at $\\alpha=1$, meaning stronger normalization flattens the covariance toward the hyperplane of the population centers.","The Bayes-error diagnostic and finite-sample simulations show that stronger normalization is preferred in sparser or more imbalanced networks, while adjacency spectral clustering ($\\alpha=0$) is preferred when block probabilities are large and communities are balanced."],"supporting_citations":[{"why":"Establishes the row-wise CLT for adjacency spectral embedding under RDPG, the template this paper extends to the degree-α family.","marker":"Athreya et al. (2016)"},{"why":"Proves limit theorems for the normalized Laplacian embedding, covering the $\\alpha=1/2$ case and supplying the proof strategy for the spectral expansion.","marker":"Tang and Priebe (2018)"},{"why":"Provides asymptotic normality for a broad class of generalized Laplacians including the degree-α family, serving as the general theory against which the explicit RDPG covariance is specialized.","marker":"Fan et al. (2026)"},{"why":"First studied the degree-power family in the statistical community-detection literature, with a detectability criterion that motivates the comparison of normalizations.","marker":"Ali and Couillet (2018)"},{"why":"Uses Chernoff information to compare adjacency and symmetric Laplacian embeddings, the benchmark that the projected-Gaussian Bayes-error diagnostic extends across the continuous family.","marker":"Cape et al. (2019a)"},{"why":"Defines the degree-corrected stochastic block model, the class of latent-position distributions on which the population-geometry and covariance specialization is carried out.","marker":"Karrer and Newman (2011)"}],"fun_headline_variants":["Gaussian limits proven for all degree-normalized spectral embeddings","Why stronger degree normalization wins in sparse networks","Explicit covariance for spectral embeddings at every alpha","No universal best normalization for spectral clustering","Degree power trade-off: sparsity favors stronger normalization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is Condition 1(c): every latent position must have a dot product with the mean latent position bounded uniformly away from zero, so no node can be nearly orthogonal to the typical direction; otherwise the inverse degree normalization is unstable and the Gaussian limit may break down.","fun_headline_variants_meta":{"raw":{"variants":["Gaussian limits proven for all degree-normalized spectral embeddings","Why stronger degree normalization wins in sparse networks","Explicit covariance for spectral embeddings at every alpha","No universal best normalization for spectral clustering","Degree power trade-off: sparsity favors stronger normalization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000646,"raw_usage":{"total_tokens":2998,"prompt_tokens":1002,"completion_tokens":1996,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":1925}},"tokens_in":618,"tokens_out":1996,"duration_ms":13393,"temperature":1.0,"reasoning_tokens":1925,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:56:49.112761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate an RDPG with latent positions drawn from a distribution satisfying Condition 1, fix $\\alpha=0.75$, and for one node with known $\\xi_i$ compute the scaled residual $n^{\\alpha+1/2}\\rho_n^\\alpha[(\\hat X_\\alpha Q_n)_{i*}-(X_\\alpha)_{i*}]^T$ across many replicates, then test its empirical distribution against $N(0,\\Sigma_\\alpha(\\xi_i))$; a systematic mismatch would show the covariance formula or the CLT is wrong.","supporting_citations":[{"cited_title":"Asymptotic Theory of Eigenvectors for Latent Embeddings With Generalized","cited_arxiv_id":null,"evidence_quote":"Provides asymptotic normality for a broad class of generalized Laplacians including the degree-α family, serving as the general theory against which the explicit RDPG covariance is specialized."}],"review_version":1}