{"id":"ddac37ae-a7d6-4f93-818b-218e35b979c7","arxiv_id":"2501.01351","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Derives an explicit Gaussian central limit theorem for the giant component in a supercritical finite-type stochastic block model, via the excursion representation of the breadth-first walk.","lead":"This paper proves a central limit theorem for the size of the giant connected component in a supercritical stochastic block model. It gives an explicit Gaussian fluctuation formula, extending the classical Erdős-Rényi result to networks with several vertex types.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central proof is gated by the unproved positivity reduction in Section 3.1: Assumption 1.1 permits zero entries of K, while the imported excursion representation of [6] requires κ_{i,i}>0, and the one-sentence Janson perturbation is not shown to preserve the n^{1/2}-fluctuation limit.","rationale":"I agree partially with the reader. The invalid Perron-Frobenius inference in Lemma 3.5 is a real local error, but that lemma is not used in the CLT proof: Section 3.6 only needs R_n→t0, which follows from Proposition 1.2 and Lemma 3.6. The title/vector mismatch is a scope issue rather than a correctness issue. The true load-bearing point is the bridge from Assumption 1.1 to the positive-diagonal excursion representation. If the perturbation is chosen at the correct o(n^{-1/2}) scale, the total variation distance tends to 0 and the limit is continuous, so the gap is likely repairable; but the paper currently asserts rather than proves it. I also checked the scaling in Corollary 3.2: after the change of variable, the exponential rate Exp(v^T C_n(l)) is correct, and the jump sizes n^{-1}K C_n(l) follow from Theorem 3.1. The main proof is otherwise coherent. The recommended verdict remains conditional acceptance, pending a written justification of the positivity reduction.","tokens_in":9654,"tokens_out":37175,"duration_ms":355326,"concrete_test":"Take the allowed singular two-type case K=[[0,a],[b,0]] (or K with zero diagonal) with μ such that λ1(KM)>1. Set K_ε=K+ε I with ε=n^{-2/3} and compute the Hellinger distance between the product Bernoulli laws of SBM_n(K) and SBM_n(K_ε); verify it tends to 0. Then, for the same λ and β, compute the right-hand side of Theorem 1.3 with K_ε in place of K and take ε→0; verify it converges to the stated expression for K. If either check fails, the WLOG positivity reduction is invalid and the present proof does not cover the full Assumption 1.1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Assumption 1.1 only requires K=(κ_{i,j}) nonnegative, symmetric, irreducible with λ1(KM)>1; individual entries, including diagonal entries, may be zero. The proof mechanism, however, is the excursion representation of [6] (Theorem 3.1 and Corollary 3.2), which is stated only for models with strictly positive diagonal κ_{i,i}. Section 3.1 disposes of this by writing: 'we will assume that κ_{i,j}^{(n)} > 0. This can be done without loss of generality by considering a perturbation sufficiently small and using the general equivalence of Janson [11].' No perturbation rate or continuity argument is supplied. This matters because Janson-type equivalence for sparse graphs is scale-sensitive: adding a fixed δ>0 to κ changes p by δ/n, and the Hellinger affinity over n^2 pairs behaves like exp(-c δ^2 n), so the laws are not asymptotically equivalent. A valid perturbation must be o(n^{-1/2}) in κ, and then one must check that the limiting expression J^{-1}Kζ - J^{-1}(KB+ΛM)ρ - ΛMρ is continuous as the perturbation vanishes. Without that check, the theorem is proved only for positive-diagonal K, a strict subset of the stated Assumption 1.1. Since every subsequent lemma uses the positivity to define the X/Z processes, this is the most load-bearing gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proves a central limit theorem for the rescaled vector of class counts in the largest connected component of a finite-type stochastic block model, allowing first-order perturbations in the class sizes and in the connection probability matrix. Under an irreducibility and supercriticality condition, the limit is Gaussian with an explicit covariance formula. The proof uses the excursion representation of [6] for a degree-corrected SBM and follows the approach of [7], reducing the fluctuations to those of empirical distribution functions of exponential variables and a delta-method step.","tokens_in":9948,"tokens_out":31051,"duration_ms":251005,"significance":"If the result is correct, it is a valuable contribution: it gives an explicit Gaussian fluctuation formula for the component size vector in a finite-type SBM with perturbations, recovers Stepanov's CLT for Erdős-Rényi graphs, and complements the functional CLT of Bhamidi et al. [2] with a concise single-time covariance. The proof strategy is elegant and the analytic steps (Donsker, delta method) are largely transparent. The main obstacle is the treatment of zero entries in the connection matrix K, which the current manuscript does not adequately resolve.","major_comments":[{"comment":"The reduction to κ^{(n)}_{i,j} > 0 is not justified. The sentence 'This can be done without loss of generality by considering a perturbation sufficiently small and using the general equivalence of Janson [11]' is insufficient: for sparse graphs, adding a fixed δ>0 to κ changes the edge probabilities by δ/n, and the Hellinger affinity over n^2 pairs is exp(-c δ^2 n), so the laws are not asymptotically equivalent. A perturbation that preserves the n^{1/2}-fluctuation limit must be o(n^{-1/2}) in κ, and one must then prove both that the excursion representation of [6] applies uniformly in n to the perturbed model and that the limiting expression in Theorem 1.3 is continuous as the perturbation vanishes. Without this, Theorem 1.3 is proved only for matrices K with strictly positive entries, a strict subset of Assumption 1.1. This point is load-bearing because the processes X^{(n)} and Z^{(n)} and Corollary 3.2 require κ_{i,i}>0.","section":"Section 3.1"}],"minor_comments":[{"comment":"The Perron-Frobenius step is invalid: the inequality KM v_y ≤ λ_1 v_y is derived only up to an O(||v_y||^2) term that is not controlled coordinate-wise, and Perron-Frobenius theory does not imply proportionality from such an approximate inequality. The lemma is not used in Sections 3.5-3.6, so it should be corrected or removed.","section":"Lemma 3.5"},{"comment":"The first displayed equation has an erroneous square: the factor should be 1/(1-c(1-ρ)) without the square, otherwise the variance does not match σ^2(c). The second displayed equation writes 'KJ^{-1}' where 'K^{-1}J^{-1}' is meant. These typos obscure the verification of the Erdős-Rényi reduction.","section":"Section 1.1.2"},{"comment":"The text says 'The second inequality follows from t = n^{-1/3} diag(K)s', but the displayed transformation is an equality; the word 'inequality' should be 'equality'.","section":"Section 3.2"},{"comment":"There are several typographical errors: 'proof for of' in the abstract should be 'proof of', 'stocahstic' should be 'stochastic', and 'Perron-Frobinous'/'eigevnector' should be 'Perron-Frobenius'/'eigenvector'.","section":"Abstract and Section 3.1"},{"comment":"The convergence n^{1/2} Z^{(n)}(R_n) ⇒ 0 is said to follow 'immediately'; the proof should spell out that the jump of Z^{(n)} over the giant component cancels exactly in the limit, so that Z^{(n)}(R_n) ≈ Z^{(n)}(L_n-). This would make the argument easier to verify.","section":"Proof of Lemma 3.6"}],"recommendation":"major_revision","confidential_remarks":"The paper depends on two recent preprints by the same group ([6], [7]); if these are not yet published, the editor may wish to verify their status. The main gap, the positivity reduction in Section 3.1, is substantial but likely fixable with a careful perturbation argument. The recovery of Stepanov's result in Section 1.1.2 contains typos that should be corrected. Overall, the central idea is sound and the result, once repaired, would be a nice contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. The main result is a CLT for the K-weighted giant component vector in a finite-type SBM, with an explicit Gaussian covariance that includes first-order perturbations in class sizes and connection probabilities. That is genuinely new: Neal's work covers the unperturbed case, and Bhamidi-Budhiraja-Sakanaveeti give a functional CLT with the covariance buried in an SDE system. The proof strategy is clean, the rescaling computations check out, and the d=1 reduction to Stepanov's theorem is correct. Lemma 3.7 is a nice elementary estimate.\n\nTwo soft spots keep it from being finished. First, Section 3.1's WLOG reduction to strictly positive κ_{i,i} is unjustified as written. The excursion representation of [6] requires positive diagonal entries, but Assumption 1.1 allows zeros. The one-line appeal to Janson's asymptotic equivalence doesn't work at this scale: a fixed δ>0 perturbation changes edge probabilities by δ/n, and the Hellinger affinity over n^2 pairs decays like exp(-c δ^2 n), so the laws are not asymptotically equivalent. You would need a perturbation of order o(n^{-1/2}) and then a continuity argument for the limiting expression. Neither appears. This is load-bearing because every subsequent lemma uses positivity to define the processes.\n\nSecond, the proof of Lemma 3.5 makes an invalid Perron-Frobenius inference. The coordinate-wise inequality KM v_y ≤ λ_1 v_y does not imply v_y is a scalar multiple of the Perron eigenvector a; there can be slack in some coordinates. The lemma may be true, but the proof doesn't show it. There is also a minor mismatch between the title/abstract, which promise a CLT for the number of vertices in the giant, and the theorem, which states the result for K C_n^(1). For d>1 the raw vertex count is not directly covered.\n\nI don't think any of this sinks the underlying claim; the result is very likely true and the method is right. But the paper as it stands proves the theorem only for positive diagonal K and even there has a defective hitting-time lemma. It deserves a serious referee who can ask for a proper treatment of the positivity reduction and a corrected (or cited) proof of Lemma 3.5. Send it out, with those conditions.","headline":"A credible and mostly explicit CLT for the K-weighted giant in a finite-type SBM, but the proof as written has a load-bearing gap in the positivity reduction and a wrong argument in Lemma 3.5.","tokens_in":10501,"tokens_out":5922,"would_cite":false,"duration_ms":55260,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60F05","05C80"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that the vector of giant-component counts in a finite-type stochastic block model has Gaussian fluctuations at the $n^{-1/2}$ scale, with an explicit covariance formula.","keywords":["stochastic block model","giant component","central limit theorem","breadth-first walk","excursion representation","delta method","random graphs","supercritical phase"],"falsifier":"Simulate a two-type stochastic block model with $n=10^5$, fixed $\\mu$, supercritical $K$, nonzero $\\beta$ and $\\Lambda$, and estimate the covariance matrix of $n^{1/2}(n^{-1} K C_n^{(1)} - K M \\rho)$ across many repetitions; if the estimated covariance disagrees with the theorem's formula beyond sampling error, Theorem 1.3 is false.","tokens_in":9404,"feed_emoji":"🎲","tokens_out":8309,"duration_ms":75268,"temperature":0.7,"pith_summary":"This paper proves a central limit theorem for the vector of class-wise counts of the largest connected component (the giant) in a finite-type stochastic block model. Under mild conditions on class sizes and connection probabilities, it shows that $n^{1/2}(n^{-1} K C_n^{(1)} - K M \\rho)$ converges in distribution to an explicit Gaussian vector. The limit combines intrinsic sampling noise with first-order perturbations coming from class-size imbalances and edge-probability shifts. The proof is short because it encodes the random graph by a breadth-first walk whose largest jump is the rescaled giant, reducing the computation to an invariance principle and the delta method. A reader should care because the result gives a concise, explicit fluctuation formula for the giant in a multi-type random graph, not just a qualitative Gaussian limit.","feed_headline":"Giant component in block models gets Gaussian limit law","feed_subtitle":"Explicit fluctuations at the square-root scale cover class-size and edge-rate perturbations, recovering the classical random-graph CLT.","key_machinery":"The load-bearing object is the breadth-first walk $T(y,v,Z^{(n)})$ built from entrywise Poisson-like processes $Z_{i,j}^{(n)}$; the associated excursion representation identifies the jumps of this walk with $n^{-1} K C_n^{(l)}$, so the largest jump is exactly the rescaled giant. The proof then uses the classical invariance principle for empirical distribution functions to write $Z^{(n)}=\\phi + n^{-1/2}\\Psi + o(n^{-1/2})$, where $\\phi(t)_i=-t_i+\\sum_j \\kappa_{ij}\\mu_j(1-e^{-t_j})$ has a zero set containing the deterministic point $t_0=KM\\rho$. The delta method with Jacobian $J=KM(I-\\mathrm{diag}(\\rho))-I$ converts fluctuations of $\\phi$ at $t_0$ into fluctuations of the walk's endpoint, giving the stated Gaussian limit.","core_discovery":"On the paper's own terms, the central claim is Theorem 1.3: under its Assumption 1.1, if $C_n^{(1)}$ is the vector of class counts of the largest component, $K=(\\kappa_{ij})$, $M=\\mathrm{diag}(\\mu)$, and $\\rho$ is the unique positive fixed point of $1-\\exp(-\\sum_j \\kappa_{ij}\\mu_j\\rho_j)=\\rho_i$, then $n^{1/2}(n^{-1} K C_n^{(1)} - K M \\rho)$ converges in distribution to $J^{-1}K\\zeta - J^{-1}(KB + \\Lambda M)\\rho - \\Lambda M\\rho$, where $J=KM(I-\\mathrm{diag}(\\rho))-I$, $B=\\mathrm{diag}(\\beta)$, $\\Lambda=(\\lambda_{ij})$, and the $\\zeta_j$ are independent centered normals with variance $\\mu_j\\rho_j(1-\\rho_j)$. The parameter $\\beta_j$ describes the $n^{1/2}$ perturbation of class size $n_j$, and $\\lambda_{ij}$ the $n^{1/2}$ perturbation of connection rates. In the one-class case the formula reduces to the classical central limit theorem for the giant in the sparse random graph $G(n,c/n)$, and with a nonzero $\\lambda$ it matches the previously known perturbed single-type result.","pith_inferences":["One can pursue the same excursion representation to derive CLTs for other additive functionals of the giant, such as internal edge counts, by replacing the jump functional $K C_n^{(1)}$ with a different component-level statistic.","The explicit covariance formula could be used to build asymptotic confidence intervals for class sizes in a fitted stochastic block model, an application the paper does not discuss.","The unproved continuity when zero connection rates are perturbed to positive values suggests a numerical check at a boundary where one $\\kappa_{ij}=0$ but the matrix $KM$ remains irreducible; if the limit's covariance changes discontinuously there, the proof's reduction would need repair."],"forward_implications":["When the model has a single class and the perturbation parameters vanish, the theorem's limit reduces to the classical CLT for the giant in $G(n,c/n)$ with variance $\\sigma^2(c)$.","For a finite-type stochastic block model, the theorem gives an explicit Gaussian law for the class-wise counts of the giant, including first-order corrections from class-size imbalances and edge-probability perturbations.","At the $n^{1/2}$ scale, all smaller components are negligible: the largest of them is $O_P(\\log n)$, so only the giant's fluctuations contribute to the limit.","The limit's covariance is controlled by the matrix $J=KM(I-\\mathrm{diag}(\\rho))-I$, and the proof implies this matrix is invertible throughout the supercritical regime.","The proof strategy shows the CLT follows from the invariance principle for empirical distribution functions plus the delta method, avoiding component-specific generating function analysis."],"supporting_citations":[{"why":"Supplies the excursion representation of the breadth-first walk that identifies its jumps with rescaled connected component vectors.","marker":"[6]"},{"why":"Provides the recent random-walk template for proving giant-component CLTs that this paper follows.","marker":"[7]"},{"why":"The classic single-type CLT for the giant that the present theorem recovers when the model has one class.","marker":"[17]"},{"why":"Gives the weak law for the giant and the operator framework $T_K,\\Phi_K$ used to define $\\rho$ and prove invertibility of $J$.","marker":"[3]"},{"why":"Defines the first-hitting-time construction used to build the breadth-first walk and its excursion intervals.","marker":"[5]"},{"why":"Justifies the reduction to strictly positive $\\kappa_{i,j}^{(n)}$ via asymptotic equivalence of random graph models.","marker":"[11]"},{"why":"Provides the perturbed single-type CLT that matches the $d=1$ case with nonzero $\\lambda$.","marker":"[15]"}],"fun_headline_variants":["CLT proves Gaussian size for block-model giant component","Stochastic block model giant follows Gaussian limit theorem","Giant component in SBM: CLT gives Gaussian size","Block-model giant size becomes Gaussian via new proof"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof assumes that the imported excursion representation is valid for this model and that perturbing any zero connection rates to small positive values does not change the limit; the second reduction is asserted rather than proved.","fun_headline_variants_meta":{"raw":{"variants":["CLT proves Gaussian size for block-model giant component","Stochastic block model giant follows Gaussian limit theorem","Giant component in SBM: CLT gives Gaussian size","Block-model giant size becomes Gaussian via new proof"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000603,"raw_usage":{"total_tokens":2795,"prompt_tokens":905,"completion_tokens":1890,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":1827}},"tokens_in":521,"tokens_out":1890,"duration_ms":13605,"temperature":1.0,"reasoning_tokens":1827,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:30:52.645347+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a two-type stochastic block model with $n=10^5$, fixed $\\mu$, supercritical $K$, nonzero $\\beta$ and $\\Lambda$, and estimate the covariance matrix of $n^{1/2}(n^{-1} K C_n^{(1)} - K M \\rho)$ across many repetitions; if the estimated covariance disagrees with the theorem's formula beyond sampling error, Theorem 1.3 is false.","supporting_citations":[{"cited_title":"Degree corrected stochastic block model: excursion representation","cited_arxiv_id":"2409.18894","evidence_quote":"Supplies the excursion representation of the breadth-first walk that identifies its jumps with rescaled connected component vectors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The classic single-type CLT for the giant that the present theorem recovers when the model has one class."},{"cited_title":"Bollob´ as, S","cited_arxiv_id":null,"evidence_quote":"Gives the weak law for the giant and the operator framework $T_K,\\Phi_K$ used to define $\\rho$ and prove invertibility of $J$."},{"cited_title":"Chaumont and M","cited_arxiv_id":null,"evidence_quote":"Defines the first-hitting-time construction used to build the breadth-first walk and its excursion intervals."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies the reduction to strictly positive $\\kappa_{i,j}^{(n)}$ via asymptotic equivalence of random graph models."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the perturbed single-type CLT that matches the $d=1$ case with nonzero $\\lambda$."}],"review_version":1}