{"id":"6d83c117-9f97-4e27-862d-0318e5041d6e","arxiv_id":"2502.10170","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"It adapts multiple comparison with the best to egocentric network randomized trials to identify subgroups of index participants with the largest spillover effects, with power and sample size formulas.","lead":"This paper proposes a statistical test for finding which types of trained peer educators most improve the behavior of their friends and partners in network-based health trials. The method can also help researchers calculate how large such a trial needs to be.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 1's printed variance for \\hat\\delta omits the 1/K factor and uses the wrong ICC inverse denominator, so standard errors, MCB confidence sets, and sample-size formulas as written cannot be correct; a reader implementing from the paper cannot reproduce valid inference.","rationale":"The reader's CONDITIONAL verdict is appropriate. I did not find a reason to reject the central idea: the causal identification in Theorem 1 is standard under Assumptions 1-3, the MCB construction follows Hsu (1996), and the extension to multiple best subgroups is plausible. The reader's weakest_assumption concerned Assumptions 1 and 2, namely non-overlapping egonetworks and neighborhood interference. Those assumptions are indeed load-bearing, but they are stated explicitly and their fragility is acknowledged in Section 8; the methodological claim is conditional on them. My own concern is more immediate and more easily falsifiable: the printed variance formula in Lemma 1 cannot be correct as written because it lacks the 1/K scaling required for a sqrt(K)-consistent estimator, and the inverse compound-symmetric covariance uses 1+n\\rho instead of 1+(n-1)\\rho. Since Lemma 2, Theorem 2, the MCB confidence set in (6), and the sample-size formulas in Section 5 all consume these variances, a reader following the printed formulas cannot produce valid inference even if every causal assumption holds. The simulation table suggests the code may have used a corrected expression, which is why this is a manuscript-correctness issue rather than a fundamental flaw. The paper should be corrected and ideally accompanied by code before the method is used for trial planning, so the reader's CONDITIONAL verdict stands unchanged.","tokens_in":19380,"tokens_out":14600,"duration_ms":159142,"concrete_test":"Re-derive Var(\\hat\\delta_h) for a single subgroup from Equation (4). With V_k^{-1}=cI+dJ, c=1/(1-\\rho), d=-\\rho/((1-\\rho)(1+(n-1)\\rho)), and b=n/(1+(n-1)\\rho), the information matrix for (\\zeta_h,\\delta_h) in subgroup h is b[gK, pgK; pgK, pgK]; inverting gives Var(\\hat\\delta_h)=\\sigma^2/[b p(1-p)g_h K]. Then plug the paper's simulation values (K=5000, n=5, \\rho=0.8, \\sigma^2=5, p=0.5, g=0.25) into both the printed Lemma 1 expression and the corrected expression: the printed formula has no 1/K and therefore gives a standard error about sqrt(K) times too large, while a variant with 1/K but the printed (1+n\\rho) denominator gives a different ICC adjustment. This one-page algebraic check, or a K=100 simulation comparing coverage under the two variance formulas, settles whether Lemma 1 is a harmless typo or a substantive error in the method as published.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The MCB procedure's FWER guarantee (Theorem 2), the critical values in Lemma 2, and the power and sample-size calculations in Section 5 all flow through the standard errors of \\hat\\delta_h. Lemma 1 (Section 3.2) prints Var(\\hat\\delta_h) = \\sigma^2(1-p)/(\\bar b p g_h) with \\bar b = n/((1-\\rho)(1+n\\rho)). This has two concrete problems. First, the expression has no dependence on the number of egonetworks K, but \\hat\\delta is sqrt(K)-consistent, so the asymptotic covariance must scale as 1/K; the printed formula is therefore off by a factor of roughly K. Second, the inverse of the compound-symmetric working covariance V_k = \\sigma^2[(1-\\rho)I + \\rho J] is cI + dJ with c = 1/(1-\\rho) and d = -\\rho/((1-\\rho)(1+(n-1)\\rho)), not the denominator 1+n\\rho used in Lemma 1. The correct cluster-summary quantity is consequently n/(1+(n-1)\\rho), not n/((1-\\rho)(1+n\\rho)). The MCB confidence set in (6) uses \\hat\\sigma\\sqrt{v_{jh}}; if v_{jh} is wrong by a factor of K, the simultaneous intervals, the membership of \\hat B_0, and the sample-size curves in Figures 1-2 are all invalid. This undermines the central claim even when Assumptions 1-3 hold exactly. The simulation table suggests the code may have used a corrected 1/K term, but the manuscript as written does not, and no code is provided.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a confirmatory method for identifying subgroups of index participants whose treatment produces the largest spillover effect on their network members in an Egocentric Network-based Randomized Trial (ENRT). The authors define a subgroup-specific spillover estimand δ(h), identify it under assumptions of non-overlapping egonetworks and neighborhood interference, estimate it via a GEE fit to the linear mixed model in equation (3), and then apply a Multiple Comparisons with the Best (MCB) procedure to construct simultaneous confidence intervals, an overall p-value, and power and sample-size calculations. The proposal is illustrated with a simulation study and with an application to the STEP into Action HIV prevention study.","tokens_in":19741,"tokens_out":17172,"duration_ms":167778,"significance":"The research question is practically important, and the combination of an ENRT design with the MCB framework is a sensible contribution if the inferential machinery is correct. The paper would give applied researchers a multiple-comparisons-adjusted procedure for selecting key influencer subgroups and for planning such trials, and the extension of MCB to a non-equicorrelated covariance structure and to multiple best subgroups is useful. However, the central variance lemma contains algebraic errors as printed, and the simulation results do not appear consistent with the stated data-generating process. No code or supplementary material is provided, so the numerical claims cannot currently be audited. The contribution is therefore conditional on substantial correction and re-validation.","major_comments":[{"comment":"The printed variance formula for \\hat\\delta_h is incorrect and this error propagates into the MCB confidence set, the critical values, and the power calculations. First, under the stated asymptotics (\\sqrt K(\\hat\\theta-\\theta)\\to N(0,\\Sigma)), Var(\\hat\\delta_h) must be O(1/K), but no K appears in the printed expression. Second, the inverse of the compound-symmetric working covariance V_k=\\sigma^2[(1-\\rho)I+\\rho J] is aI+bJ with b=-\\rho/[\\sigma^2(1-\\rho)(1+(n-1)\\rho)], so the cluster-summary constant is n/[1+(n-1)\\rho], not n/[(1-\\rho)(1+n\\rho)]. Third, the treatment assignment probability enters through p(1-p), not through p alone. With balanced networks, the correct expression is Var(\\hat\\delta_h)=\\sigma^2[1+(n-1)\\rho]/[n K p(1-p)g_h]. Because \\hat\\sigma\\sqrt{v_{jh}} is used in the MCB set defined around equation (6), and because Lemma 2 and the power formulas in Section 5 build on v_{jh}, the entire downstream procedure is affected. The authors should restate Lemma 1 and re-derive all standard errors, critical values, and sample-size formulas from the corrected expression.","section":"Section 3.2, Lemma 1"},{"comment":"The simulation results in Table 1 are not consistent with the stated design. For K=5000, n=5, \\sigma^2=5, \\rho=0.8, p=0.5, and g_h=0.25, the corrected variance formula gives \\mathrm{sd}(\\hat\\delta_h)\\approx0.116. The printed Lemma 1 gives \\mathrm{sd}\\approx2 without inserting a 1/K factor and \\mathrm{sd}\\approx0.028 if a 1/K factor is simply inserted. The table reports StdE(eStdE)\\approx0.052 for every \\delta_h, which is roughly a factor of two smaller than the corrected value. This discrepancy suggests that either the data were not generated with the specified cluster-level variance \\sigma_u^2=4, or the standard error calculation omits part of the 1/[K p(1-p)] factor. Because the simulation is the main evidence that the MCB procedure controls coverage at the nominal level, the simulation must be rerun and the reported standard errors, coverage rates, and power values must be reconciled with the corrected formulas. The authors should also make the simulation code available.","section":"Section 6, Table 1"}],"minor_comments":[{"comment":"The indexing is internally inconsistent: the egonetwork is defined with i=1,\\ldots,n_k, model (3) is written for i=2,\\ldots,n_k+1, and network members are later described as i>1. Please harmonize the notation throughout.","section":"Section 2.1"},{"comment":"The theorem defines simultaneous intervals for \\delta_h-\\max_{j\\ne h}\\delta_j, but the second sentence of the statement writes the coverage event as \\delta_h-\\min_{j\\ne h}\\delta_j. This typo should be corrected.","section":"Section 4.2.1, Theorem 2"},{"comment":"The displayed power integral has garbled limits (\"Z\\infty\\infty\" and \"Z u^*0\") and uses r(u) for the density of \\hat\\sigma/\\sigma after Lemma 2 used \\gamma(u). Please rewrite the power formula with consistent notation and correct integration limits.","section":"Section 5.2, equation (9)"},{"comment":"The estimator \\hat\\sigma and its degrees of freedom \\nu are used in the critical value calculation and the p-value formula, but they are never explicitly defined. The authors should state how \\hat\\sigma is computed from the GEE fit and what distribution is assumed for \\nu\\hat\\sigma^2/\\sigma^2.","section":"Section 4.2, Lemma 2"},{"comment":"The final displayed integral for the overall p-value has mismatched parentheses, mixes the variables x and z, and contains an incomplete square-root expression. Please rederive and display this formula cleanly.","section":"Section 4.2.2, p-value formula"},{"comment":"The paper repeatedly refers to supplementary material S1 and S4 for proofs of Theorems 1 and 2 and for additional analyses, but no supplement was included with the manuscript. Please provide the supplementary file or move the proofs into the main text.","section":"Supplementary material"},{"comment":"There are numerous typos, including \"remina\" in the abstract, \"Casual Inference\" in the keywords, \"identifing\" in the Introduction, and \"stead\" in Section 4.2. A careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The Lemma 1 error is serious but clearly fixable, so I recommend major revision rather than rejection. However, the simulation inconsistency in Table 1 is more than a typo: the reported standard errors are about half the magnitude implied by the stated design, which suggests the simulation code may not have implemented the cluster-level random effects or the variance formula as described. I would ask for corrected formulas, a rerun simulation, and deposition of code before considering the paper acceptable. There is also a scope question for the editor: the paper is a fairly standard methodological extension of MCB to a new design, and its contribution would be strengthened by clearer guidance on when the assumptions in Assumptions 1 and 2 are credible in practice."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what's worth knowing: this is a sensible and genuinely new application of Hsu's MCB to identify subgroups of index participants with the largest spillover effects in egocentric network randomized trials. It extends MCB to a general covariance structure and to multiple best subgroups, and provides power and sample-size machinery. Theorem 1's identification of δ(h) is clean and correct under the stated assumptions. If the formulas were right, this would be a useful confirmatory tool for implementation science.\n\nThe problem is that they aren't, as printed. Lemma 1 in Section 3.2 gives Var(δ̂h) = σ²(1−p)/(b̄ p g_h) with b̄ = n/((1−ρ)(1+nρ)). That has no 1/K, so it cannot be the variance of a √K-consistent estimator; and the inverse of the compound-symmetric working covariance leads to 1+(n−1)ρ in the denominator, not 1+nρ. The standard errors in Table 1 (0.052) match the corrected formula with 1/K, which means the simulation code used something different from what the paper prints. Since the MCB confidence set in (6), the critical values in Lemma 2, and the sample-size curves in Figures 1–2 all flow through these variances, the manuscript as written is not implementable. No code or data are provided, so a reader cannot reconstruct the intended calculation. That is a load-bearing flaw, not a typo in a peripheral display.\n\nThe application also overstates the evidence. Table 3 shows intervals for δh − min_j δj with lower bounds 0 for subgroups 1 and 2, and negative for 3–5; none of the lower bounds is strictly positive, so no subgroup is actually 'significantly inferior' to the best. The text claims subgroups 1 and 2 are significantly inferior, which contradicts its own table.\n\nWhat's good: the causal framing is careful, the extension to multiple best subgroups is a real gap in the MCB literature, and the simulation design is aligned with the corrected formulas. The self-citations to prior work by the same group are appropriate given the direct lineage of the ENRT estimands.\n\nBottom line: this deserves a serious referee and major revision. I would not cite it in its current form, and I would want to see the Lemma fixed, code posted, and the application re-read before using it for trial planning.","headline":"Useful MCB extension for egocentric network trials, but Lemma 1's printed variance is wrong and the application misreads its own Table 3; fix before use.","tokens_in":20266,"tokens_out":3270,"would_cite":false,"duration_ms":31018,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H15","62J12","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a Multiple Comparison with Best procedure applied to GEE estimates can identify the subgroup of index participants with the largest spillover effect in egocentric network randomized trials while controlling…","keywords":["causal inference","interference","spillover effects","multiple comparisons with the best","egocentric network randomized trial","key influencers","generalized estimating equations"],"falsifier":"Simulate an ENRT in which a small fraction of network members are linked to two index participants and outcomes depend on both indices' treatments; if MCB simultaneous coverage falls measurably below $1-\\alpha$ or the estimated $\\delta(h)$ shows bias growing with that fraction, the identifying assumptions are load-bearing and the central claim fails.","tokens_in":19176,"feed_emoji":"🎯","tokens_out":8945,"duration_ms":78321,"temperature":0.7,"pith_summary":"The paper aims to give researchers a confirmatory test for finding which kinds of people, when trained in a peer-education trial, most change the behavior of their social network members. In an egocentric network randomized trial, only index participants are randomized; their nominated network members stay untreated, so the paper defines the spillover effect of subgroup $h$ as the average difference in network-member outcomes between treated and untreated indices in that subgroup. The paper claims this effect is identified by a simple mean contrast (Theorem 1), estimable by GEE from a linear mixed model clustering by egonetwork, and that Multiple Comparison with Best (MCB) then produces simultaneous confidence intervals that single out the best subgroup(s) while controlling the family-wise error rate (Theorem 2). It also derives power and sample-size formulas so future ENRTs can be sized to detect key influencers, and illustrates the procedure on a peer HIV-prevention trial.","feed_headline":"Test singles out the most influential peer subgroups","feed_subtitle":"MCB intervals flag which peer educators' training changes network members' outcomes.","key_machinery":"The load-bearing object is the heterogeneous spillover effect $\\delta(h)$, the average effect of a treated index participant of subgroup $h$ on an untreated network member, identified as $E[Y_{ik}|Z_{1k}=1, X_{1k}=h, R_{ik}=0] - E[Y_{ik}|Z_{1k}=0, X_{1k}=h, R_{ik}=0]$ under non-overlapping egonetworks and neighborhood interference. This contrast is estimated through the linear mixed model $Y_{ik} = \\sum_{h=1}^H \\zeta_h S_{kh} + \\sum_{h=1}^H \\delta_h G_{ik} S_{kh} + u_k + \\epsilon_{ik}$, with GEE and a working covariance that accounts for within-egonetwork correlation. The MCB machinery then builds simultaneous confidence intervals using subgroup-specific critical values $c_\\alpha^h$ computed from a double-integral identity, and the power formula combines interval coverage with interval narrowness to size the trial.","core_discovery":"The paper's central claim is that MCB, applied to the GEE estimator from model (3), identifies the subgroup(s) of index participants with the largest spillover effect on their network members while controlling the family-wise error rate. For each subgroup $h$, MCB tests whether $\\delta_h$ is at least as large as the best of the other subgroups and builds simultaneous confidence intervals for $\\delta_h - \\max_{j \\neq h} \\delta_j$. Theorem 2 states that, as the number of egonetworks grows, these intervals cover all true differences with probability at least $1-\\alpha$, and exactly $1-\\alpha$ when the best subgroup is unique. The paper further claims that its power definition and sample-size calculations extend MCB to multiple best subgroups, and that in the STEP into Action HIV-prevention trial the method identifies the mid-age and college-educated subgroups as the key influencers.","pith_inferences":["Editorial extension: the same MCB-on-GEE template could be carried to binary or count network-member outcomes through generalized estimating equations, with the critical-value computation updated accordingly.","Editorial extension: because subgroups are fixed before analysis from baseline covariates, the procedure is confirmatory and will not discover influencer types that were not pre-specified.","Editorial extension: a natural stress test would re-analyze data under a growing fraction of network members connected to more than one index participant; the coverage guarantee should degrade smoothly as that fraction grows, revealing how much validity depends on Assumption 1."],"forward_implications":["An ENRT can be pre-sized with the provided power formulas to have a chosen probability of detecting a specified difference between the best and second-best subgroups.","MCB outputs a confidence set of subgroups statistically indistinguishable from the best, so implementers can target a defensible set of peer educators rather than relying on a single point estimate.","Under the overall null of no heterogeneity, the probability that all subgroups enter the best set is at least $1-\\alpha$, so false claims of a key influencer are controlled.","In the STEP into Action application, the method indicates that older and college-educated index participants have the largest beneficial spillover effects on HIV risk behavior, which would guide peer-educator selection.","Compared with the Wald heterogeneity test, MCB requires more egonetworks for the same power, but it answers the targeting question the Wald test leaves open."],"supporting_citations":[{"why":"Supplies the MCB framework and the computationally efficient critical-value computation that the paper extends to a general covariance structure.","marker":"Hsu (1996)"},{"why":"Gives the constrained simultaneous confidence intervals that become Equation (6).","marker":"Hsu (1984)"},{"why":"Establishes the potential-outcome framework and the consistency assumption used to define and identify spillover effects.","marker":"Rubin (1974, 2005)"},{"why":"Provides the ENRT spillover setting and causal estimands that the paper builds on.","marker":"Buchanan et al. (2018)"},{"why":"Supplies the neighborhood-interference assumption restricting outcomes to depend on direct neighbors' treatments.","marker":"Forastiere et al. (2021, 2022)"},{"why":"Gives the GEE asymptotic normality conditions used for the estimator and the simultaneous intervals.","marker":"Zeger et al. (1988)"},{"why":"Provides the STEP into Action trial data used in the illustrative application.","marker":"Tobin et al. (2011)"}],"fun_headline_variants":["MCB test identifies peer subgroups with largest spillover","Pinpoint influential peers with MCB intervals","New method finds key influencers in peer networks","Which peer educators change network outcomes? MCB answers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes each network member is connected to exactly one index participant and that outcomes are affected only by a treated direct neighbor; if either fails, the estimated subgroup contrast is no longer a spillover effect from subgroup $h$.","fun_headline_variants_meta":{"raw":{"variants":["MCB test identifies peer subgroups with largest spillover","Pinpoint influential peers with MCB intervals","New method finds key influencers in peer networks","Which peer educators change network outcomes? MCB answers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000246,"raw_usage":{"total_tokens":1548,"prompt_tokens":962,"completion_tokens":586,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":527}},"tokens_in":578,"tokens_out":586,"duration_ms":5809,"temperature":1.0,"reasoning_tokens":527,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T19:08:24.223801+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate an ENRT in which a small fraction of network members are linked to two index participants and outcomes depend on both indices' treatments; if MCB simultaneous coverage falls measurably below $1-\\alpha$ or the estimated $\\delta(h)$ shows bias growing with that fraction, the identifying assumptions are load-bearing and the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the MCB framework and the computationally efficient critical-value computation that the paper extends to a general covariance structure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the constrained simultaneous confidence intervals that become Equation (6)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the potential-outcome framework and the consistency assumption used to define and identify spillover effects."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ENRT spillover setting and causal estimands that the paper builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the neighborhood-interference assumption restricting outcomes to depend on direct neighbors' treatments."},{"cited_title":"L., K.-Y","cited_arxiv_id":null,"evidence_quote":"Gives the GEE asymptotic normality conditions used for the estimator and the simultaneous intervals."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the STEP into Action trial data used in the illustrative application."}],"review_version":1}