{"id":"b13e30e2-298d-437f-95c1-83e0b6513a22","arxiv_id":"2411.17905","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A hybrid sampling design that reuses clusters but not individuals across time can substantially lower the variance of difference-in-differences estimators relative to repeated cross-sectional sampling.","lead":"Repeatedly sampling the same clusters (for example, the same villages or hospitals) at each survey wave, but different people within them, can sharply reduce the variance of before-after treatment effect estimators compared with drawing entirely new clusters each wave.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the variance reduction is real under Eq. (3), and the paper explicitly discloses the time-invariant cluster-effect scope condition.","rationale":"The paper's aim is to show that sampling the same clusters but different individuals reduces variance relative to re-sampling clusters. The analytic derivation in Appendix B is a straightforward variance calculation under the stated linear mixed model; I checked the algebra: for RCS, each of the four group-time means contributes 2(mτ²+σ²)/n and the four are independent, giving 8(mτ²+σ²)/n; for DISC, the within-cluster difference removes α_i, leaving only ε variance, giving 8σ²/n. The ratio and ICC reparameterization are correct. The simulation results are consistent with the theory in Figure 3, and Figure 4 shows the DRDID point-estimate variance also improves, which is a useful robustness check. The main limitation is that the variance cancellation requires α_i to be constant over time and the same clusters to be available at both waves. The paper explicitly acknowledges this in the Discussion, so the conclusion is conditional rather than overclaimed. The reader's weakest-assumption identification matches this boundary condition, but because it is disclosed and the analytic claim is explicitly derived from the model, I do not see a reason to change the accept verdict. No formal verification or independent code was provided, but the parameter-free analytic result and reproducible simulation structure are sufficient for this low-risk correctness assessment.","tokens_in":9358,"tokens_out":15289,"duration_ms":149699,"concrete_test":"Simulate from Y_ijk = μ_j + δ X_ij + α_i + γ_it + ε_ijk with α_i ~ N(0,τ²), γ_i1,γ_i2 iid N(0,τ_γ²), and ε ~ N(0,σ²); compute the empirical variance of Eq. (2) under RCS and DISC for τ_γ²/τ² = 0, 0.1, 0.5, 1. If the variance ratio is 1+mρ/(1-ρ) only at τ_γ²=0 and falls toward 1 as τ_γ² grows, the time-invariance condition is exactly the boundary of the claimed advantage, confirming that the paper's scoping is accurate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central variance comparison is internally correct. Under Eq. (3) with a single time-invariant cluster effect α_i, the Appendix B derivation gives Var(δ_RCS)=8(mτ²+σ²)/n and Var(δ_DISC)=8σ²/n because the same-cluster difference cancels α_i. The ratio 1+mρ/(1-ρ) follows algebraically. The only way this advantage collapses is if cluster effects are not constant across time or if the cluster population changes between waves. If α_i is replaced by α_i + γ_it with Var(γ_i2 - γ_i1)=v², the DISC variance gains a term 8m v²/n, and the advantage over RCS shrinks toward 1 as v² grows relative to τ². The paper explicitly states this limitation in the Discussion ('changes over the study period in terms of the total number of clusters or the sizes of the clusters'), so the claim is appropriately scoped. Minor reporting inconsistencies (simulation text states ICC=0.2 while figure captions state ICC=0.1) do not affect the analytic result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript proposes the DISC design, in which the same clusters are sampled at each wave but different individuals are drawn within clusters, and studies it in a two-period difference-in-differences setting. Under the linear mixed model in Eq. (3), the authors derive Var(δ̂_RCS)=8(mτ²+σ²)/n and Var(δ̂_DISC)=8σ²/n, yielding a variance ratio of 1 + mρ/(1−ρ). A simulation study with 1,000 replicates confirms the variance formulas for the simple estimator and shows similar gains for a doubly-robust DID estimator that adjusts for covariates. The paper also discusses extensions to other longitudinal designs. The mathematical derivation is correct under the stated assumptions, and the simulation results support the theory, although the simulation text and figure captions disagree on the ICC, and an illustration promised in the abstract is absent from the body.","tokens_in":9565,"tokens_out":8806,"duration_ms":83783,"significance":"If the result holds, the DISC design provides a practically meaningful precision improvement for cluster-sampled longitudinal and DID studies without requiring follow-up of the same individuals, and the analytic variance ratio is simple enough to be used in power calculations. The derivation is explicit and self-contained: the ICC enters only as a reparameterization of τ² and σ², and the simulation checks the formulas rather than fitting them. The authors also disclose the design's limitation when cluster populations change and connect the work to existing practice in cluster randomized trials. The contribution is incremental but useful for applied researchers planning cluster-based longitudinal studies.","major_comments":[{"comment":"The simulation text states 'THe value of τ was chosen to yield an ICC of 0.2,' but the captions of Figures 3 and 4 both state an ICC of 0.1. Since Figure 3 is presented as confirming the analytic variance formulas, the simulation parameter must be unambiguous; please correct the inconsistency and report the Monte Carlo standard errors for the simulated variances so the visual agreement can be assessed.","section":"Section 4, Figures 3 and 4"},{"comment":"The abstract's Results paragraph says 'we illustrate DISC sampling using a household survey dataset from South Africa,' but no such illustration appears in Sections 1-5 or Appendix B of the manuscript. As written, the abstract promises content the paper does not deliver; either add the application or remove the claim.","section":"Abstract and body"},{"comment":"The variance reduction in Eq. (5) relies entirely on the time-invariant cluster effect α_i in model (3) canceling when the same clusters are differenced. The Discussion's limitation paragraph mentions changes in the number or size of clusters but does not mention the equally important case where cluster-level effects drift over time; if α_i is replaced by α_i + γ_it with time-varying γ_it, the DISC variance gains an additional term and the advantage shrinks. Please add an explicit caveat that the claimed precision gain assumes stable cluster effects, not just stable cluster populations.","section":"Section 3 and Discussion"}],"minor_comments":[{"comment":"The text contains a typo: 'THe value of τ' should read 'The value of τ'.","section":"Section 4"},{"comment":"With 1,000 simulation replicates, the relative standard error of a variance estimate is roughly √(2/999) ≈ 4.5%; adding Monte Carlo standard errors or confidence bands around the simulated variance points would make the agreement with the analytic curves more interpretable.","section":"Figures 3 and 4"},{"comment":"The y-axis label 'Total variance' is used after scaling Var(δ̂_DISC)=1, so the plotted quantity is actually the variance ratio; please relabel the axis accordingly.","section":"Figure 2"},{"comment":"The statement that the variance ratio increases as the number of clusters decreases would benefit from a brief reference back to the formula 1 + mρ/(1−ρ) with n = mC, making the dependence on C immediate.","section":"Section 5"}],"recommendation":"minor_revision","confidential_remarks":"The paper is within scope for stat.ME and makes a useful, if incremental, contribution. The main concerns for the editor are that the abstract overpromises a South Africa illustration that is not in the body and that the simulation ICC is reported inconsistently between text and figure captions; neither issue affects the correctness of the central variance derivation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read Downey and Kenny's DISC paper. The core result is a clean, correct variance comparison. Under their two-period model with a time-invariant cluster effect, drawing fresh individuals from the same clusters across waves cancels the cluster effect, giving Var(δ_DISC)=8σ²/n versus Var(δ_RCS)=8(mτ²+σ²)/n. The ratio 1+mρ/(1−ρ) follows algebraically and matches simulation. That is the main contribution, and it holds.\n\nWhat's genuinely new is the explicit ratio and its framing for quasi-experimental DID designs; the design itself is acknowledged as cross-sectional sampling in cluster trials, so the novelty is modest but real. The Appendix B derivation is straightforward but correct, and the simulation confirms the theoretical variance. The DRDID extension is a useful check—it shows the precision gain persists with covariate adjustment, though that part is empirical rather than analytic.\n\nSoft spots, in proportion. The variance formulas rest on a single additive time-invariant cluster effect. If cluster effects drift or the cluster population changes between waves, the cancellation weakens. The paper explicitly flags the population-change limitation in the Discussion, but the drift case is only implied. The RCS derivation assumes zero cluster overlap across waves, a simplification that is reasonable with many clusters. There's a minor reporting inconsistency: the text says ICC=0.2 for the first simulation while Figure 3's caption says ICC=0.1. It doesn't affect the analytic result, but it should be fixed. Also, no Monte Carlo standard errors around simulated variances, so the match with theory is visual rather than formally tested.\n\nThe citation pattern looks fine; the paper positions itself honestly. Who benefits: researchers planning cluster-based longitudinal surveys, especially in global health and implementation science. This deserves a serious referee. It's not groundbreaking, but it is practically useful and methodologically sound, and the explicit variance formula fills a real gap in the literature.","headline":"A correct, practically useful variance ratio for the DISC design; the main claim holds up, with the expected caveats about time-invariant cluster effects.","tokens_in":10030,"tokens_out":2164,"would_cite":true,"duration_ms":18196,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05"],"pacs":[],"model":"deepseek-v4-flash","headline":"A study design that samples the same clusters but different individuals at each wave produces difference-in-differences estimators with substantially lower variance than repeated cross-sectional sampling, with the gain captured by an…","keywords":["difference-in-differences","DISC design","repeated cross-sectional sampling","cohort sampling","cluster sampling","intraclass correlation","estimator variance","longitudinal surveys"],"falsifier":"Simulate data with a cluster-by-time interaction, for example $\\alpha_{it} \\sim \\mathrm{N}(0, \\tau^2)$ with $\\mathrm{Corr}(\\alpha_{i1}, \\alpha_{i2}) = 0.8$ under the model used in Section 3, and compare the empirical variances of $\\hat{\\delta}_{\\mathrm{DISC}}$ and $\\hat{\\delta}_{\\mathrm{RCS}}$; if the DISC variance rises above $8\\sigma^2/n$ and the ratio falls below $1 + m\\rho/(1-\\rho)$, the time-invariant cluster-effect assumption is doing the work. The same test can be run with clusters entering and leaving the frame between waves.","tokens_in":9166,"feed_emoji":"📊","tokens_out":11203,"duration_ms":89050,"temperature":0.7,"pith_summary":"This paper proposes a hybrid longitudinal sampling design, DISC (Different Individuals, Same Clusters), that redraws a fresh cross-sectional sample of individuals within each selected cluster at every wave, while returning to the same clusters. The paper claims that for difference-in-differences estimators, this one change removes the contribution of time-invariant cluster effects from the estimator's variance, making DISC substantially more precise than a repeated cross-sectional (RCS) design under the same total sample size. The claimed gain is summarized by a closed-form variance ratio $1 + m\\rho/(1-\\rho)$, which grows with both cluster size and the intraclass correlation $\\rho$. If the claim holds, a study that cannot follow individuals over time can still reduce DID estimator variance substantially by committing to a fixed set of clusters across waves.","feed_headline":"Same clusters, new individuals: treatment-effect variance drops 74%","feed_subtitle":"A hybrid survey design keeps cluster effects from inflating difference-in-differences estimates, with an exact formula for the gain.","key_machinery":"The load-bearing object is the variance identity comparing $\\hat{\\delta}_{\\mathrm{RCS}}$ and $\\hat{\\delta}_{\\mathrm{DISC}}$ under the model above. Because the same clusters appear at both time points under DISC, the cluster-level term $\\alpha_i$ enters the before-after difference within each cluster as $\\alpha_i - \\alpha_i = 0$, leaving only individual noise; this cancellation yields $\\operatorname{Var}(\\hat{\\delta}_{\\mathrm{DISC}}) = 8\\sigma^2/n$, while RCS retains both noise and cluster variation, giving $\\operatorname{Var}(\\hat{\\delta}_{\\mathrm{RCS}}) = 8(m\\tau^2+\\sigma^2)/n$. The ratio $1 + m\\rho/(1-\\rho)$, with $\\rho = \\tau^2/(\\tau^2+\\sigma^2)$, then converts the identity into a sample-size planning statement.","core_discovery":"The central discovery is an exact variance comparison. Under the generative model $Y_{ijk} = \\mu_j + \\delta X_{ij} + \\alpha_i + \\epsilon_{ijk}$ with i.i.d. cluster effects $\\alpha_i \\sim \\mathrm{N}(0, \\tau^2)$ and individual noise $\\epsilon_{ijk} \\sim \\mathrm{N}(0, \\sigma^2)$, the simple DID estimator has variance $8(m\\tau^2+\\sigma^2)/n$ under RCS and $8\\sigma^2/n$ under DISC, so the RCS variance is $1 + m\\rho/(1-\\rho)$ times larger. Simulations reproduce the formula, and a simulation with a doubly-robust DID estimator that adjusts for covariates shows the precision gain persists with covariate adjustment. The authors present the design as a practical middle ground: sample the same clusters each wave, take a different sample of individuals within each cluster, and analyze with existing longitudinal estimators.","pith_inferences":["Beyond the paper: the variance gain is bought at the price of wave-by-wave representativeness; when clusters enter, dissolve, or change composition during the study, the fixed cluster list describes the subpopulation present across the whole study, and the DISC estimand is cleaner for treatment effects but not for population-level change.","Beyond the paper: a testable extension is the three-level sampling case using the paper's S-S-D and S-D-D notation, where one would predict an intermediate variance reduction whose size depends on which level's cluster effect cancels.","Beyond the paper: because clusters are fixed across waves, cluster-level covariates can be measured once, so a natural comparison is whether analysis-phase adjustment for cluster covariates can recover part of the DISC gain without repeating cluster samples."],"forward_implications":["With 40 clusters and 25 individuals per cluster, the repeated cross-sectional variance is 2.3 times larger than DISC at an intraclass correlation of 0.05, 3.8 times larger at 0.1, and 7.3 times larger at 0.2.","Since variance is inversely proportional to sample size, a study powered at 90 percent could be run with roughly half the sample at an ICC of 0.05, and roughly one quarter of the sample at an ICC of 0.1, if DISC replaces RCS.","The precision gain holds for a doubly-robust DID estimator that uses covariate information, not just for the simple unadjusted linear estimator, as shown in simulation.","By a simple modification of the Section 3 variance argument, the same variance ratio applies to uncontrolled before-and-after comparisons, and the design transfers to cluster randomized trials, stepped wedge trials, and interrupted time series designs."],"supporting_citations":[{"why":"Supplies the doubly-robust DID estimator used in the simulation study and the causal conditions under which the DID estimand identifies the ATT.","marker":"Sant’Anna and Zhao (2020)"},{"why":"Establishes the conventional result that cohort designs are more efficient than cross-sectional designs, the gap DISC aims to close without individual follow-up.","marker":"Diehr et al. (1995)"},{"why":"Provides the probability-proportional-to-size cluster sampling framework assumed in the analytic variance derivations.","marker":"Rosén (1997)"},{"why":"Justifies equal weighting of individuals under probability-proportional-to-size sampling, supporting the participant-average estimand analyzed in Section 3.","marker":"Makela et al. (2018)"},{"why":"Gives the empirical intraclass correlation range (0.02 to 0.1) used to calibrate the numerical examples of the variance ratio.","marker":"Korevaar et al. (2021)"},{"why":"Provides the simulation engine used to produce the empirical variance curves that confirm the analytic formulas.","marker":"Kenny and Wolock (2024)"}],"fun_headline_variants":["Same clusters, new people: DID variance cut by 74%","DISC design: reuse clusters, swap individuals, cut DID variance by 74%","Hybrid survey design: same clusters, fresh individuals, 74% variance drop","Reuse clusters, resample individuals: DID variance down 74%","Precision boost for DID: sample same clusters, new individuals each wave"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each cluster's effect is a single time-invariant additive term that is identical at both surveys and independent of treatment and sampling, so the entire DISC variance gain is the cancellation of that term — a premise that fails if cluster effects drift, clusters change composition, or new clusters enter the population.","fun_headline_variants_meta":{"raw":{"variants":["Same clusters, new people: DID variance cut by 74%","DISC design: reuse clusters, swap individuals, cut DID variance by 74%","Hybrid survey design: same clusters, fresh individuals, 74% variance drop","Reuse clusters, resample individuals: DID variance down 74%","Precision boost for DID: sample same clusters, new individuals each wave"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001026,"raw_usage":{"total_tokens":4380,"prompt_tokens":1054,"completion_tokens":3326,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":670,"completion_tokens_details":{"reasoning_tokens":3226}},"tokens_in":670,"tokens_out":3326,"duration_ms":20366,"temperature":1.0,"reasoning_tokens":3226,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:42:12.129941+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate data with a cluster-by-time interaction, for example $\\alpha_{it} \\sim \\mathrm{N}(0, \\tau^2)$ with $\\mathrm{Corr}(\\alpha_{i1}, \\alpha_{i2}) = 0.8$ under the model used in Section 3, and compare the empirical variances of $\\hat{\\delta}_{\\mathrm{DISC}}$ and $\\hat{\\delta}_{\\mathrm{RCS}}$; if the DISC variance rises above $8\\sigma^2/n$ and the ratio falls below $1 + m\\rho/(1-\\rho)$, the time-invariant cluster-effect assumption is doing the work. The same test can be run with clusters entering and leaving the frame between waves.","supporting_citations":[{"cited_title":"Doubly robust difference-in-differences estimators","cited_arxiv_id":null,"evidence_quote":"Supplies the doubly-robust DID estimator used in the simulation study and the causal conditions under which the DID estimand identifies the ATT."},{"cited_title":"Optimal survey design for community intervention evaluations: cohort or cross-sectional? Journal of clinical epidemiology, 48 0 (12): 0 1461--1472, 1995","cited_arxiv_id":null,"evidence_quote":"Establishes the conventional result that cohort designs are more efficient than cross-sectional designs, the gap DISC aims to close without individual follow-up."},{"cited_title":"Bayesian inference under cluster sampling with probability proportional to size","cited_arxiv_id":null,"evidence_quote":"Justifies equal weighting of individuals under probability-proportional-to-size sampling, supporting the participant-average estimand analyzed in Section 3."},{"cited_title":"Intra-cluster correlations from the clustered outcome dataset bank to inform the design of longitudinal cluster trials","cited_arxiv_id":null,"evidence_quote":"Gives the empirical intraclass correlation range (0.02 to 0.1) used to calibrate the numerical examples of the variance ratio."},{"cited_title":"SimEngine : A modular framework for statistical simulations in R","cited_arxiv_id":null,"evidence_quote":"Provides the simulation engine used to produce the empirical variance curves that confirm the analytic formulas."}],"review_version":1}