REVIEW 4 major objections 4 minor 1 cited by
In small-sample propensity score analyses, this paper argues, sandwich variance estimators systematically understate uncertainty, while bootstrapping with propensity score re-estimation gives reliable coverage.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
In small samples, percentile bootstrap with propensity-score re-estimation beats sandwich and fixed-PS variance estimators, which can under-cover badly.
T0 review reviewed 2026-08-03 challenge →
load-bearing objection Useful small-sample simulation guidance, but the fixed-treatment-count benchmark and a few reporting gaps mean the headline numbers should be read with caution — worth refereeing but needing solid revision. the 4 major comments →
Improving Variance and Confidence Interval Estimation in Small-Sample Propensity Score Analyses: Bootstrap vs. Asymptotic Methods
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that the choice between sandwich and bootstrap variance estimation matters most when the propensity score is re-estimated rather than treated as fixed, and in small samples the numerical sandwich estimator is severely anti-conservative: for inverse probability of treatment weighting with 10% treatment prevalence it can have coverage below 85% even at n=1000. Meanwhile, the standard bootstrap with PS re-estimation (stdB-Est) yields standard errors that start above the empirical SE and converge to it, and the percentile bootstrap (Pct stdB-Est) gives robustly conservative 95% confidence intervals. The paper also shows that treating the PS as fixed is not necessarily conser
What carries the argument
The load-bearing objects are the propensity score (the probability of treatment given covariates) and how a variance estimator handles its estimation. The paper compares sandwich estimators (model-based, fully empirical, and numerical) that account for PS estimation against bootstrap procedures that either keep the PS fixed or re-estimate it inside each resample, and it introduces a 'stratified bootstrap' that resamples within treated and control groups to avoid quasi-separation. The key identity is the sandwich variance formula as an M-estimator, and the key contrast is whether the PS is treated as known or estimated; that contrast determines whether SEs are under- or over-estimated in smal
Load-bearing premise
The simulation always fixes the treatment fraction in each Monte Carlo sample, fits a correctly specified logistic propensity score model, and drops the smallest/imbalanced cells where quasi-separation occurs; the recommended ordering (bootstrap > sandwich > fixed) is established only under those favorable conditions.
What would settle it
Run the paper's exact simulation at n=1000, Pr(Z=1)=10%, with correctly specified PS and binary outcome, and compute the numerical sandwich CI coverage over 10,000 replications. If the numerical sandwich 95% CI covers the true ATE more than 90% of the time (rather than below 85%), the paper's central claim about severe under-coverage fails.
If this is right
- For standard IPTW in small or imbalanced samples, use bootstrap with PS re-estimation for standard errors and percentile bootstrap for confidence intervals; sandwich methods should be avoided.
- Treating the PS as fixed is not reliably conservative in small samples, so analysts cannot rely on that assumption.
- Overlap weights (ATO) yield stable inference at smaller sample sizes than IPTW, and may be a better choice when extreme weights are a concern.
- Augmenting with a correctly specified outcome model (AIPW) can make sandwich CIs even narrower and more anti-conservative; bootstrap remains robust.
- The differences are large enough to change statistical conclusions in the LIMIT-JIA example.
Where Pith is reading between the lines
- The paper's simulations fix the treatment proportion and use a correctly specified propensity score model; under PS misspecification or random treatment proportions, the ordering might differ, so the recommendation is not universal.
- For very rare treatments (prevalence below 10%), quasi-separation prevents the bootstrap from being applied; the paper excludes those cells, so the guidance may not extend to the most extreme imbalances.
- The results suggest that in single-arm trials with external controls (e.g., 50-100 treated patients), asymptotic CIs from software defaults could be misleading; a bootstrap with PS re-estimation should be the default.
- The stratified bootstrap's benefit suggests a general principle: preserve the treatment structure in resampling for small samples, which could apply to other calibration problems in causal inference.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper compares sandwich and bootstrap-based standard error and confidence interval estimators for propensity score-weighted treatment effect estimators (IPTW, ATO, and AIPW) in small samples via a large Monte Carlo simulation. The main findings are that the numerical sandwich estimator (NS) frequently underestimates standard errors and produces below-nominal coverage, especially under treatment imbalance, while the standard bootstrap with PS re-estimation (stdB-Est) yields conservative and reliable variances, with the percentile bootstrap (Pct stdB-Est) giving robustly conservative confidence intervals. The paper also introduces a 'purely empirical sandwich' estimator (PES) and applies the methods to the LIMIT-JIA case study.
Significance. If the conclusions are robust, the paper provides practically relevant guidance for a common and increasingly important problem: variance estimation in small-sample propensity score analyses. Strengths include an extensive factorial simulation (10,000 replicates, B=1000), a wide range of sample sizes and prevalences, and comparison of multiple CI construction methods; the GitHub code is a reproducibility asset. However, the generalizability of the recommendations is limited by the use of a correctly specified PS model throughout and by the conditional sampling benchmark, which may not match the sampling model underlying the variance estimators being compared. The abstract's emphasis on the stratified bootstrap is not aligned with the methods actually recommended in the main text.
major comments (4)
- [Section 3 (opening) and Section 3.3 vs. Section 2.1.1/2.1.2] The Monte Carlo design fixes the number of treated subjects in every replicate via stratified sampling, and the empirical SE is computed as the SD of point estimates across these replicates (Section 3.3). This is a conditional sampling variance. However, the sandwich estimators in Section 2.1.1 (Eqs. 3–4) and the standard bootstrap in Section 2.1.2 are derived for unconditional i.i.d. sampling; the standard bootstrap allows the treatment count to vary across resamples. Thus the SE ratios in Figures 2 and 4 compare unconditional variance estimators to a conditional benchmark. The standard bootstrap will tend to look conservative because it includes treatment-count variability that the stratified design removes, while sandwich estimators will look anti-conservative even if they correctly estimate the unconditional variance. This mismatch can inflate the headline ordering 'bootstrap reliabl
- [Abstract vs. Sections 4 and 6] The abstract states: 'A stratified bootstrap avoids quasi-separation and performs well.' But the main text recommends and evaluates the standard bootstrap (stdB-Est, Pct stdB-Est) in Sections 4.1, 4.2, and 6. The stratified bootstrap (strB-Est) appears only in the supplementary appendix (e.g., Figure 11), with no main-text coverage results. Moreover, Section 3.2 excludes scenarios that frequently produce quasi-separation, and no failure-handling rule is given for bootstrap replicates that encounter separation during PS re-estimation. The abstract's claim about the stratified bootstrap is therefore not actually demonstrated by the experiments shown. Please align the abstract with the recommended methods and either present the stratified bootstrap results in the main text or temper the claim.
- [Section 3.1/3.3 and Section 6] The treatment model is always a correctly specified logistic regression, since the simulation generates treatment from the same linear predictor used to fit the PS model (Section 3.1, Treatment Model). The only misspecification considered is in the outcome model for AIPW (Section 3.3). The recommendation to 'abandon unreliable asymptotic methods and instead use the bootstrap' (Section 6) is broad, but the simulations do not cover PS model misspecification, where bootstrap re-estimation of a wrong model can behave differently and sandwich estimators may have different robustness properties. Please add a PS misspecification scenario, or explicitly restrict the conclusions to correctly specified PS models.
- [Section 4.4] The AIPW SE ratio uses the non-augmented estimator Δ̂_ATE as the benchmark for the augmented estimator (Section 4.4, paragraph beginning 'A key aspect of our simulation design...'). This makes the ratio a combined measure of variance estimation and efficiency gain. The paper interprets a ratio below one for the correctly specified outcome model as evidence that the variance estimators are 'not underestimating.' That inference is problematic: the augmented estimator is expected to have a smaller true variance, so the ratio only validates the variance estimator if the benchmark is the empirical SD of Δ̂_Aug itself. Coverage results are not affected, but the SE-ratio conclusions in this section rest on a mismatched benchmark. Please also report the empirical SD of the augmented estimator, or restate the interpretation accordingly.
minor comments (4)
- [Section 2.2] The 'Double Bootstrap Interval' is defined as (2Δ̂ − Q0.975, 2Δ̂ − Q0.025). This is the basic/hybrid bootstrap interval, not Efron's double bootstrap, which is an iterated procedure. Please rename or correct the description.
- [Appendix 9.2.1] In the sentence about within-cluster differences, the text reads 'stfB-Est' — likely a typo for 'strB-Est'.
- [Throughout] Several typos and inconsistent abbreviations: 'boostrap' (Section 3.3), 'covaraites' (Section 2), 'Converage Probablity' in figure labels, and 'PSE' vs 'PES' in the Appendix. Please copyedit.
- [Sections 2.1.1 and Appendix 9.1] The new PES estimator is defined and derived but its performance is not discussed in the main text; the Appendix states that all Sandwich-Est methods behave similarly. A sentence in the main text clarifying its role would help.
Circularity Check
No significant circularity: simulation benchmarks are independent of the fitted estimators, and the PES/sandwich derivations are standard M-estimation rather than outputs of the simulations.
full rationale
The paper's central claims come from Monte Carlo simulation, not from fitting parameters and then presenting those fitted values as predictions. Section 3.1 explicitly simulates a super-population of N=1,000,000 with known treatment and outcome models; Section 3.3 computes the empirical SE as the standard deviation of 10,000 point estimates and evaluates coverage against the known true ATE. No variance estimator is calibrated to the simulation outcomes, and the new PES estimator in Section 2.1.1 and Appendix 9.1 is derived from stacked M-estimating equations, independent of the simulation results. Self-citations such as PSweight and overlap-weight references are used as computational tools or published methodology, not as an unverified uniqueness theorem that forces the conclusions. The stratified-sampling feature in the Monte Carlo design is a potential external-validity concern, since it conditions on treatment counts while the compared variance estimators target unconditional sampling, but this is not circularity in the defined sense: the comparisons are still against an independently simulated truth, and the claimed ordering is not an algebraic identity with the paper's inputs. No specific circular step can be exhibited, so the appropriate score is low.
Axiom & Free-Parameter Ledger
free parameters (4)
- Treatment-model intercept α0,treat =
calibrated by bisection for Pr(Z=1) 0.1-0.5
- Outcome-model intercept α0,outcome =
calibrated by bisection for Pr(Y(0)=1) 0.1-0.5
- Treatment-effect coefficient αtreat =
adjusted so true ATE = -0.02
- Simulation DGP constants (correlation 0.2, dichotomization percentiles, N=1,000,000) =
0.2; 10th-50th percentiles; 1,000,000
axioms (5)
- domain assumption Causal identification assumptions SUTVA, unconfoundedness, and positivity hold
- domain assumption The PS model is correctly specified as logistic in the 10 covariates
- domain assumption Treatment proportion is fixed at the target by stratified sampling
- standard math Estimating-equation sandwich theory and M-estimation asymptotics apply to the simulation estimators
- domain assumption The super-population DGP (covariates from MVN with quantile-dichotomized binary variables, logistic treatment and outcome) is representative of small-sample PS applications
invented entities (1)
-
Purely Empirical Sandwich estimator (Σ_PES)
no independent evidence
Cite this review
Pith. "Pith review of Improving Variance and Confidence Interval Estimation in Small-Sample Propensity Score Analyses: Bootstrap vs. Asymptotic Methods." pith.science (2026). https://pith.science/paper/Z6YMZAEK
@misc{pith2026251110911,
author = {Pith},
title = {Pith review of: Improving Variance and Confidence Interval Estimation in Small-Sample Propensity Score Analyses: Bootstrap vs. Asymptotic Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/Z6YMZAEK}},
note = {Machine review of arXiv:2511.10911}
}
read the original abstract
Propensity score (PS) methods are widely used to estimate treatment effects in non-randomized studies. Variance is typically estimated using sandwich or bootstrap methods, which can either treat the PS as estimated or fixed. The latter is thought to be conservative. Comparisons between the sandwich and bootstrap estimators have been compared in moderate to large sample sizes, favoring the bootstrap estimator. With the growing interest in treatments for rare disease and externally controlled clinical trials, very small sample sizes are not uncommon and the asymptotic properties of sandwich estimators may not hold. Bootstrap methods that allow for PS re-estimation can also generate problems with quasi-separation in small samples. It is unclear whether it is safe to prefer sandwich estimators or to assume that treating the PS as fixed is conservative. We conducted a Monte Carlo simulation to compare the performance of bootstrap versus sandwich variance and CI estimators for average treatment effects estimated with PS methods. We systematically evaluated the impact of treating the PS as fixed versus re-estimating it. These methodological comparisons were performed using Inverse Probability of Treatment Weighting (IPTW) and Augmented Inverse Probability of Treatment Weighting (AIPW) estimators. Simulations assessed performance under various conditions, including small sample sizes and different outcome and treatment prevalences. We illustrate the differences in our motivating example, the LIMIT-JIA trial. We show that the sandwich estimators can perform quite poorly in small samples, and fixed PS methods are not necessarily conservative. A stratified bootstrap avoids quasi-separation and performs well. Differences were large enough to alter statistical conclusions in our motivating example, LIMIT-JIA.
Forward citations
Cited by 1 Pith paper
-
Doubly Robust Estimators of Quantile Treatment Effects With Semiparametric Cumulative Probability Models
Quantile treatment effects can be estimated doubly robustly with a cumulative probability model, giving reliable distributional comparisons for skewed outcomes.
Reference graph
Works this paper leans on
-
[1]
The central role of the propensity score in observational studies for causal effects
Rosenbaum PR, Rubin DB. The central role of the propensity score in observational studies for causal effects. Biometrika 1983; 70(1): 41–55
1983
-
[2]
Balancing covariates via propensity score weighting
Li F, Morgan KL, Zaslavsky AM. Balancing covariates via propensity score weighting. Journal of the American Statistical Association 2018; 113(521): 390–400
2018
-
[3]
Addressing extreme propensity scores via the overlap weights
Li F, Thomas LE, Li F. Addressing extreme propensity scores via the overlap weights. American journal of epidemiology 2019; 188(1): 250–257
2019
-
[4]
Stratification and weighting via the propensity score in estimation of causal treatment effects: a comparative study
Lunceford JK, Davidian M. Stratification and weighting via the propensity score in estimation of causal treatment effects: a comparative study. Statistics in medicine 2004; 23(19): 2937–2960
2004
-
[5]
Bootstrap vs asymptotic variance estimation when using propensity score weighting with continuous and binary outcomes
Austin PC. Bootstrap vs asymptotic variance estimation when using propensity score weighting with continuous and binary outcomes. Statistics in medicine 2022; 41(22): 4426–4443
2022
-
[6]
An introduction to the bootstrap
Efron B, Tibshirani RJ. An introduction to the bootstrap . Chapman and Hall/CRC . 1994. Baoshan Zhang ET AL 19 Confounder LIMIT−JIA Part II Subgroup Sex Treatment Control Age ATE 95% CI Number of Active Joint SE p−value Uveitis Strata / Male Female < 6 >= 6 < 2 >= 2 Risk Lower Risk Higher 81 26 55 36 45 39 42 33 48 722 248 474 330 392 562 160 257 465 −0.2...
1994
-
[7]
Matching on the estimated propensity score
Abadie A, Imbens GW. Matching on the estimated propensity score. Econometrica 2016; 84(2): 781–807
2016
-
[8]
On variance of the treatment effect in the treated when estimated by inverse probability weighting
Reifeis SA, Hudgens MG. On variance of the treatment effect in the treated when estimated by inverse probability weighting. American Journal of Epidemiology 2022; 191(6): 1092–1097
2022
-
[9]
Propensity score weighting analysis and treatment effect discovery
Mao H, Li L, Greene T. Propensity score weighting analysis and treatment effect discovery. Statistical methods in medical research 2019; 28(8): 2439–2454
2019
-
[10]
On variance estimation of the inverse probability-of-treatment weighting estimator: A tutorial for different types of propensity score weights
Kostouraki A, Hajage D, Rachet B, et al. On variance estimation of the inverse probability-of-treatment weighting estimator: A tutorial for different types of propensity score weights. Statistics in Medicine 2024; 43(13): 2672–2694
2024
-
[11]
Comparative effectiveness of azathioprine and mycophenolate mofetil for myasthenia gravis (PROMISE-MG): a prospective cohort study
Narayanaswami P, Sanders DB, Thomas L, et al. Comparative effectiveness of azathioprine and mycophenolate mofetil for myasthenia gravis (PROMISE-MG): a prospective cohort study. The Lancet Neurology 2024; 23(3): 267–276
2024
-
[12]
ClinicalTrials.gov, Identifier NCT03841357
Preventing Extension of Oligoarticular Juvenile Idiopathic Arthritis (Limit-JIA). ClinicalTrials.gov, Identifier NCT03841357; . Accessed: 2025-08-29
2025
-
[13]
Management and outcomes in patients with moderate or severe functional mitral regurgitation and severe left ventricular dysfunction
Samad Z, Shaw LK, Phelan M, et al. Management and outcomes in patients with moderate or severe functional mitral regurgitation and severe left ventricular dysfunction. European heart journal 2015; 36(40): 2733–2741. 20 Baoshan Zhang ET AL
2015
-
[14]
Stürmer T, Joshi M, Glynn RJ, Avorn J, Rothman KJ, Schneeweiss S. A review of the application of propensity score methods yielded increasing use, advantages in specific settings, but not substantially different estimates compared with conventional multivariable methods. Journal of clinical epidemiology 2006; 59(5): 437–e1
2006
-
[15]
Doubly robust estimation in missing data and causal inference models
Bang H, Robins JM. Doubly robust estimation in missing data and causal inference models. Biometrics 2005; 61(4): 962– 973
2005
-
[16]
Variance estimation for the average treatment effects on the treated and on the controls
Matsouaka RA, Liu Y, Zhou Y. Variance estimation for the average treatment effects on the treated and on the controls. Statistical Methods in Medical Research 2023; 32(2): 389–403
2023
-
[17]
Causal inference in statistics, social, and biomedical sciences
Imbens GW, Rubin DB. Causal inference in statistics, social, and biomedical sciences . Cambridge university press . 2015
2015
-
[18]
Improving propensity score weighting using machine learning
Lee BK, Lessler J, Stuart EA. Improving propensity score weighting using machine learning. Statistics in medicine 2010; 29(3): 337–346
2010
-
[19]
Improving propensity score estimators’ robustness to model misspecification using super learner
Pirracchio R, Petersen ML, Van Der Laan M. Improving propensity score estimators’ robustness to model misspecification using super learner. American journal of epidemiology 2015; 181(2): 108–119
2015
-
[20]
Efficient estimation of average treatment effects using the estimated propensity score
Hirano K, Imbens GW, Ridder G. Efficient estimation of average treatment effects using the estimated propensity score. Econometrica 2003; 71(4): 1161–1189
2003
-
[21]
The calculus of M-estimation
Stefanski LA, Boos DD. The calculus of M-estimation. The American Statistician 2002; 56(1): 29–38
2002
-
[22]
PSweight: An R package for propensity score weighting analysis
Zhou T, Tong G, Li F, Thomas LE. PSweight: An R package for propensity score weighting analysis. arXiv preprint arXiv:2010.08893 2020
arXiv 2010
-
[23]
Interval estimation for treatment effects using propensity score matching
Hill J, Reiter JP. Interval estimation for treatment effects using propensity score matching. Statistics in medicine 2006; 25(13): 2230–2256
2006
-
[24]
Estimation of regression coefficients when some regressors are not always observed
Robins JM, Rotnitzky A, Zhao LP. Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association 1994; 89(427): 846–866
1994
-
[25]
Semiparametric theory and missing data
Tsiatis AA. Semiparametric theory and missing data . 4. Springer . 2006
2006
-
[26]
Variance estimation when using inverse probability of treatment weighting (IPTW) with survival analysis
Austin PC. Variance estimation when using inverse probability of treatment weighting (IPTW) with survival analysis. Statistics in medicine 2016; 35(30): 5642–5655
2016
-
[27]
bread" matrix 𝐀(𝜃) = − 𝐸[̇ 𝜓𝑖(𝜃)] and the
Hajage D, Chauvet G, Belin L, Lafourcade A, Tubach F, De Rycke Y. Closed-form variance estimator for weighted propensity score estimators with survival outcome. Biometrical Journal 2018; 60(6): 1151–1163. Baoshan Zhang ET AL 21 9 APPENDIX 9.1 Purely Empirical Sandwich Estimator for ATE In this section, we develop the purely empirical sandwich estimator ((...
2018
-
[28]
This estimator uses only empirical averages and does not rely on further model assumptions beyond the structure of the estimating equations themselves
= 1 𝑛 𝑛∑ 𝑖=1 { 𝐂(̄𝐀22,𝑖)−1(𝛙𝜇 𝑖 − ̄𝐇(̄𝐀11)−1𝜓𝛽 𝑖 } , (7) where𝜓𝜇 𝑖 = [𝜓𝜇1 𝑖 (̂𝜃),𝜓 𝜇0 𝑖 (̂𝜃)]⊤. This estimator uses only empirical averages and does not rely on further model assumptions beyond the structure of the estimating equations themselves. Its key advantage is robustness: it provides asymptotically valid variance estimates even if the PS model 𝑒(𝐗...
2004
-
[29]
(Marked as □ with red in figures)
Cluster 1: Fixed PS ( Fixed): Includes all methods, both asymptotic and bootstrap, that treat the PS as a fixed, known quantity, thereby ignoring estimation uncertainty. (Marked as □ with red in figures)
-
[30]
(Marked as ○ with green in figures)
Cluster 2: Estimated PS - Sandwich ( Sandwich-Est): Includes all asymptotic (sandwich) variance methods that analytically account for PS model estimation. (Marked as ○ with green in figures). Baoshan Zhang ET AL 23
-
[31]
(Marked as △ with blue in figures)
Cluster 3: Estimated PS - Bootstrap ( Bootstrap-Est): Includes all bootstrap methods where the PS model is re- estimated within each bootstrap resample to capture estimation uncertainty. (Marked as △ with blue in figures). This consistent, three-group structure strongly guided our selection of representative methods for the main manuscript. To clearly illu...
This paper was first reviewed by deepseek-v4-flash on August 3, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.