Pith. sign in

REVIEW 4 major objections 3 minor 7 references

Sensitivity of weighted least squares estimators to omitted variables

T0 review · 4 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The bias of weighted-regression treatment estimates from an omitted variable is exactly a product of two weighted partial $R^2$ values.

desk verdict A clean, honest extension of Cinelli–Hazlett to weighted regression; the semi-weighted benchmarking and translator term are the genuinely new parts, and the fixed-weight target is a disclosed judgment call, not a hidden flaw. read the letter →

arxiv 2508.02954 v1 pith:OGMFEOCV submitted 2025-08-04 stat.ME

classification stat.ME MSC 62J0562F3562D20
keywords weightedleastsquaresomittedvariablebiassensitivityanalysisunobservedconfoundingpartialR-squaredrobustnessvaluecausalinferenceweightingestimators
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to give users of weighted regression a simple way to say how robust their causal conclusions are to unobserved confounding. It claims the bias of the treatment coefficient, measured against the weighted regression that also includes the omitted variable, is fully controlled by two weighted partial $R^2$ values: one for how well the omitted variable explains treatment given the covariates, and one for how well it explains the outcome given treatment and covariates. If true, the same sensitivity analysis works for inverse-propensity, matching, balancing, and stratification weights without modeling how the weights themselves would change. The paper also provides robustness values, bounds by comparison with observed covariates, and bootstrap confidence intervals, implemented in the weightsense R package.

What carries the argument

The central object is the weighted omitted-variable-bias decomposition of Equation (21): the coefficient difference is a product of two weighted partial correlations, $R_w(Y \sim Z \mid D,X)$ and $R_w(D \sim Z \mid X)$, divided by $\sqrt{1-R^2_w(D \sim Z \mid X)}$, times a ratio of weighted residual standard deviations. Weighted partial $R^2$ is defined as the share of remaining weighted variance in one variable explained by another after partialing out the conditioning variables in the weighted empirical distribution. The additional machinery is a semi-weight benchmarking device: because weighting often makes treatment and covariates nearly uncorrelated, the treatment-side benchmark uses weights estimated without the benchmark covariate, with a translator term converting semi-weighted strength to weighted strength; the bootstrap supplies adjusted inference.

What would settle it

Simulate data with a known confounder $Z$, compute the two weighted partial $R^2$ values, and compare the observed difference between the weighted least squares estimate and the weighted regression that adds $Z$ against the right-hand side of Equation (21); any systematic mismatch beyond sampling error refutes the bias formula.

Watch

Extended reading notes

Core claim

In the paper's notation, the omitted-variable bias of the weighted least squares estimator relative to the weighted regression that also includes the unobserved variable is $$\mathrm{dbias}(\hat{\tau}_{\mathrm{wls}}) = R_w(Y \sim Z \mid D,X)\, \frac{R_w(D \sim Z \mid X)}{\sqrt{1-$R^{2}$_w(D \sim Z \mid X)}}\, \frac{\hat{\mathrm{sd}}_w($Y^{{\perp_w X,D}}$)}{\hat{\mathrm{sd}}_w($D^{{\perp_w X}}$)}.$$ The only unknown quantities on the right are two weighted partial $R^2$ values: how much of the leftover weighted variation in treatment and in outcome an omitted variable explains after conditioning on the observed covariates. The paper calls these the sensitivity parameters, derives robustness values, extreme-scenario thresholds, and covariate benchmarking bounds from them, and supplies a percentile bootstrap for adjusted confidence intervals.

Load-bearing premise

The bias formula assumes the reference estimator only adds the omitted variable to the weighted outcome regression and leaves the weights unchanged; if real omitted confounding would also change how the weights are constructed, the formula captures only the outcome-model part of the bias.

Editorial extensions

If this is right

  • Any weighting scheme can be assessed with the same two weighted partial $R^2$ parameters, because the bias formula leaves the origin of the weights unspecified.
  • Investigators can report a robustness value: the equal strength of confounding in treatment and outcome needed to reduce the estimate by a specified fraction or render it insignificant, without assuming a distribution for the omitted variable.
  • With observed covariates as benchmarks, a covariate as strong as a chosen variable (or a stated multiple) yields formal upper bounds on both sensitivity parameters, so users can see whether plausible confounding would change conclusions.
  • A percentile bootstrap, with cluster or fixed-weight variants, provides confidence intervals for the adjusted estimate at chosen sensitivity parameter values.
  • For weights that exactly balance covariate means, the tools apply directly to the weighted difference in means as well as to the weighted regression.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension would be to invert the bias formula and plot the boundary of the sensitivity-parameter grid on which the adjusted estimate crosses a policy-relevant threshold such as sign, minimum effect, or significance.
  • When full weights and semi-weights diverge strongly, the translator term can exceed 7; one could automate flagging of such datasets by correlating full and semi-weights, a diagnostic the paper leaves informal.
  • Because the bias formula makes no distributional assumption on $Z$, it could also cover omitted treatment-by-covariate interactions, turning the sensitivity parameters into statements about heterogeneity rather than mean confounding; this is not worked out in the paper.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The paper proposes a sensitivity-analysis framework for weighted least squares estimators of treatment effects. It defines a reference estimator tau_hat_target that augments the weighted outcome regression with an omitted scalar covariate Z while holding the weights fixed, and proves that the sample difference between tau_hat_wls and tau_hat_target equals the product of two weighted partial correlations normalized by residual standard deviations (Eq. 21). On this basis it develops robustness values, an extreme-scenario diagnostic, covariate benchmarking bounds using semi-weights, a percentile bootstrap for adjusted inference, and an extension to weighted difference-in-means estimators. The methods are applied to the Darfur data with inverse-propensity-score, matching, and balancing weights, and supported by simulation appendices.

Significance. If the framework is valid, it would be a practically useful extension of Cinelli-Hazlett to weighting estimators, with interpretable R2 sensitivity parameters and no distributional assumptions on Z. The central algebraic identity (Eq. 21, Appendix B.1) is correct and exact for the sample difference it defines. The semi-weight benchmarking idea addresses a real difficulty (zero weighted correlation between D and X), and the application covers several common weighting schemes. However, the benchmarking bound in Eq. (26) is not proven as stated, the causal interpretation rests on a strong fixed-weight existence argument, and the adjusted-inference claims for matching with replacement outrun the available theory. These issues affect load-bearing parts of the paper, so the current version needs substantial revision.

major comments (4)
  1. [3.2.4, Eq. (26), Appendix B.4.1] The claimed identity R2_w(D~Z|X) = κ_w/w(-j)(D) * R2_w(-j)(D~X(j)|X(-j)) / (1 - R2_w(D~X(j)|X(-j))) is not correct for the κ defined in Eq. (23). The proof's 'without loss of generality' replacement of Z by a variable orthogonal to X is not WLOG: the numerator R2_w(D~Z|X(-j)) (and hence κ) changes under that replacement. Concrete counterexample in the unweighted case (a special case of this framework, Section 2.2.2): let X(-j) be empty, let X and E be independent standard normal, set D = X + E and Z = X - E. Then R2(D~X) = 0.5, R2(D~Z) = 0, so κ = 0, while the actual partial R2 of D on Z given X is 1 (the residuals are E and -E). Eq. (26) therefore returns 0, neither the true value nor a conservative upper bound. The same issue propagates to the multi-covariate bound in Eq. (51) and to the expression for R2_w(Z~X(1:j)|D,X(-j)) in Eq. (53). The benchmarking section needs a corrected derivation, a redefined κ, or a valid conservative bound; the numerical benchmark results in Tables 2-5 need to be revisited accordingly.
  2. [3.1, Expression (19), Section 3.3, footnote 10] The bias decomposition in Eq. (21) is an exact identity for the difference between tau_hat_wls and tau_hat_target, but the paper's central interpretive claim is that tau_hat_target can be regarded as an unbiased reference. This rests on footnote 10's construction of a univariate Z using the true potential-outcome regression functions, with weights assumed to be functions of D and X alone. That is an existence result in population; it is not a finite-sample guarantee when weights are estimated, and a user's chosen sensitivity parameters need not correspond to such a Z. Since an investigator who actually observed Z would typically re-estimate the weights, the tools quantify sensitivity of the weighted outcome regression to an omitted regressor, not the full effect of unobserved confounding on the weighting procedure. The manuscript should either provide explicit conditions under which Eq. (20) is the bias of tau_hat_wls relative to the target estimand, or consistently frame the contribution as sensitivity of the fixed-weight regression and qualify the causal language in the abstract and Section 3.
  3. [3.2.2, Appendix A.1, Section 5] The percentile bootstrap for adjusted inference is presented as a contribution, but for matching with replacement the standard bootstrap is known to be inconsistent (Abadie and Imbens 2008), and the paper's own Figure 7 shows coverage declining with n for the bootstrap that re-estimates the weights. The fixed-weight variant that achieves nominal coverage in the simulations is not supported by any theorem; the manuscript itself states in Section 5 that the properties of the bootstrap in the matching setting 'remain understudied.' As the paper stands, the main text overstates the status of the adjusted-inference tool. Please either add a validity result for the proposed (or fixed-weight) bootstrap under explicit conditions, or clearly label this part of the procedure as heuristic and move the caveat from the limitations section into Section 3.2.2.
  4. [3.2.1, Eq. (21), Tables 2-5] The sensitivity parameters are defined as the two R2 values, but Eq. (21) requires the signed partial correlations R_w(Y~Z|D,X) and R_w(D~Z|X). For fixed R2 values the sign of the bias is not determined, so the 'adjusted estimate' reported in Tables 2-5 is not uniquely defined by the stated inputs. The text says the correlations are 'freely varying in both magnitude and sign' but does not state which sign convention is used to produce the adjusted estimates and confidence intervals. Please specify the convention (for example, the worst-case direction, or the sign implied by the benchmark covariate), and state it wherever adjusted estimates are reported.
minor comments (3)
  1. [Footnote 1 and Abstract] Footnote 1 states that the weightsense package is 'to be made available upon acceptance,' while the abstract says the tools are 'made available'; please align these statements and provide a repository or version if one exists.
  2. [Appendix A.1, Figures 5-7] The figure captions and notes contain the typo 'Fixed Boostrap' (also 'Boostrap' in the note to Figure 7); should read 'Fixed Bootstrap.'
  3. [Section 3.2.2, Step 3] Step 3 says to 'choose values for the sensitivity parameters,' but Eq. (21) also depends on the signs of the partial correlations; please clarify that the user must choose signs as well as R2 values, or state a worst-case sign convention.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the weighted OVB decomposition is an exact algebraic identity and the sensitivity parameters are user-specified inputs, not fitted predictions.

full rationale

The paper's central result, Eq. 21, is an algebraic identity derived in Appendix B.1 that exactly decomposes the difference between tau_hat_wls and tau_hat_target, where tau_hat_target is the weighted regression with Z added and weights held fixed. The two weighted partial R2 parameters are treated as unknowns that the user specifies, not as quantities fitted to data and then relabeled as predictions. The robustness values, extreme-scenario thresholds, and benchmarking bounds are derived from the same identity using algebra and inequalities in Appendices B.2-B.4; none of these derivations assumes the conclusion it seeks to establish. Although the paper adapts machinery from Cinelli and Hazlett (2020), a work sharing an author, the load-bearing steps here are proved in this paper rather than imported unexamined: the conservative claim in Section 3.3 is accompanied by an explicit explanation of why R2w(D~Z|X) is no smaller than the partial R2 of the linear combination Z, and the bootstrap procedure is validated against an external DGP in Appendix A.1. The fixed-weight reference estimator is a stated modeling choice and a limitation, not a circular reduction: the paper does not claim that real confounders never change the weights, only that its tools quantify the omitted-variable bias in the weighted outcome regression for any user-specified Z. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no known result is merely relabeled. The derivation is self-contained and the sensitivity parameters are inputs, so the paper has no significant circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No free parameters are fitted to data. The sensitivity parameters R2_w(D~Z|X) and R2_w(Y~Z|D,X) and the benchmark multiples kappa are user-specified inputs, not quantities estimated from data to make the derivation work. The covariates X, weights w, and unobserved variable Z are standard causal-inference objects, not new postulated entities.

assumptions (5)
  • domain assumption SUTVA and consistency: potential outcomes for unit i do not depend on other units' treatments, and Y = Y(D) is observed.
    Stated in Section 2. They are standard causal inference assumptions and are not the target of the sensitivity analysis.
  • domain assumption Weights are functions of D and X only and are treated as fixed when defining the target estimator tau_hat_target.
    Section 3.1, Expressions 19-20, and footnote 10. The entire framework holds the weights fixed; if weights would be re-estimated with Z, the bias formula has a different interpretation.
  • domain assumption A univariate linear Z can represent the combined effect of omitted variables and misspecification so that including it in the weighted regression (with fixed weights) yields an unbiased estimate of the target estimand.
    Section 3.3 and footnote 10. The construction defines Z in terms of E[Y(d)|X,Z_tilde], which is assumed correctly specified and linearly entering the regression.
  • domain assumption Sample tuples (X_i, D_i, Y_i(d)) are i.i.d., and weights are strictly positive and sum to n.
    Section 2. Needed for the probability-limit interpretation of weighted regressions and for bootstrap inference.
  • standard math Frisch-Waugh-Lovell decomposition / omitted variable bias algebra for weighted least squares.
    Used in Appendix B proofs; holds algebraically with any weights.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sensitivity of weighted least squares estimators to omitted variables." pith.science (2026). https://pith.science/paper/OGMFEOCV

@misc{pith2026250802954,
  author       = {Pith},
  title        = {Pith review of: Sensitivity of weighted least squares estimators to omitted variables},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OGMFEOCV}},
  note         = {Machine review of arXiv:2508.02954}
}
abstract

This paper introduces tools for assessing the sensitivity, to unobserved confounding, of a common estimator of the causal effect of a treatment on an outcome that employs weights: the weighted linear regression of the outcome on the treatment and observed covariates. We demonstrate through the omitted variable bias framework that the bias of this estimator is a function of two intuitive sensitivity parameters: (i) the proportion of weighted variance in the treatment that unobserved confounding explains given the covariates and (ii) the proportion of weighted variance in the outcome that unobserved confounding explains given the covariates and the treatment, i.e., two weighted partial $R^2$ values. Following previous work, we define sensitivity statistics that lend themselves well to routine reporting, and derive formal bounds on the strength of the unobserved confounding with (a multiple of) the strength of select dimensions of the covariates, which help the user determine if unobserved confounding that would alter one's conclusions is plausible. We also propose tools for adjusted inference. A key choice we make is to examine only how the (weighted) outcome model is influenced by unobserved confounding, rather than examining how the weights have been biased by omitted confounding. One benefit of this choice is that the resulting tool applies with any weights (e.g., inverse-propensity score, matching, or covariate balancing weights). Another benefit is that we can rely on simple omitted variable bias approaches that, for example, impose no distributional assumptions on the data or unobserved confounding, and can address bias from misspecification in the observed data. We make these tools available in the weightsense package for the R computing language.

Figures

Figures reproduced from arXiv: 2508.02954 by the authors.

Figure 1
Figure 1. Inverse propensity score weights and semi-weights for estimating the ATE [PITH_FULL_IMAGE:figures/full_fig_p020_1.png] view at source ↗
Figure 2
Figure 2. Comparison of propensity score matching weights and semi-weights for estimating the [PITH_FULL_IMAGE:figures/full_fig_p021_2.png] view at source ↗
Figure 3
Figure 3. Exact matching weights and semi-weights for estimating the ATT [PITH_FULL_IMAGE:figures/full_fig_p023_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Mean balancing weights and semi-weights for estimating the ATT [PITH_FULL_IMAGE:figures/full_fig_p025_4.png]
Figure 5
Figure 5. Figure 5: Percentile bootstrap coverage rates in DGP 1 for inverse propensity score weights for the [PITH_FULL_IMAGE:figures/full_fig_p034_5.png]
Figure 6
Figure 6. Figure 6: Percentile bootstrap coverage rates in DGP 1 for entropy balancing weights for the ATT [PITH_FULL_IMAGE:figures/full_fig_p034_6.png]
Figure 7
Figure 7. Figure 7: Percentile bootstrap coverage rates in DGP 1 for propensity score matching for the ATT [PITH_FULL_IMAGE:figures/full_fig_p035_7.png]
Figure 8
Figure 8. Figure 8: Percentile cluster-bootstrap coverage rates in DGP 2 for inverse propensity score weights [PITH_FULL_IMAGE:figures/full_fig_p036_8.png]
Figure 9
Figure 9. Figure 9: a, it is apparent that Z and D are moderately correlated overall (at approximately 0.218), but are highly correlated for |Z| ≤ 1 (at approximately 0.758). Thus, were X used to benchmark the strength of Z, weights that neglect (i.e., set wi ≈ 0) units with |Zi | > 1 wou…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

7 extracted references · 6 canonical work pages

  1. [1]

    Abadie, A., Diamond, A., and Hainmueller, J. (2010). Synthetic control methods for compara- tive case studies: Estimating the effect of California’s tobacco control program. Journal of the American Statistical Association, 105(490):493–505. Abadie, A. and Gardeazabal, J. (2003). The economic costs of conflict: A case study of the Basque country. American ...

  2. [4]

    Note that when θ2 = 0, the error is homoscedastic

    where (δD +ϵ) makes up a combined error term, and θ2∈{ 0, 4, 16} determines the extent of the error’s heteroscedasticity. Note that when θ2 = 0, the error is homoscedastic. Furthermore, there is no treatment effect (i.e., the ATE, ATT, and ATC are all 0). We apply the percentile bootstrap procedure proposed in Section 3.2.2 to make 95% confidence interval...

  3. [6]

    Bootstrap

    32 Figure 5: Percentile bootstrap coverage rates in DGP 1 for inverse propensity score weights for the ATE n=500 n=1000 n=2000 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Default Bootstrap Fixed Bootstrap (a) θ2 = 0 n=500 n=1000 n=2000 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 (b) θ2 = 4 n=500 n=1000 n=2000 0.50 0.55 0.60 0.65 0.70...

  4. [7]

    Bootstrap

    Figure 6: Percentile bootstrap coverage rates in DGP 1 for entropy balancing weights for the ATT n=500 n=1000 n=2000 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Default Bootstrap Fixed Bootstrap (a) θ2 = 0 n=500 n=1000 n=2000 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 (b) θ2 = 4 n=500 n=1000 n=2000 0.50 0.55 0.60 0.65 0.70 0.75 0.80...

  5. [8]

    Bootstrap

    33 Figure 7: Percentile bootstrap coverage rates in DGP 1 for propensity score matching for the ATT n=500 n=1000 n=2000 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Default Bootstrap Fixed Bootstrap (a) θ2 = 0 n=500 n=1000 n=2000 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 (b) θ2 = 4 n=500 n=1000 n=2000 0.50 0.55 0.60 0.65 0.70 0.75 0...

  6. [472]

    Tan, Z. (2006). A distributional approach for causal inference using propensity scores. Journal of the American Statistical Association , 101(476):1619–1637. Tan, Z. (2010). Bounded, efficient and doubly robust estimation with inverse weighting.Biometrika, 97(3):661–682. van der Laan, M. J. and Rubin, D. (2006). Targeted maximum likelihood learning.The In...

  7. [866]

    Rosenbaum, P. R. (1987). Sensitivity analysis for certain permutation inferences in matched ob- servational studies. Biometrika, 74(1):13–26. 30 Rosenbaum, P. R. (2002). Sensitivity to hidden bias. In Observational studies , pages 105–170. Springer, second edition. Rosenbaum, P. R. and Rubin, D. B. (1983). The central role of the propensity score in obser...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.