{"id":"f8797e3e-cb8d-4b2f-a33d-30cfac4b6f9c","arxiv_id":"2508.02954","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The bias of a weighted regression estimate due to an omitted confounder is exactly a function of two weighted partial R-squared values, yielding simple sensitivity statistics for IPW, matching, and balancing weights.","lead":"This paper develops sensitivity analysis tools for weighted regression estimates of causal effects, showing a weighted estimator's bias from an omitted confounder is governed by two weighted partial R-squared values. It extends a popular unweighted sensitivity framework to inverse-propensity, matching, and balancing weights without distributional assumptions on the omitted variable.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The weighted OVB formula in Eq. 21 is exact, but the fixed-weight reference estimator (Expression 19) is the load-bearing assumption; footnote 10's existence argument needs finite-sample verification when weights would also change.","rationale":"The reader's CONDITIONAL verdict identifies the fixed-weight reference estimator as the weakest assumption, and I agree that this is the most load-bearing assumption for the central claim. However, the paper's footnote 10 provides a strong existence argument: for binary D, a scalar Z can be constructed from the true outcome regression functions so that the fixed-weight augmented regression is unbiased for ATE (and similar constructions work for ATT/ATC), because the model becomes correctly specified and WLS with any weights depending on the regressors recovers the coefficient on D. This substantially mitigates the concern that re-estimating weights would change the bias. The residual concern is not mathematical but empirical and interpretational: the constructed Z depends on unknown functions, and the sensitivity parameters the user specifies (R2_w(D~Z|X), R2_w(Y~Z|D,X)) are not guaranteed to correspond to any real Z that makes tau_hat_target unbiased. If the user sets them based on the actual confounder's strength, the analysis is conservative per Section 3.3, but this is argued rather than tested. The paper also self-identifies limitations in the benchmarking translator and the bootstrap inference, which are acknowledged in Section 5 and Appendix A.1. These do not invalidate the exact bias formula, but they justify keeping the verdict CONDITIONAL rather than ACCEPT. My concrete test directly checks whether the fixed-weight bias decomposition matches the fully adjusted bias in a scenario where weights would change, which would settle whether the central claim holds beyond the sample identity.","tokens_in":32667,"tokens_out":28399,"duration_ms":289359,"concrete_test":"Simulate a DGP with Y = tau D + beta X + eta Z + epsilon and P(D=1|X,Z) = logit(gamma X + delta Z). Estimate IPW weights from a propensity score model using X only, and compute tau_hat_wls. Construct Z* = beta X + eta Z (the footnote-10 variable) and compute tau_hat_target by weighted regression on D, X, Z* with the original weights; verify whether the coefficient on D equals tau. Then compute the fully adjusted estimator (re-estimating weights with X and Z, then weighted regression on D, X, Z) and compare its bias relative to tau_hat_wls with the dbias given by Eq. 21. If tau_hat_target is unbiased and dbias matches the fully adjusted bias, the fixed-weight concern is resolved; if not, the tools understate confounding that changes the weights.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The bias decomposition in Eq. 21 is an exact sample identity for the difference between tau_hat_wls and tau_hat_target, where tau_hat_target is the weighted regression with Z added and weights held fixed (Expression 19). The central interpretive claim is that this difference captures the true omitted-variable bias because a univariate Z can always be chosen so that tau_hat_target is unbiased (footnote 10). That existence argument requires (i) weights that are functions of D and X alone, and (ii) a constructed Z that makes E[Y|D,X,Z] linear in (1, D, X, Z) with coefficient tau on D. For IPW, balancing, and matching weights, (i) holds. However, the construction uses the true outcome regression functions, which are unknown; the sensitivity parameters the user specifies are not guaranteed to correspond to such a Z. In practice, if Z were observed, the weights themselves would be re-estimated (e.g., propensity score including Z), and the fixed-weight augmented estimator is not the estimator an investigator would use. The paper argues the fixed-weight estimator can still be unbiased in population, but this is a formal existence claim, not a finite-sample guarantee when weights are estimated from data. Thus the two-R2 parameterization fully determines the sensitivity of tau_hat_wls relative to tau_hat_target, but the mapping from those parameters to the actual bias of tau_hat_wls as an estimator of the true effect depends on the strength of the footnote 10 argument, which is not empirically validated in the manuscript.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a sensitivity-analysis framework for weighted least squares estimators of treatment effects. It defines a reference estimator tau_hat_target that augments the weighted outcome regression with an omitted scalar covariate Z while holding the weights fixed, and proves that the sample difference between tau_hat_wls and tau_hat_target equals the product of two weighted partial correlations normalized by residual standard deviations (Eq. 21). On this basis it develops robustness values, an extreme-scenario diagnostic, covariate benchmarking bounds using semi-weights, a percentile bootstrap for adjusted inference, and an extension to weighted difference-in-means estimators. The methods are applied to the Darfur data with inverse-propensity-score, matching, and balancing weights, and supported by simulation appendices.","tokens_in":33063,"tokens_out":26485,"duration_ms":279534,"significance":"If the framework is valid, it would be a practically useful extension of Cinelli-Hazlett to weighting estimators, with interpretable R2 sensitivity parameters and no distributional assumptions on Z. The central algebraic identity (Eq. 21, Appendix B.1) is correct and exact for the sample difference it defines. The semi-weight benchmarking idea addresses a real difficulty (zero weighted correlation between D and X), and the application covers several common weighting schemes. However, the benchmarking bound in Eq. (26) is not proven as stated, the causal interpretation rests on a strong fixed-weight existence argument, and the adjusted-inference claims for matching with replacement outrun the available theory. These issues affect load-bearing parts of the paper, so the current version needs substantial revision.","major_comments":[{"comment":"The claimed identity R2_w(D~Z|X) = κ_w/w(-j)(D) * R2_w(-j)(D~X(j)|X(-j)) / (1 - R2_w(D~X(j)|X(-j))) is not correct for the κ defined in Eq. (23). The proof's 'without loss of generality' replacement of Z by a variable orthogonal to X is not WLOG: the numerator R2_w(D~Z|X(-j)) (and hence κ) changes under that replacement. Concrete counterexample in the unweighted case (a special case of this framework, Section 2.2.2): let X(-j) be empty, let X and E be independent standard normal, set D = X + E and Z = X - E. Then R2(D~X) = 0.5, R2(D~Z) = 0, so κ = 0, while the actual partial R2 of D on Z given X is 1 (the residuals are E and -E). Eq. (26) therefore returns 0, neither the true value nor a conservative upper bound. The same issue propagates to the multi-covariate bound in Eq. (51) and to the expression for R2_w(Z~X(1:j)|D,X(-j)) in Eq. (53). The benchmarking section needs a corrected derivation, a redefined κ, or a valid conservative bound; the numerical benchmark results in Tables 2-5 need to be revisited accordingly.","section":"3.2.4, Eq. (26), Appendix B.4.1"},{"comment":"The bias decomposition in Eq. (21) is an exact identity for the difference between tau_hat_wls and tau_hat_target, but the paper's central interpretive claim is that tau_hat_target can be regarded as an unbiased reference. This rests on footnote 10's construction of a univariate Z using the true potential-outcome regression functions, with weights assumed to be functions of D and X alone. That is an existence result in population; it is not a finite-sample guarantee when weights are estimated, and a user's chosen sensitivity parameters need not correspond to such a Z. Since an investigator who actually observed Z would typically re-estimate the weights, the tools quantify sensitivity of the weighted outcome regression to an omitted regressor, not the full effect of unobserved confounding on the weighting procedure. The manuscript should either provide explicit conditions under which Eq. (20) is the bias of tau_hat_wls relative to the target estimand, or consistently frame the contribution as sensitivity of the fixed-weight regression and qualify the causal language in the abstract and Section 3.","section":"3.1, Expression (19), Section 3.3, footnote 10"},{"comment":"The percentile bootstrap for adjusted inference is presented as a contribution, but for matching with replacement the standard bootstrap is known to be inconsistent (Abadie and Imbens 2008), and the paper's own Figure 7 shows coverage declining with n for the bootstrap that re-estimates the weights. The fixed-weight variant that achieves nominal coverage in the simulations is not supported by any theorem; the manuscript itself states in Section 5 that the properties of the bootstrap in the matching setting 'remain understudied.' As the paper stands, the main text overstates the status of the adjusted-inference tool. Please either add a validity result for the proposed (or fixed-weight) bootstrap under explicit conditions, or clearly label this part of the procedure as heuristic and move the caveat from the limitations section into Section 3.2.2.","section":"3.2.2, Appendix A.1, Section 5"},{"comment":"The sensitivity parameters are defined as the two R2 values, but Eq. (21) requires the signed partial correlations R_w(Y~Z|D,X) and R_w(D~Z|X). For fixed R2 values the sign of the bias is not determined, so the 'adjusted estimate' reported in Tables 2-5 is not uniquely defined by the stated inputs. The text says the correlations are 'freely varying in both magnitude and sign' but does not state which sign convention is used to produce the adjusted estimates and confidence intervals. Please specify the convention (for example, the worst-case direction, or the sign implied by the benchmark covariate), and state it wherever adjusted estimates are reported.","section":"3.2.1, Eq. (21), Tables 2-5"}],"minor_comments":[{"comment":"Footnote 1 states that the weightsense package is 'to be made available upon acceptance,' while the abstract says the tools are 'made available'; please align these statements and provide a repository or version if one exists.","section":"Footnote 1 and Abstract"},{"comment":"The figure captions and notes contain the typo 'Fixed Boostrap' (also 'Boostrap' in the note to Figure 7); should read 'Fixed Bootstrap.'","section":"Appendix A.1, Figures 5-7"},{"comment":"Step 3 says to 'choose values for the sensitivity parameters,' but Eq. (21) also depends on the signs of the partial correlations; please clarify that the user must choose signs as well as R2 values, or state a worst-case sign convention.","section":"Section 3.2.2, Step 3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a solid extension of Cinelli–Hazlett to weighted regression, and the new benchmarking machinery is worth knowing even if you don't use weights. The central identity (Eq. 21) is exact: the difference between the weighted regression and the weighted regression with Z added, weights fixed, is fully determined by two weighted partial R2 values. The proof is clean, and the claim is honestly scoped.\n\nWhat's genuinely new: the semi-weighted benchmarking with the translator term. When D and X are balanced by weights, the usual covariate-benchmarking denominator collapses, and they solve that by comparing Z to X(j) in a semi-weighted distribution and then translating back. The translator can be large—their Appendix A.3 example gives roughly 7.7—and they're upfront that this is the part that needs judgment. That's a real contribution, not a repackaging.\n\nThe soft spots are real but not disqualifying. The reference estimator keeps the weights fixed. If an actual confounder would change how weights are estimated, the bias decomposition captures only part of the story. The authors know this; they argue (footnote 10) that a univariate Z can always be constructed so that the fixed-weight target is unbiased, but that's an existence argument built on the true outcome regressions. It doesn't guarantee that the user-chosen sensitivity parameters correspond to a real Z of that kind. The paper would be stronger with a small simulation showing how far the fixed-weight target can diverge from a reweighted estimator when the propensity score depends on Z. Second, the adjusted-inference bootstrap for matching with replacement is supported only by simulations; they cite Abadie–Imbens and acknowledge the gap, but the fixed-weight version is an empirical fix without theory. Minor: the software package isn't available in the preprint, despite being prominently announced.\n\nThe citation pattern is fine—C&H is the direct ancestor, and the rest of the literature is covered. Nothing smells circular.\n\nWho it's for: anyone doing applied causal work with IPW, matching, or balancing weights who wants a quick sensitivity check, and methodologists working on OVB-style bounds. It deserves a serious referee. I'd engage with it.\n\nRecommendation: send to peer review. The core identity stands, the limitations are disclosed, and the benchmarking idea is worth developing even if the translator needs more theory.","headline":"A clean, honest extension of Cinelli–Hazlett to weighted regression; the semi-weighted benchmarking and translator term are the genuinely new parts, and the fixed-weight target is a disclosed judgment call, not a hidden flaw.","tokens_in":33458,"tokens_out":2040,"would_cite":true,"duration_ms":22131,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J05","62F35","62D20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The bias of weighted-regression treatment estimates from an omitted variable is exactly a product of two weighted partial $R^2$ values.","keywords":["weighted least squares","omitted variable bias","sensitivity analysis","unobserved confounding","weighted partial R-squared","robustness value","causal inference","weighting estimators"],"falsifier":"Simulate data with a known confounder $Z$, compute the two weighted partial $R^2$ values, and compare the observed difference between the weighted least squares estimate and the weighted regression that adds $Z$ against the right-hand side of Equation (21); any systematic mismatch beyond sampling error refutes the bias formula.","tokens_in":32457,"feed_emoji":"⚖️","tokens_out":11139,"duration_ms":103095,"temperature":0.7,"pith_summary":"The paper aims to give users of weighted regression a simple way to say how robust their causal conclusions are to unobserved confounding. It claims the bias of the treatment coefficient, measured against the weighted regression that also includes the omitted variable, is fully controlled by two weighted partial $R^2$ values: one for how well the omitted variable explains treatment given the covariates, and one for how well it explains the outcome given treatment and covariates. If true, the same sensitivity analysis works for inverse-propensity, matching, balancing, and stratification weights without modeling how the weights themselves would change. The paper also provides robustness values, bounds by comparison with observed covariates, and bootstrap confidence intervals, implemented in the weightsense R package.","feed_headline":"Two R-squared values pin down weighted regression bias","feed_subtitle":"The bias formula applies to any weighting scheme, from propensity scores to matching weights.","key_machinery":"The central object is the weighted omitted-variable-bias decomposition of Equation (21): the coefficient difference is a product of two weighted partial correlations, $R_w(Y \\sim Z \\mid D,X)$ and $R_w(D \\sim Z \\mid X)$, divided by $\\sqrt{1-R^2_w(D \\sim Z \\mid X)}$, times a ratio of weighted residual standard deviations. Weighted partial $R^2$ is defined as the share of remaining weighted variance in one variable explained by another after partialing out the conditioning variables in the weighted empirical distribution. The additional machinery is a semi-weight benchmarking device: because weighting often makes treatment and covariates nearly uncorrelated, the treatment-side benchmark uses weights estimated without the benchmark covariate, with a translator term converting semi-weighted strength to weighted strength; the bootstrap supplies adjusted inference.","core_discovery":"In the paper's notation, the omitted-variable bias of the weighted least squares estimator relative to the weighted regression that also includes the unobserved variable is $$\\mathrm{dbias}(\\hat{\\tau}_{\\mathrm{wls}}) = R_w(Y \\sim Z \\mid D,X)\\, \\frac{R_w(D \\sim Z \\mid X)}{\\sqrt{1-$R^{2}$_w(D \\sim Z \\mid X)}}\\, \\frac{\\hat{\\mathrm{sd}}_w($Y^{{\\perp_w X,D}}$)}{\\hat{\\mathrm{sd}}_w($D^{{\\perp_w X}}$)}.$$ The only unknown quantities on the right are two weighted partial $R^2$ values: how much of the leftover weighted variation in treatment and in outcome an omitted variable explains after conditioning on the observed covariates. The paper calls these the sensitivity parameters, derives robustness values, extreme-scenario thresholds, and covariate benchmarking bounds from them, and supplies a percentile bootstrap for adjusted confidence intervals.","pith_inferences":["A natural extension would be to invert the bias formula and plot the boundary of the sensitivity-parameter grid on which the adjusted estimate crosses a policy-relevant threshold such as sign, minimum effect, or significance.","When full weights and semi-weights diverge strongly, the translator term can exceed 7; one could automate flagging of such datasets by correlating full and semi-weights, a diagnostic the paper leaves informal.","Because the bias formula makes no distributional assumption on $Z$, it could also cover omitted treatment-by-covariate interactions, turning the sensitivity parameters into statements about heterogeneity rather than mean confounding; this is not worked out in the paper."],"forward_implications":["Any weighting scheme can be assessed with the same two weighted partial $R^2$ parameters, because the bias formula leaves the origin of the weights unspecified.","Investigators can report a robustness value: the equal strength of confounding in treatment and outcome needed to reduce the estimate by a specified fraction or render it insignificant, without assuming a distribution for the omitted variable.","With observed covariates as benchmarks, a covariate as strong as a chosen variable (or a stated multiple) yields formal upper bounds on both sensitivity parameters, so users can see whether plausible confounding would change conclusions.","A percentile bootstrap, with cluster or fixed-weight variants, provides confidence intervals for the adjusted estimate at chosen sensitivity parameter values.","For weights that exactly balance covariate means, the tools apply directly to the weighted difference in means as well as to the weighted regression."],"supporting_citations":[{"why":"Supplies the conditional ignorability assumption and the propensity-score weighting rationale that motivate the estimators studied.","marker":"Rosenbaum and Rubin (1983)"},{"why":"Provides the omitted-variable bias formulas, robustness values, and benchmarking bounds that the paper extends to the weighted setting.","marker":"Cinelli and Hazlett (2020)"},{"why":"Supplies the percentile bootstrap approach adapted for adjusted inference on the sensitivity-adjusted estimate.","marker":"Zhao et al. (2019)"},{"why":"Provides a related percentile bootstrap and interpretable sensitivity parameterization for balancing weights.","marker":"Soriano et al. (2021)"},{"why":"Defines the entropy balancing weights used in the application and in the exact-balance extension to weighted difference in means.","marker":"Hainmueller (2012)"},{"why":"Documents the inconsistency of the standard bootstrap for matching with replacement, motivating the fixed-weight bootstrap variant.","marker":"Abadie and Imbens (2008)"}],"fun_headline_variants":["Two R2 numbers expose weighted regression bias","Weighted regression bias pinned by two R2 values","Bias formula works for any weighting scheme","Sensitivity to confounding in two R2 statistics","Weighted regression: bias from two R2 values"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The bias formula assumes the reference estimator only adds the omitted variable to the weighted outcome regression and leaves the weights unchanged; if real omitted confounding would also change how the weights are constructed, the formula captures only the outcome-model part of the bias.","fun_headline_variants_meta":{"raw":{"variants":["Two R2 numbers expose weighted regression bias","Weighted regression bias pinned by two R2 values","Bias formula works for any weighting scheme","Sensitivity to confounding in two R2 statistics","Weighted regression: bias from two R2 values"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000144,"raw_usage":{"total_tokens":1212,"prompt_tokens":1016,"completion_tokens":196,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":125}},"tokens_in":632,"tokens_out":196,"duration_ms":2717,"temperature":1.0,"reasoning_tokens":125,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:36:14.122643+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate data with a known confounder $Z$, compute the two weighted partial $R^2$ values, and compare the observed difference between the weighted least squares estimate and the weighted regression that adds $Z$ against the right-hand side of Equation (21); any systematic mismatch beyond sampling error refutes the bias formula.","supporting_citations":[],"review_version":2}