Pith. sign in

REVIEW 4 major objections 4 minor 49 references

Generalizing causal effects with noncompliance: Application to deep canvassing experiments

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper proves that the target complier average causal effect is identified as a ratio of two target-population averages, and gives estimators and sensitivity bounds.

desk verdict Worthwhile transported CACE paper with a fixable but load-bearing formal gap in the exclusion restriction assumption. read the letter →

arxiv 2506.00149 v2 pith:VDPTGAVN submitted 2025-05-30 stat.ME

classification stat.ME MSC 62D2062F12
keywords targetcomplieraveragecausaleffectgeneralizabilityinstrumentalvariablesnoncomplianceinverseprobabilityweightingmultiplyrobustestimationsensitivityanalysisdeepcanvassing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks how to estimate the effect of a treatment on people who comply with it when the study sample does not look like the target population. It introduces the target complier average causal effect (T-CACE), the average effect among compliers in the target population, and proves that under four assumptions this effect is identified as a ratio of two target-population averages: the mean outcome difference between assigned-treatment and assigned-control groups divided by the mean difference in treatment received. The key move is a population-level exclusion restriction that lets a randomized encouragement inside the study act as an instrument in the target population, without needing principal ignorability. The paper then supplies inverse-weighted, weighted-least-squares, and multiply robust estimators, derives their asymptotic normality, and adds a sensitivity analysis for unmeasured confounding. A reanalysis of a deep canvassing experiment shows the method can generalize the complier effect to participants who skipped the follow-up survey, while effects on non-participants remain uncertain.

What carries the argument

The central object is the target complier average causal effect (T-CACE), defined as $E[Y(1)-Y(0) \mid C=1, S=0]$, the effect among compliers in the target population. The machinery is the Wald ratio reformulation: T-CACE equals the ratio of the generalized intent-to-treat effect to the generalized first stage, each expressed as a target-covariate average of study-sample conditional means. That ratio turns noncompliance into a denominator that must also be generalized, and Assumption 4 (mean exchangeability of the first stage) is what makes the denominator transportable.

What would settle it

If one could obtain a random sample from the target population and measure, under an encouragement design, the actual difference in treatment receipt between those assigned to encouragement and those not, then comparing the target average of $E[D(1)-D(0)|X]$ with the study-sample analog would directly test Assumption 4; a large discrepancy would imply the proposed T-CACE estimators are biased.

Watch

Extended reading notes

Core claim

The central claim is Theorem 3.1: under treatment ignorability, mean exchangeability of treatment effect heterogeneity, mean exchangeability of the first stage, and monotonicity plus valid-instrument assumptions, the target complier average causal effect is identified as $\tau_{T-CACE} = \frac{E_{X|S=0}[\mu_{y1}(X)-\mu_{y0}(X)]}{E_{X|S=0}[\mu_{d1}(X)-\mu_{d0}(X)]}$, where the numerator is the generalized intent-to-treat effect and the denominator is the generalized first stage. Identification follows because the exclusion restriction lets potential outcomes depend only on treatment received, so the complier effect equals the Wald ratio in the target population, and the two mean-exchangeability assumptions replace target-population outcome and compliance means by study-sample conditional means averaged over the target covariate distribution. This reframes generalization with noncompliance as a weighted instrumental-variables problem rather than a principal-stratification problem.

Load-bearing premise

The load-bearing premise is that, after conditioning on observed covariates, the average compliance response to the encouragement is the same in the target population as in the study sample.

Editorial extensions

If this is right

  • Estimating the T-CACE reduces to computing two weighted means, so standard inverse-probability-of-sampling weights can be adapted for noncompliance.
  • All three proposed estimators are consistent and asymptotically normal, and the weighted-least-squares version can reduce the variance inflation caused by reweighting.
  • When partial compliance information is available in the target population, the first-stage exchangeability assumption can be relaxed or checked against data.
  • The sensitivity analysis produces a partial-identification interval whose width grows with the sensitivity parameter $\Gamma$, giving a concrete threshold at which the conclusion would reverse.
  • Applied to deep canvassing, the method finds evidence that the complier effect on reducing exclusionary attitudes generalizes to participants who did not complete the follow-up survey, but not to registered voters who never entered the experiment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because identification is a Wald ratio, the estimators likely inherit weak-instrument sensitivity: when target compliance is rare, the denominator is near zero and even unbiased estimates will be unstable.
  • A testable extension would use target-population proxies of compliance, such as prior-year treatment uptake, to bound or validate the first-stage exchangeability assumption before trusting the T-CACE.
  • The multiply robust construction points toward a doubly robust ratio estimator in which only the outcome model or only the first-stage model needs to be correct, which the paper's Theorem 4.4 already hints at.
  • The same identification strategy may extend to cluster-encouragement designs with partial compliance, as long as the exclusion restriction holds at the cluster level.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper develops methods for transporting the complier average causal effect (CACE) from a study sample to a target population when there is noncompliance. It defines the target complier average causal effect (T-CACE), identifies it under treatment ignorability, IV assumptions (monotonicity, exclusion, relevance), and two mean exchangeability assumptions (Assumptions 2 and 4), and expresses it as the ratio of two target-population expectation terms (Theorem 3.1). The paper then proposes inverse-probability weighted, weighted least squares, and multiply robust estimators with consistency and asymptotic normality results, considers settings with partial compliance information, and introduces a Rosenbaum-style sensitivity analysis implemented through linear programming. The methods are evaluated in extensive simulations and applied to a deep canvassing experiment on reducing exclusionary attitudes.

Significance. If the identification and proofs were fully correct, the paper would make a useful contribution: it provides an alternative to principal-ignorability-based transportability of the CACE, with a simple Wald-ratio identification formula and standard M-estimator machinery. The estimation strategy is straightforward to implement in RCTs with known treatment assignment probabilities, and the multiply robust estimator and sensitivity analysis are practically valuable additions. The paper's strengths include the clear identification formula, the extensive simulation comparisons that include violations of principal ignorability and of the exclusion restriction, and the real-data application. However, the current version contains a formal gap in the statement of the exclusion restriction, an ambiguity in the definition of the WLS weights, an unsupported asymptotic claim for observational studies, and an incomplete linear-programming reformulation of the sensitivity bounds. These issues are fixable but affect the results as stated.

major comments (4)
  1. [§3, Assumption 3(b), Theorem 3.1, Appendix B.1] The formal statement of the exclusion restriction is too weak for the identification theorem. As written, Assumption 3(b) requires Y(z,D(z))=Y(z',D(z')) only when D(z)=D(z'), so under monotonicity it is vacuous for compliers. The proof of Theorem 3.1 uses the condition to discard always-takers and never-takers and then equates the remaining complier contrast with the treatment-received effect, which requires the pointwise exclusion Y(z,d)=Y(z',d) for all z,z',d. The prose in Section 3 states the strong version, so this is a formal statement/proof gap rather than a conceptual one, but it is load-bearing: under the weak version the ratio in (3.2) identifies an assignment effect rather than the treatment-received CACE. For example, with perfect compliance and Y=D+λZ, all stated assumptions hold but the Theorem 3.1 ratio equals 1+λ, not the CACE of 1. The formal assumption and the proof in Appendix B.1 should be corrected to the strong exclusion restriction.
  2. [§4.2, Eq. (4.6), Appendix C.5] The weighted least squares estimator is defined in (4.6) with a single weight pw_i(X_i), but the consistency proof in Appendix C.5 treats the weight as w_1(X_i) for treated units and w_0(X_i) for control units and then uses identities for E[S w_1 Z X^T] and E[S w_0(1-Z) X^T]. Under the displayed definition with one weight, these identities do not follow, and the claim that the standard weighted estimator is a special case of the WLS estimator with intercept-only covariates holds only when w_1=w_0, which fails in general observational settings. The authors should either clarify that the WLS weight is w_{Z_i}(X_i) (i.e., w_1(X_i) for Z_i=1 and w_0(X_i) for Z_i=0) and adjust (4.6) accordingly, or supply the correct consistency proof for a single weight.
  3. [§4.1, Assumption 6, Theorem 4.3, §F.3] The abstract and introduction claim the method applies to observational studies, but the asymptotic theory for the weighted estimator assumes that the treatment assignment probability P(Z=1|S=1,X) is known (Assumption 6; Theorems 4.1-4.3). Section F.3 presents an observational simulation in which Z is generated from a logistic model depending on X, but the paper does not state how this treatment propensity is estimated or how its estimation is incorporated into the variance calculations. To support the observational claim, the authors should add an explicit treatment-propensity estimation step to the M-estimator and derive the resulting asymptotic distribution, or restrict the formal claims to randomized experiments with known assignment probabilities and present the observational extension as heuristic.
  4. [§5.2, Eq. (5.4), Appendix E.3] The Charnes-Cooper reformulation of the sensitivity bounds is incomplete. After defining \bar r = t r and imposing the normalization \sum \bar r_i pw_i = 1, the constraints listed are t>0, Γ^{-1} ≤ r_i ≤ Γ, and the normalization, but no constraint links \bar r_i to r_i and t. As written, the objective is a function of \bar r alone while the box constraints are on r alone, so the two optimization problems are not equivalent to the fractional program in (E.1), and the interval in (5.4) is not guaranteed to contain the true T-CACE. The linking constraints (e.g., \bar r_i = t r_i, or equivalently Γ^{-1} t ≤ \bar r_i ≤ Γ t) should be added and the equivalence proof supplied.
minor comments (4)
  1. [§5.1] The text refers to 'mean exchangeability of the first stage—Assumption 2'; the first-stage exchangeability is Assumption 4, while Assumption 2 is mean exchangeability of treatment effect heterogeneity.
  2. [Appendix A.1] The sentence 'Theorem 4.3 states that pτwls is consistent' appears to be a cross-reference error; consistency of the WLS estimator is established in Theorem A.1, not Theorem 4.3.
  3. [Table 1] The multiply robust estimator's coverage is reported as low as 84.6% (N=10000, ratio 0.55), which is a noticeable undercoverage; the text's statement that coverage 'fluctuates around the nominal level' is too generous and should be qualified.
  4. [§G.1] The sensitivity analysis in Section 5.2 is formalized for the weighted estimator, but Figure 7 applies it to the WLS estimator. The paper should either provide the formal extension to the WLS estimator or state explicitly that the displayed sensitivity intervals pertain to the weighted estimator and are shown for the WLS application only as an approximation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the T-CACE identification and estimators are self-contained, and self-citations are contextual rather than load-bearing.

full rationale

The central identification claim (Theorem 3.1) takes the target estimand as E[Y(1)-Y(0)|C=1,S=0] and derives the Wald-ratio expression E[Y(1)-Y(0)|S=0]/E[D(1)-D(0)|S=0] from stated assumptions: treatment ignorability (Assumption 1), mean exchangeability of treatment effect heterogeneity (Assumption 2), monotonicity/valid instrument/exclusion (Assumption 3), and mean exchangeability of the first stage (Assumption 4). None of these assumptions is defined in terms of the fitted estimators or of the estimand itself. The weighted, WLS, and multiply robust estimators are conventional M-estimators or doubly robust estimators whose consistency proofs (Theorems 4.1, A.1, 4.4) reduce to moment conditions derived from Theorem 3.1; there is no fitted parameter that is renamed as a prediction. The sensitivity analysis uses a user-specified Gamma and a marginal sensitivity model following Zhao et al. (2019), not an outcome-fitted quantity. Self-citations (e.g., Huang et al. 2023; Huang 2024a,b; Huang and Pimentel 2025; Hartman and Huang 2023; Chen et al. 2025a,b) appear mainly as context, efficiency discussion, or calibration suggestions, and none is load-bearing for the identification or consistency results. One formal gap exists that is not circular: Assumption 3(b) as written, Y(z,D(z))=Y(z',D(z')) only for z,z' with D(z)=D(z'), is vacuous for compliers, while the surrounding prose asserts the stronger claim that the instrument affects the outcome only through D. This is an assumption-strength and interpretation gap regarding whether the identified quantity is the treatment-received CACE, not a reduction of the conclusion to the inputs. The paper's own sensitivity analysis explicitly treats Assumptions 2 and 4 as tenuous, which is an honest limitation rather than a circular step.

Assumptions & free parameters 3 free parameters · 7 assumptions · 0 invented entities

The identification result rests on exchangeability and IV assumptions; no invented entities are introduced. The free parameters are nuisance selection model coefficients, the sensitivity parameter Γ, and nuisance regression functions for the multiply robust estimator.

free parameters (3)
  • Selection model coefficients β (logistic regression for P(S=1|X)) = Estimated by MLE from combined study and target covariate data
    Nuisance parameters in the inverse probability weights; Assumption 6 specifies the logistic model.
  • Sensitivity parameter Γ = User-specified (e.g., Γ* = 1.95 in the application)
    Defines the marginal sensitivity model (Assumption 9); not fitted to the outcome but chosen by the analyst.
  • Outcome and treatment-received regression functions μ_yz and μ_dz = Estimated by linear regression in simulations; method unspecified in general
    Nuisance functions used in the multiply robust estimator; consistency requires correct specification of either the weights or these models.
assumptions (7)
  • domain assumption Treatment ignorability (Assumption 1): Z independent of potential outcomes and potential treatment given X within the study.
    Holds by design in RCTs; needed to identify conditional outcome and treatment means from the study sample.
  • domain assumption Mean exchangeability of selection and treatment effect heterogeneity (Assumption 2): E[Y(1)-Y(0)|X,S=1] = E[Y(1)-Y(0)|X,S=0].
    Core assumption for transporting the numerator of the T-CACE ratio to the target population.
  • domain assumption Monotonicity, exclusion restriction, and instrument relevance (Assumption 3): no defiers, assignment affects outcome only through treatment received, and the first stage is nonzero in the target.
    Standard IV assumptions, here required to hold at the population level for both study and target.
  • domain assumption Mean exchangeability of the first stage (Assumption 4): E_{X|S=0}E[D(1)-D(0)|S=0,X] = E_{X|S=0}E[D(1)-D(0)|S=1,X].
    Identifies the denominator of the T-CACE ratio; fragile when compliance patterns differ after conditioning on X.
  • domain assumption Overlap (Assumption 5): 0 < P(S=1|X) and 0 < P(Z=1|S=1,X) almost surely.
    Necessary for inverse probability weighting to be well defined.
  • domain assumption Logistic selection model (Assumption 6): P(S=1|X) = σ(β^T X), and P(Z=1|S=1,X) is known.
    Used for the asymptotic distribution of the weighted estimator when the selection model is unknown; the known treatment propensity holds only in RCTs.
  • domain assumption Marginal sensitivity model (Assumption 9): the odds ratio of the oracle weights to estimable weights is bounded by Γ.
    Underlies the proposed sensitivity analysis; Γ is user-specified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Generalizing causal effects with noncompliance: Application to deep canvassing experiments." pith.science (2026). https://pith.science/paper/VDPTGAVN

@misc{pith2026250600149,
  author       = {Pith},
  title        = {Pith review of: Generalizing causal effects with noncompliance: Application to deep canvassing experiments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VDPTGAVN}},
  note         = {Machine review of arXiv:2506.00149}
}
read the original abstract

Standard approaches in generalizability often focus on generalizing the intent-to-treat (ITT). However, in practice, a more policy-relevant quantity is the generalized impact of an intervention across compliers. While instrumental variable (IV) methods are commonly used to estimate the complier average causal effect (CACE) within samples, standard approaches cannot be applied to a target population with a different distribution from the study sample. This paper makes several key contributions. First, we introduce a new set of identifying assumptions in the form of a population-level exclusion restriction that allows for identification of the target complier average causal effect (T-CACE) in both randomized experiments and observational studies. This allows researchers to identify the T-CACE without relying on standard principal ignorability assumptions. Second, we propose a class of inverse-weighted estimators for the T-CACE and derive their asymptotic properties. We provide extensions for settings in which researchers have access to auxiliary compliance information across the target population. Finally, we introduce a sensitivity analysis for researchers to evaluate the robustness of the estimators in the presence of unmeasured confounding and extend existing tests to evaluate instrument validity in this context. We illustrate our proposed method through extensive simulations and a study evaluating the impact of deep canvassing on reducing exclusionary attitudes.

Figures

Figures reproduced from arXiv: 2506.00149 by the authors.

Figure 1
Figure 1. Boxplots of the bias of the weighted estimator, WLS estimator, multiply robust estimator, [PITH_FULL_IMAGE:figures/full_fig_p024_1.png] view at source ↗
Figure 2
Figure 2. Plot of various estimates and 95% confidence intervals for the deep canvassing applica [PITH_FULL_IMAGE:figures/full_fig_p027_2.png] view at source ↗
Figure 3
Figure 3. Comparison of bias of the WLS estimator and the PS estimator when principal ignorability [PITH_FULL_IMAGE:figures/full_fig_p057_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Comparison of bias of the WLS estimator and the PS estimator when exclusion restriction [PITH_FULL_IMAGE:figures/full_fig_p059_4.png]
Figure 5
Figure 5. Figure 5: Plot of the estimated T-CACE with respect to the sensitivity parameter [PITH_FULL_IMAGE:figures/full_fig_p064_5.png]
Figure 6
Figure 6. Figure 6: Plot of Γ for each of the 6 selected covariates, arranged in increasing order. The [PITH_FULL_IMAGE:figures/full_fig_p065_6.png]
Figure 7
Figure 7. Figure 7: Plot of the estimated T-CACE with respect to the sensitivity parameter [PITH_FULL_IMAGE:figures/full_fig_p066_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

49 extracted references · 42 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type url volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := #1 'mid.sentence := ...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Angrist, J. D. , Imbens, G. W. and Rubin, D. B. (1996). Identification of causal effects using instrumental variables. Journal of the American statistical Association, 91 444--455

  4. [4]

    Aronow, P. M. and Green, D. P. (2013). Sharp bounds for complier average potential outcomes in experiments with noncompliance and incomplete reporting. Statistics & Probability Letters, 83 677--679

  5. [5]

    Aronow, P. M. and Lee, D. K. (2013). Interval estimation of population means under unknown but bounded probabilities of sample selection. Biometrika, 100 235--240

  6. [6]

    and Wager, S

    Athey, S. and Wager, S. (2019). Estimating treatment effects with causal forests: An application. Observational studies, 5 37--51

  7. [7]

    Buchanan, A. L. , Hudgens, M. G. , Cole, S. R. , Mollan, K. R. , Sax, P. E. , Daar, E. S. , Adimora, A. A. , Eron, J. J. and Mugavero, M. J. (2018). Generalizing evidence from randomized trials using inverse probability of sampling weights. Journal of the Royal Statistical Society Series A: Statistics in Society, 181 1193--1209

  8. [8]

    Carlin, B. P. and Gelfand, A. E. (1990). Approaches for empirical bayes confidence intervals. Journal of the American Statistical Association, 85 105--114

Show all 49 references
  1. [9]

    Carroll, R. J. , Ruppert, D. , Stefanski, L. A. and Crainiceanu, C. M. (2006). Measurement error in nonlinear models: a modern perspective. Chapman and Hall/CRC

  2. [10]

    , Kalla, J

    Chen, Z. , Kalla, J. , Le, Q. , Nakamura-Sakai, S. , Sekhon, J. and Wang, R. (2025 a ). A framework to assess the persuasion risks large language model chatbots pose to democratic societies. arXiv preprint arXiv:2505.00036

  3. [11]

    , Tian, L

    Chen, Z. , Tian, L. and Olshen, R. A. (2025 b ). An empirical bayes approach for constructing confidence intervals for clonality and entropy. Journal of Applied Statistics 1--18

  4. [12]

    , Chetverikov, D

    Chernozhukov, V. , Chetverikov, D. , Demirer, M. , Duflo, E. , Hansen, C. , Newey, W. and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters

  5. [13]

    Clark, J. M. , Rott, K. W. , Hodges, J. S. and Huling, J. D. (2024). Transportability of principal causal effects. arXiv preprint arXiv:2405.04419

  6. [14]

    Cole, S. R. and Stuart, E. A. (2010). Generalizing evidence from randomized clinical trials to target populations: the actg 320 trial. American journal of epidemiology, 172 107--115

  7. [15]

    Cooper, A. C. W. et al. (1962). Programming with linear fractional functionals. Naval Research logistics quarterly, 9 181--186

  8. [16]

    Crump, R. K. , Hotz, V. J. , Imbens, G. W. and Mitnik, O. A. (2009). Dealing with limited overlap in estimation of average treatment effects. Biometrika, 96 187--199

  9. [17]

    Dahabreh, I. J. , Robertson, S. E. and Hern \'a n, M. A. (2022). Generalizing and transporting inferences about the effects of treatment assignment subject to non-adherence. arXiv preprint arXiv:2211.04876

  10. [18]

    Deaton, A. S. (2009). Instruments of development: Randomization in the tropics, and the search for the elusive keys to economic development. Tech. rep., National bureau of economic research

  11. [19]

    and Lu, J

    Ding, P. and Lu, J. (2017). Principal stratification analysis using principal scores. Journal of the Royal Statistical Society Series B: Statistical Methodology, 79 757--777

  12. [20]

    and Hartman, E

    Egami, N. and Hartman, E. (2023). Elements of external validity: Framework, design, and analysis. American Political Science Review, 117 1070--1088

  13. [21]

    , Mealli, F

    Feller, A. , Mealli, F. and Miratrix, L. (2017). Principal score methods: Assumptions, extensions, and practical considerations. Journal of Educational and Behavioral Statistics, 42 726--758

  14. [22]

    Freedman, D. A. and Berk, R. A. (2008). Weighting regressions by propensity scores. Evaluation review, 32 392--409

  15. [23]

    Funk, M. J. , Westreich, D. , Wiesen, C. , St \"u rmer, T. , Brookhart, M. A. and Davidian, M. (2011). Doubly robust estimation of causal effects. American journal of epidemiology, 173 761--767

  16. [24]

    , Grieve, R

    Hartman, E. , Grieve, R. , Ramsahai, R. and Sekhon, J. S. (2015). From sample average treatment effect to population average treatment effect on the treated: combining experimental with observational studies to estimate population treatment effects. Journal of the Royal Statis...

  17. [25]

    and Huang, M

    Hartman, E. and Huang, M. (2023). Improving precision through design and analysis in experiments with noncompliance. Political Science Research and Methods 1--16

  18. [26]

    Hotz, V. J. , Imbens, G. and Mortimer, J. H. (1999). Predicting the efficacy of future training programs using past experiences

  19. [27]

    (2024 a )

    Huang, M. (2024 a ). Overlap violations in external validity. arXiv preprint arXiv:2403.19504

  20. [28]

    , Egami, N

    Huang, M. , Egami, N. , Hartman, E. and Miratrix, L. (2023). Leveraging population outcomes to improve the generalization of experimental results: Application to the jtpa study. The Annals of Applied Statistics, 17 2139--2164

  21. [29]

    and Pimentel, S

    Huang, M. and Pimentel, S. D. (2025). Variance-based sensitivity analysis for weighting estimators results in more informative bounds. Biometrika, 112 asae040

  22. [30]

    Huang, M. Y. (2024 b ). Sensitivity analysis for the generalization of experimental results. Journal of the Royal Statistical Society Series A: Statistics in Society qnae012

  23. [31]

    Huber, P. J. (1992). Robust estimation of a location parameter. In Breakthroughs in statistics: Methodology and distribution. Springer, 492--518

  24. [32]

    Huber, P. J. et al. (1967). The behavior of maximum likelihood estimates under nonstandard conditions. In Proceedings of the fifth Berkeley symposium on mathematical statistics and probability, vol. 1. Berkeley, CA: University of California Press

  25. [33]

    Jo, B. (2002). Estimation of intervention effects with noncompliance: Alternative model specifications. Journal of Educational and Behavioral Statistics, 27 385--409

  26. [34]

    and Stuart, E

    Jo, B. and Stuart, E. A. (2009). On the use of propensity scores in principal causal effect estimation. Statistics in medicine, 28 2857--2875

  27. [35]

    Kalla, J. L. and Broockman, D. E. (2020). Reducing exclusionary attitudes through interpersonal conversation: Evidence from three field experiments. American Political Science Review, 114 410--425

  28. [36]

    Kern, H. L. , Stuart, E. A. , Hill, J. and Green, D. P. (2016). Assessing methods for generalizing experimental impact estimates to target populations. Journal of research on educational effectiveness, 9 103--127

  29. [37]

    and Babu, G

    Li, B. and Babu, G. J. (2019). A graduate course on statistical inference. Springer

  30. [38]

    , Thomas, L

    Li, F. , Thomas, L. E. and Li, F. (2019). Addressing extreme propensity scores via the overlap weights. American journal of epidemiology, 188 250--257

  31. [39]

    , Chen, Z

    Liu, Q. , Chen, Z. and Wong, W. H. (2024). An encoding generative modeling approach to dimension reduction and covariate adjustment in causal inference with observational studies. Proceedings of the National Academy of Sciences, 121 e2322376121

  32. [40]

    Miratrix, L. W. , Sekhon, J. S. , Theodoridis, A. G. and Campos, L. F. (2018). Worth weighting? how to think about and use weights in survey experiments. Political Analysis, 26 275--291

  33. [41]

    , Imbens, G

    Nie, X. , Imbens, G. and Wager, S. (2021). Covariate balancing sensitivity analysis for extrapolating randomized trials across locations. arXiv preprint arXiv:2112.04723

  34. [42]

    Ottoboni, K. N. and Poulos, J. V. (2020). Estimating population average treatment effects from experiments with noncompliance. Journal of Causal Inference, 8 108--130

  35. [43]

    Robins, J. M. , Rotnitzky, A. and Zhao, L. P. (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association, 89 846--866

  36. [44]

    Rosenbaum, P. R. (1987). Sensitivity analysis for certain permutation inferences in matched observational studies. Biometrika, 74 13--26

  37. [45]

    Rosenbaum, P. R. and Rubin, D. B. (1983). Assessing sensitivity to an unobserved binary covariate in an observational study with binary outcome. Journal of the Royal Statistical Society: Series B (Methodological), 45 212--218

  38. [46]

    Rudolph, K. E. and Laan, M. J. (2017). Robust estimation of encouragement design intervention effects transported across sites. Journal of the Royal Statistical Society Series B: Statistical Methodology, 79 1509--1525

  39. [47]

    , Rippel, O

    Snoek, J. , Rippel, O. , Swersky, K. , Kiros, R. , Satish, N. , Sundaram, N. , Patwary, M. , Prabhat, M. and Adams, R. (2015). Scalable bayesian optimization using deep neural networks. In International conference on machine learning. PMLR

  40. [48]

    Sovey, A. J. and Green, D. P. (2011). Instrumental variables estimation in political science: A readers’ guide. American Journal of Political Science, 55 188--200

  41. [49]

    , Small, D

    Zhao, Q. , Small, D. S. and Bhattacharya, B. B. (2019). Sensitivity analysis for inverse probability weighting estimators via the percentile bootstrap. Journal of the Royal Statistical Society Series B: Statistical Methodology, 81 735--761

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.