Pith. sign in

REVIEW 3 major objections 4 minor 46 references

Sensitivity analysis methods for outcome missingness using substantive-model-compatible multiple imputation and their application in causal inference

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read When the outcome causes its own missingness, imputation compatibility still matters—two delta-adjusted approaches avoid the bias of naive NARFCS.

desk verdict A solid, useful extension of compatible MI to delta-adjustment sensitivity analysis, with an honest simulation study and one unproven compatibility step that deserves a careful response. read the letter →

arxiv 2411.13829 v1 pith:3DWVY7AC submitted 2024-11-21 stat.ME

classification stat.ME
keywords multipleimputationsensitivityanalysismissingnotatrandomdelta-adjustmentsubstantivemodelcompatibilitycausalinferenceg-computationNARFCS
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Multiple imputation is only trustworthy when the imputation model is compatible with the analysis model, and this paper extends that requirement to sensitivity analysis for missing-not-at-random outcomes. The paper proposes two methods, NAR-SMCFCS and NAR-SMC-stack, that bring delta-adjustment for outcome missingness into two existing substantive-model-compatible imputation frameworks. In simulations based on a real cohort study, the proposed methods estimate the average causal effect with relative bias below 6% even when the substantive model includes strong exposure-confounder interactions, whereas a naive implementation of NARFCS shows relative bias as large as -31.8%. The central claim is that ignoring compatibility in sensitivity analysis can make the results of the sensitivity analysis itself unreliable.

What carries the argument

The central object is the compatible target distribution for imputing each incomplete non-outcome variable $V_j$, proportional to $f(Y|X,Z_1,Z_2,\theta)\,f(V_j|V_{-j},S,\lambda_j)$, where $f(Y|X,Z_1,Z_2,\theta)$ is the substantive outcome model. NAR-SMCFCS samples from this target within a chained-equations algorithm; NAR-SMC-stack imputes $V_j$ from the proposal $f(V_j|V_{-j},S,\lambda_j)$ and then assigns each stacked record an importance weight proportional to $f(Y|X,Z_1,Z_2,M_Y=0,\theta')$. The outcome itself is imputed from the delta-adjusted model $f(Y|X,Z_1,Z_2,M_Y,\theta',\delta)$, whose fixed sensitivity parameters $\delta$ shift the imputed values for records with $M_Y=1$. This division of labour is what keeps the non-outcome imputations compatible with an interaction-containing substantive model while allowing the outcome's own missingness to be reflected in the imputation.

What would settle it

Generate data under the paper's missingness DAG but add a direct arrow from the outcome missingness indicator $M_Y$ to a non-outcome variable $V$, so that $f(V|Y, M_Y=1)$ differs from $f(V|Y)$, while keeping strong exposure-confounder interactions in the outcome model; if NAR-SMCFCS or NAR-SMC-stack then shows average causal effect relative bias above 10%, the core compatibility claim would be falsified.

Watch

Extended reading notes

Core claim

Under the missingness mechanism where the outcome causes its own missingness, the average causal effect is not recoverable, so analysts must assess how conclusions change under alternative assumptions. The paper argues that the standard sensitivity-analysis tool, NARFCS, is misspecified when the substantive analysis uses g-computation with exposure-confounder interactions, because its univariate imputation models omit those interactions. The proposed NAR-SMCFCS and NAR-SMC-stack methods impute the outcome from a delta-adjusted model, $f(Y|X,Z_1,Z_2,M_Y,\theta',\delta)$, while keeping the substantive outcome model inside the target distribution used to impute the non-outcome variables. In the simulation scenario with strong exposure-confounder interactions, naive NARFCS gave mean ACE estimates of 0.21 and 0.20 (relative bias -30.6% and -31.8%), whereas the proposed approaches kept relative bias below 6% and achieved near-nominal coverage for NAR-SMCFCS. The same pattern appeared for binary outcomes, and a case study illustrated that naive NARFCS can make an ACE estimate look more sensitive to missingness assumptions than it really is.

Load-bearing premise

The load-bearing premise is that the delta-adjusted outcome model fitted to observed outcomes, together with the unadjusted substantive model used in the target distribution for non-outcome variables, remains a compatible description for records with missing outcomes; the paper does not prove this and its simulation may not exercise violations of it.

Editorial extensions

If this is right

  • In sensitivity analyses for outcome missingness, using NAR-SMCFCS or NAR-SMC-stack removes imputation incompatibility as a source of bias, so the sensitivity curve reflects the assumed delta values rather than misspecification of the imputation model.
  • With strong exposure-confounder interactions, naive NARFCS can misstate the average causal effect by roughly a third even when the true sensitivity parameter is used, which means conclusions drawn from such sensitivity analyses may be wrong in either direction.
  • NAR-SMCFCS achieves approximately nominal 95% coverage when the true delta is used, while NAR-SMC-stack slightly undercovers (around 90-93%), indicating a lingering variance-estimation issue for the stacked approach.
  • The proposed methods are not limited to g-computation; because compatibility is built into the target distributions, they apply to any substantive analysis based on a regression model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct extension would replace exposure-confounder interactions with quadratic terms, since a quadratic term is an interaction of a variable with itself; the same target-distribution construction should remove the bias that a main-effects NARFCS would introduce there.
  • In a real application with modest interaction strength, the difference between naive and compatible NARFCS may be small, but with a strong interaction whose delta effect has the opposite sign, naive NARFCS could reverse the direction of the sensitivity conclusion.
  • Because NAR-SMC-stack undercovers consistently across all scenarios, practitioners should prefer NAR-SMCFCS for reporting confidence intervals until the stacked variance estimator is improved.
  • The simulation generates missingness from a selection model while delta-adjustment is a pattern-mixture assumption, so the 'true' delta values are only approximate; a simulation that generates data directly from the pattern-mixture model would provide a cleaner test of the methods.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper addresses sensitivity analysis for outcome missingness in causal inference settings where the outcome causes its own missingness and the target estimand (the average causal effect, ACE) is non-recoverable. It extends two multiple-imputation approaches, SMCFCS and SMC-stack, by incorporating delta-adjustment for the outcome, yielding NAR-SMCFCS and NAR-SMC-stack. The proposed methods are intended to maintain compatibility between the imputation models and a substantive analysis that includes exposure-confounder interactions, as in g-computation. The paper evaluates the methods in a simulation study motivated by the Victorian Adolescent Health Cohort Study, across continuous and binary outcomes, simple and complex outcome-missingness mechanisms, and null/weak/strong exposure-confounder interactions. It reports that a naive NARFCS implementation is biased in the interaction scenarios, while the proposed approaches produce approximately unbiased ACE estimates, with NAR-SMC-stack showing some under-coverage. The methods are also applied to the case study.

Significance. If the compatibility properties claimed for the new approaches can be established or empirically bounded, the paper would fill a practical gap: it offers a way to perform delta-adjustment sensitivity analysis for g-computation with multivariable missingness, where standard NARFCS is incompatible when exposure-confounder interactions are in the substantive model. The simulation study is carefully tied to a real cohort, the results are reported transparently, and the variance shortfall of NAR-SMC-stack is explicitly acknowledged rather than hidden. The paper also builds on existing published approaches with available implementations, which supports reproducibility. However, the central unbiasedness claim rests on a compatibility step in NAR-SMCFCS that is not formally justified, and the simulation's 'true' sensitivity parameters are acknowledged approximations obtained from a misspecified model; these issues weaken the strength of the conclusions as currently stated.

major comments (3)
  1. [Section 5.3 and Section 7] The compatibility claim for NAR-SMCFCS is not justified. In step 7, each non-outcome variable Vj is imputed from the target distribution f(Y|X,Z1,Z2,theta)f(Vj|V-j,S,lambda_j), with theta estimated in step 5 from the current imputed dataset. But in step 8 the missing outcomes are imputed from the delta-shifted model f(Y|X,Z1,Z2,MY,theta',delta). The theta fitted in step 5 therefore describes a distribution of Y averaged over MY, not conditional on MY. The pattern-mixture conditional that is compatible with the outcome imputation model is f(Vj|...,Y,MY) proportional to f(Y|...,MY,theta',delta)f(Vj|...,MY,lambda), which conditions on MY. Unless f(Y|V,MY=0)=f(Y|V,MY=1) after integrating out the missingness model, which is not generally true when delta is nonzero, the step-7 target is not the compatible conditional for records with MY=1. The paper provides no argument that the bias from this mismatch is negligible, and the simulation, generated under a selection model with a misspecified outcome imputation model, does not isolate this discrepancy. Since the approximate unbiasedness claim for NAR-SMCFCS depends on this step, this is a load-bearing gap.
  2. [Section 5.4, Table 2] The simulation's 'true' sensitivity parameter values are not true data-generating deltas. They are obtained by fitting the outcome imputation model (5), which the paper acknowledges may be misspecified, to a large generated complete dataset from the selection model. The Discussion correctly states that these values represent only an approximation to the closest possible pattern-mixture representation. Consequently, evaluating the methods at these values and at multiples of them does not establish approximate unbiasedness with respect to the actual missingness mechanism; it establishes behavior along a projection of that mechanism. I would like to see either a pattern-mixture data-generating process in which delta is exactly the shift in the outcome model, or an explicit assessment of how large the projection error is. Without this, the headline claim that the proposed methods are 'approximately unbiased' at the true sensitivity parameters is overstated.
  3. [Section 7] The variance estimator for NAR-SMC-stack produces systematically below-nominal coverage, with values around 90-93% across scenarios and model-based standard errors that are substantially smaller than those from NAR-SMCFCS (e.g., 0.009-0.103 versus 0.108-0.112 for continuous outcomes). The paper notes this issue and cites importance-sampling literature, but it still concludes in Section 7 that both proposed approaches perform well. Given that coverage is a central operating characteristic for a sensitivity analysis method, and given that one of the two proposed approaches is the subject of the paper's recommendation, this needs more than a remark: either a corrected variance estimator, a clear statement that NAR-SMC-stack should be used only after further development of its variance estimator, or a substantive justification that the under-coverage is acceptable in this setting.
minor comments (4)
  1. [Section 5.3] The text says each method was implemented by setting delta_0 to the true value, twice the true value, and zero, but the complex missingness scenario has two sensitivity parameters, delta_0 and delta_1. Please clarify whether delta_1 was fixed at its true value while delta_0 was varied, and how this is reflected in Figure 2.
  2. [Section 4.1, Algorithm 1] The notation in steps 5 and 8 makes the distinction between theta and theta' difficult to follow. A sentence clarifying that theta in step 5 is the substantive-model parameter fitted to the current imputed dataset (including delta-shifted imputed outcomes) and that theta' in step 8 is the identifiable part of the delta-adjusted outcome model would help the reader.
  3. [Section 4.2, equation (3)] The weight formula is clear in substance, but the denominator uses a sum over m without making the index on the numerator explicit. Please write the numerator as f(Y_i | X_i^m, Z_{1i}, Z_{2i}^m, M_{Yi}=0, theta') to avoid ambiguity about which imputation m is being weighted.
  4. [General] The abbreviations MAR-SMCFCS and MAR-SMC-stack are introduced in Section 3.4 after SMCFCS and SMC-stack have already been used; defining both variants once in Section 1 or at first use would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the proposed estimators are benchmarked against an externally generated simulation truth and self-citations are not load-bearing.

full rationale

The paper's main claims are supported by a simulation study in which data are generated from explicit outcome and missingness models (equations (4) and the logistic missingness mechanisms of Section 5.2), with the true ACE fixed independently. The proposed NAR-SMCFCS and NAR-SMC-stack algorithms use the substantive outcome model as a factor in their imputation targets, which is the standard compatibility construction rather than a circular equation: no fitted parameter is renamed as a prediction, and the unbiasedness result is not equivalent to an input by construction. The sensitivity parameters delta are estimated from a large generated dataset and explicitly described as an approximation to the closest pattern-mixture representation, which is a simulation calibration device, not a data-driven prediction of the estimand. The paper contains self-citations, notably to Zhang et al. (2024) for non-recoverability and to Tompsett et al. (2018) for the behavior of NARFCS, but the non-recoverability result is also cited to independent work by Mohan and Pearl, and the compatibility guarantees are supported by published methods and prior simulations. The potential concern raised by a skeptical reader about Algorithm 1's use of a theta marginalized over MY is a misspecification or correctness question about the target distribution, not a circularity: it does not make any equation reduce to itself, and it does not involve fitting a parameter to the quantity being predicted. Overall, the derivation chain is self-contained and externally benchmarked, so no significant circularity is present.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The methods rest on prior compatibility theory, the standard delta-adjustment framework, and the assumption of a correctly specified substantive outcome model. The only data-fitted quantities in the evaluation are the sensitivity parameters, which are estimated from the data-generation model in the simulation rather than elicited independently, a limitation the authors acknowledge.

free parameters (2)
  • delta_0 (sensitivity parameter for outcome missingness intercept) = Set to 0, true value, or 2x true value in simulations; true value estimated by fitting model (5) to a large generated…
    The simulation's 'true' delta is not defined directly by the selection-model generator, but is obtained by fitting the potentially misspecified outcome imputation model to a large dataset. Performance of all three methods depends on this choice.
  • delta_1 (sensitivity parameter for exposure interaction in outcome missingness) = Used in complex missingness scenarios; value estimated similarly
    Same as delta_0; appears only when the missingness mechanism includes an exposure-outcome interaction (complex scenario).
assumptions (5)
  • domain assumption The substantive outcome model f(Y|X,Z1,Z2,theta) is correctly specified (Section 3.1).
    The paper states this assumption at the start; compatibility and unbiasedness claims rely on it.
  • domain assumption The missingness mechanism is described by the m-DAG in Figure 1, where the outcome causes its own missingness and no non-outcome variables cause their own missingness.
    Focus of the paper; all methods and simulations are built around this mechanism.
  • standard math The compatibility theory of SMCFCS (Bartlett et al. 2015) and SMC-stack (Beesley and Taylor 2021) is valid, including importance-sampling weights and Beesley's variance rule.
    Taken from cited published methodology with software; the paper extends these methods instead of re-deriving them.
  • domain assumption The delta-adjustment framework, where including the outcome missingness indicator with a fixed sensitivity parameter captures the outcome-missingness association, is an appropriate model.
    This is a pattern-mixture modeling choice; the paper uses it for sensitivity analysis.
  • ad hoc to paper In the simulation, the 'true' delta values estimated by fitting the outcome imputation model to a large generated dataset approximate the closest pattern-mixture representation of the selection-model mechanism.
    The authors acknowledge that the delta-adjusted imputation model is misspecified under the selection model; this assumption is needed to interpret the simulation bias results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sensitivity analysis methods for outcome missingness using substantive-model-compatible multiple imputation and their application in causal inference." pith.science (2026). https://pith.science/paper/3DWVY7AC

@misc{pith2026241113829,
  author       = {Pith},
  title        = {Pith review of: Sensitivity analysis methods for outcome missingness using substantive-model-compatible multiple imputation and their application in causal inference},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3DWVY7AC}},
  note         = {Machine review of arXiv:2411.13829}
}
read the original abstract

When using multiple imputation (MI) for missing data, maintaining compatibility between the imputation model and substantive analysis is important for avoiding bias. For example, some causal inference methods incorporate an outcome model with exposure-confounder interactions that must be reflected in the imputation model. Two approaches for compatible imputation with multivariable missingness have been proposed: Substantive-Model-Compatible Fully Conditional Specification (SMCFCS) and a stacked-imputation-based approach (SMC-stack). If the imputation model is correctly specified, both approaches are guaranteed to be unbiased under the "missing at random" assumption. However, this assumption is violated when the outcome causes its own missingness, which is common in practice. In such settings, sensitivity analyses are needed to assess the impact of alternative assumptions on results. An appealing solution for sensitivity analysis is delta-adjustment using MI, specifically "not-at-random" (NAR)FCS. However, the issue of imputation model compatibility has not been considered in sensitivity analysis, with a naive implementation of NARFCS being susceptible to bias. To address this gap, we propose two approaches for compatible sensitivity analysis when the outcome causes its own missingness. The proposed approaches, NAR-SMCFCS and NAR-SMC-stack, extend SMCFCS and SMC-stack, respectively, with delta-adjustment for the outcome. We evaluate these approaches using a simulation study that is motivated by a case study, to which the methods were also applied. The simulation results confirmed that a naive implementation of NARFCS produced bias in effect estimates, while NAR-SMCFCS and NAR-SMC-stack were approximately unbiased. The proposed compatible approaches provide promising avenues for conducting sensitivity analysis to missingness assumptions in causal inference.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

46 extracted references · 45 canonical work pages

  1. [1]

    Causal Inference: What If

    Robins JM Hernán MA. Causal Inference: What If . Chapman & Hall/CRC, Boca Ratonn, 2020

  2. [2]

    Katherine J Lee, Kate M Tilling, Rosie P Cornish, Roderick JA Little, Melanie L Bell, Els Goetghebeur, Joseph W Hogan, James R Carpenter, et al. Framework for the treatment and reporting of missing data in observational studies: The treatment and reporting of missing data in observational studies framework.Journal of clinical epidemiology, 134:79–88, 2021

  3. [3]

    Recoverability and estimation of causal effects under typical multivariable missingness mechanisms

    Jiaxin Zhang, S Ghazaleh Dashti, John B Carlin, Katherine J Lee, and Margarita Moreno-Betancur. Recoverability and estimation of causal effects under typical multivariable missingness mechanisms. Biometrical Journal, 66(3):2200326, 2024

  4. [4]

    Multiple imputation for nonresponse in surveys, volume 81

    Donald B Rubin. Multiple imputation for nonresponse in surveys, volume 81. John Wiley & Sons, 2004

  5. [5]

    Multiple imputation of discrete and continuous data by fully con- ditional specification

    Stef Van Buuren. Multiple imputation of discrete and continuous data by fully con- ditional specification. Statistical methods in medical research, 16(3):219–242, 2007

  6. [6]

    Multiple imputation and its application

    James R Carpenter, Jonathan W Bartlett, Tim P Morris, Angela M Wood, Matteo Quartagno, and Michael G Kenward. Multiple imputation and its application. John Wiley & Sons, 2023

  7. [7]

    Multiple-imputation inferences with uncongenial sources of input

    Xiao-Li Meng. Multiple-imputation inferences with uncongenial sources of input. Statistical Science, pages 538–558, 1994

  8. [8]

    On the stationary distribution of iterative imputations

    Jingchen Liu, Andrew Gelman, Jennifer Hill, Yu-Sung Su, and Jonathan Kropko. On the stationary distribution of iterative imputations. Biometrika, 101(1):155–173, 2014

Show all 46 references
  1. [9]

    Multiple imputation of covariates by fully conditional specification: accommodating the substantive model

    Jonathan W Bartlett, Shaun R Seaman, Ian R White, James R Carpenter, and Alzheimer’s Disease Neuroimaging Initiative*. Multiple imputation of covariates by fully conditional specification: accommodating the substantive model. Statisti- cal methods in medical research, 24(4):46...

  2. [10]

    A stacked approach for chained equations multiple imputation incorporating the substantive model

    Lauren J Beesley and Jeremy MG Taylor. A stacked approach for chained equations multiple imputation incorporating the substantive model. Biometrics, 77(4):1342– 1354, 2021

  3. [11]

    miss- ing at random

    Shaun Seaman, John Galati, Dan Jackson, and John Carlin. What is meant by “miss- ing at random”? Statistical Science, 28(2):257–268, 2013

  4. [12]

    Canonical causal diagrams to guide the treatment of missing data in epidemiologic studies

    Margarita Moreno-Betancur, Katherine J Lee, Finbarr P Leacy, Ian R White, Julie A Simpson, and John B Carlin. Canonical causal diagrams to guide the treatment of missing data in epidemiologic studies. American journal of epidemiology, 187(12): 2705–2715, 2018

  5. [13]

    Assumptions and analysis planning in studies with missing data in multiple vari- ables: moving beyond the mcar/mar/mnar classification

    Katherine J Lee, John B Carlin, Julie A Simpson, and Margarita Moreno-Betancur. Assumptions and analysis planning in studies with missing data in multiple vari- ables: moving beyond the mcar/mar/mnar classification. International Journal of Epidemiology, page dyad008, 2023

  6. [14]

    Graphical models for inference with missing data

    Karthika Mohan, Judea Pearl, and Jin Tian. Graphical models for inference with missing data. Advances in neural information processing systems, 26, 2013. 25

  7. [15]

    On the use of the not-at-random fully conditional specification (narfcs) procedure in practice

    Daniel Mark Tompsett, Finbarr Leacy, Margarita Moreno-Betancur, Jon Heron, and Ian R White. On the use of the not-at-random fully conditional specification (narfcs) procedure in practice. Statistics in medicine, 37(15):2338–2353, 2018

  8. [16]

    Multiple imputation under missing not at random assumptions via fully conditional specification

    Finbarr P Leacy. Multiple imputation under missing not at random assumptions via fully conditional specification. PhD thesis, University of Cambridge, 2016

  9. [17]

    A general method for elicitation, imputation, and sensitivity analysis for incomplete repeated binary data

    Daniel Tompsett, Stephen Sutton, Shaun R Seaman, and Ian R White. A general method for elicitation, imputation, and sensitivity analysis for incomplete repeated binary data. Statistics in Medicine, 39(22):2921–2935, 2020

  10. [18]

    Implementation of g- computation on a simulated data set: demonstration of a causal inference technique

    Jonathan M Snowden, Sherri Rose, and Kathleen M Mortimer. Implementation of g- computation on a simulated data set: demonstration of a causal inference technique. American journal of epidemiology, 173(7):731–738, 2011

  11. [19]

    Cannabis use and mental health in young people: cohort study

    George C Patton, Carolyn Coffey, John B Carlin, et al. Cannabis use and mental health in young people: cohort study. Bmj, 325(7374):1195–1198, 2002

  12. [20]

    The manual of cis-r

    G Lewis and AJ Pelosi. The manual of cis-r. London: Institute of Psychiatry, 1992

  13. [21]

    A new approach to causal inference in mortality studies with a sus- tained exposure period—application to control of the healthy worker survivor effect

    James Robins. A new approach to causal inference in mortality studies with a sus- tained exposure period—application to control of the healthy worker survivor effect. Mathematical modelling, 7(9-12):1393–1512, 1986

  14. [22]

    Graphical models for recovering probabilistic and causal queries from missing data

    Karthika Mohan and Judea Pearl. Graphical models for recovering probabilistic and causal queries from missing data. Advances in Neural Information Processing Systems, 27, 2014

  15. [23]

    Formalizing subjective notions about the effect of nonrespondents in sample surveys

    Donald B Rubin. Formalizing subjective notions about the effect of nonrespondents in sample surveys. Journal of the American Statistical Association , 72(359):538– 543, 1977

  16. [24]

    Recent developments in the prevention and treatment of missing data

    Craig Mallinckrodt, J Roger, C Chuang-Stein, Geert Molenberghs, M O’Kelly, B Ratitch, M Janssens, and P Bunouf. Recent developments in the prevention and treatment of missing data. Therapeutic Innovation & Regulatory Science , 48(1): 68–80, 2014

  17. [25]

    Appropriate inclusion of interactions was needed to avoid bias in multiple imputation

    Kate Tilling, Elizabeth J Williamson, Michael Spratt, Jonathan AC Sterne, and James R Carpenter. Appropriate inclusion of interactions was needed to avoid bias in multiple imputation. Journal of clinical epidemiology, 80:107–115, 2016

  18. [26]

    How should variable selection be performed with multiply imputed data? Statistics in medicine , 27(17):3227– 3246, 2008

    Angela M Wood, Ian R White, and Patrick Royston. How should variable selection be performed with multiply imputed data? Statistics in medicine , 27(17):3227– 3246, 2008

  19. [27]

    Accounting for not-at-random missingness through imputation stacking

    Lauren J Beesley and Jeremy MG Taylor. Accounting for not-at-random missingness through imputation stacking. Statistics in Medicine, 40(27):6118–6132, 2021

  20. [28]

    Bootstrap inference for multiple im- putation under uncongeniality and misspecification

    Jonathan W Bartlett and Rachael A Hughes. Bootstrap inference for multiple im- putation under uncongeniality and misspecification. Statistical methods in medical research, 29(12):3533–3546, 2020

  21. [29]

    R: A Language and Environment for Statistical Computing

    R Core Team. R: A Language and Environment for Statistical Computing. R Founda- tion for Statistical Computing, Vienna, Austria, 2019. URL https://www.R-project. org/. 26

  22. [30]

    The impact of non-response bias due to sampling in public health studies: A comparison of voluntary versus mandatory recruitment in a dutch national survey on adolescent health

    Kei Long Cheung, Peter M Ten Klooster, Cees Smit, Hein de Vries, and Marcel E Pieterse. The impact of non-response bias due to sampling in public health studies: A comparison of voluntary versus mandatory recruitment in a dutch national survey on adolescent health. BMC public ...

  23. [31]

    Multiple imputation using chained equations: issues and guidance for practice

    Ian R White, Patrick Royston, and Angela M Wood. Multiple imputation using chained equations: issues and guidance for practice. Statistics in medicine, 30(4): 377–399, 2011

  24. [32]

    Identification in missing data models represented by directed acyclic graphs

    Rohit Bhattacharya, Razieh Nabi, Ilya Shpitser, and James M Robins. Identification in missing data models represented by directed acyclic graphs. In Uncertainty in Artificial Intelligence, pages 1149–1158. PMLR, 2020

  25. [33]

    Full law identification in graph- ical models of missing data: Completeness results

    Razieh Nabi, Rohit Bhattacharya, and Ilya Shpitser. Full law identification in graph- ical models of missing data: Completeness results. In International Conference on Machine Learning, pages 7153–7163. PMLR, 2020

  26. [34]

    Adjustment criteria for recovering causal effects from missing data

    Mojdeh Saadati and Jin Tian. Adjustment criteria for recovering causal effects from missing data. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pages 561–577. Springer, 2019

  27. [35]

    Causal inference with confounders missing not at random

    Shu Yang, Linbo Wang, and Peng Ding. Causal inference with confounders missing not at random. Biometrika, 106(4):875–888, 2019

  28. [36]

    On varieties of doubly robust estimators under missingness not at random with a shadow variable

    Wang Miao and Eric J Tchetgen Tchetgen. On varieties of doubly robust estimators under missingness not at random with a shadow variable. Biometrika, 103(2):475– 482, 2016

  29. [37]

    An instrumental variable approach for identification and estimation with nonignorable nonresponse

    Sheng Wang, Jun Shao, and Jae Kwang Kim. An instrumental variable approach for identification and estimation with nonignorable nonresponse. Statistica Sinica, pages 1097–1116, 2014

  30. [38]

    A new instrumental method for dealing with endogenous selection

    Xavier d’Haultfoeuille. A new instrumental method for dealing with endogenous selection. Journal of Econometrics, 154(1):1–15, 2010

  31. [39]

    Identifiability of subgroup causal effects in randomized experiments with nonignorable missing covariates

    Peng Ding and Zhi Geng. Identifiability of subgroup causal effects in randomized experiments with nonignorable missing covariates. Statistics in Medicine , 33(7): 1121–1133, 2014

  32. [40]

    Semiparametric inference of causal effect with nonig- norable missing confounders

    Zhaohan Sun and Lan Liu. Semiparametric inference of causal effect with nonig- norable missing confounders. Statistica Sinica, 31(4):1669–1688, 2021

  33. [41]

    R vignettes: smcfcs, 2022

    Jonathan W Bartlett. R vignettes: smcfcs, 2022. URL https://cran.r-project.org/ web/packages/smcfcs/vignettes/smcfcs-vignette.html

  34. [42]

    A cautious note on auxiliary variables that can increase bias in missing data problems

    Felix Thoemmes and Norman Rose. A cautious note on auxiliary variables that can increase bias in missing data problems. Multivariate Behavioral Research, 49(5): 443–459, 2014

  35. [43]

    Multiple imputation of missing data under missing at random: including a collider as an auxiliary variable in the imputation model can induce bias

    Elinor Curnow, Kate Tilling, Jon E Heron, Rosie P Cornish, and James R Carpenter. Multiple imputation of missing data under missing at random: including a collider as an auxiliary variable in the imputation model can induce bias. medRxiv, pages 2023–06, 2023. 27

  36. [44]

    The common structure of statistical models of truncation, sample selection and limited dependent variables and a simple estimator for such models

    James J Heckman. The common structure of statistical models of truncation, sample selection and limited dependent variables and a simple estimator for such models. In Annals of economic and social measurement, volume 5, number 4 , pages 475–492. NBER, 1976

  37. [45]

    Importance sampling

    Art B Owen. Importance sampling. Monte Carlo theory, methods and examples, 9, 2013

  38. [46]

    Using simulation studies to evaluate statistical methods

    Tim P Morris, Ian R White, and Michael J Crowther. Using simulation studies to evaluate statistical methods. Statistics in medicine, 38(11):2074–2102, 2019

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.