REVIEW 3 major objections 2 minor 28 references
Discussion of "Causal and counterfactual views of missing data models" by Razieh Nabi, Rohit Bhattacharya, Ilya Shpitser, & James M. Robins
T0 review · 3 major / 2 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Under the permutation missingness model, the mean of a partially missing variable can be estimated at the parametric rate with local efficiency, and binary outcomes reduce to a posterior odds calculation.
desk verdict A useful discussion with a real new influence function for the permutation MNAR model, but the efficiency claim is under-supported as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the permutation missingness model's pair of nonparametric independence restrictions: $R_1 \perp\!\!\perp Y(1) \mid X(1)$ and $R_2 \perp\!\!\perp (Y(1), X(1)) \mid Y, R_1$. These restrictions identify the full data law and, via Bayes' rule, imply the ratio formula for $\theta$. The analytical engine is the von Mises expansion of the functional $\theta = E(\beta(X)/\alpha(X) \mid R_1 = 0, R_2 = 1)$, whose influence function is derived in Lemma 1; the remainder term is explicitly displayed, making the second-order rate transparent. In the binary case the key simplification is the odds identity $\xi(X) = \lambda(X) \times \text{odds}(Y = 1 \mid R_1 = R_2 = 1)$, where $\lambda$ is a density ratio, so the mean functional becomes the posterior probability of $Y = 1$ from prior odds times likelihood ratio, standardized to the $R_1 = 0, R_2 = 1$ group.
What would settle it
Generate data from a model that satisfies both independence restrictions, compute the proposed one-step estimator with correctly specified nuisance functions on many samples, and check that the standardized estimator is approximately standard normal; a systematic deviation would disprove the influence function.
Extended reading notes
Core claim
Under the permutation missingness model, the discussion establishes that the target mean $\psi = E(Y(1))$ is identified from observed data $(X, R, Y)$ by $\psi = P(R_1 = 1)E(Y \mid R_1 = 1) + P(R_1 = 0)\theta$, where $\theta$ is a ratio of conditional expectations involving $\zeta(Y) = P(R_2 = 1 \mid R_1 = 1, Y)$. Proposition 1 states this identification, and Lemma 1 shows that $\theta$ has the von Mises expansion $\theta(P) - \theta(P) = \int \varphi(o; P) \, d(P - P)(o) + R_\theta(P; P)$ with the displayed influence function $\varphi(O; P)$; this gives the local asymptotic minimax lower bound. For binary $Y$, $\theta = E\{\xi(X)/(\rho + \xi(X)) \mid R_1 = 0, R_2 = 1\}$, where $\xi$ is a conditional odds and $\rho$ an odds ratio, interpreted as a posterior odds combining prior information from the $R_1 = 1$ group with a likelihood ratio from the doubly observed group. Corollary 2 provides the corresponding influence function in the binary case, and the resulting one-step estimator is displayed.
Load-bearing premise
The permutation model assumes that $R_1$ is unrelated to the unobserved outcome given the first covariates, and that $R_2$ is unrelated to both unobserved variables once $Y$ and $R_1$ are known; if either is false, the identifying formula and the efficient estimator are invalid.
Editorial extensions
If this is right
- Under the permutation model, estimating $\psi$ does not require estimating the joint distribution of $(X(1), Y(1))$; the closed-form ratio in Proposition 1 suffices.
- The influence function in Lemma 1 gives the efficiency bound for $\theta$, so the one-step estimator attains local asymptotic minimax optimality.
- With binary $Y$, the estimand has a direct posterior odds reading, making the missing-not-at-random adjustment transparent to practitioners.
- Replacing the unknown odds ratio $\rho$ by its plug-in estimate preserves $\sqrt{n}$-consistency and asymptotic normality, adding only an extra asymptotically linear term.
- SWIGs (m-SWIGs) can deliver counterfactual independence conditions such as $A(1) \perp\!\!\perp (Y^{a(1)}, R^{a(1)}) \mid X$ that are not directly visible from m-DAG d-separation.
Reading between the lines
- The same ratio-of-expectations structure suggests that double/debiased machine learning with flexible nuisance estimators could be applied directly; the discussion only sketches the one-step estimator.
- The posterior odds interpretation implies the estimand is monotone in the prior odds and in the density ratio $\lambda(x)$, which could be used to construct sensitivity bounds under partial violations of the independence assumptions.
- The explicit remainder in the von Mises expansion shows products of second-order nuisance errors, the structure required for rate double robustness; the discussion does not develop this point.
- A natural extension is to models with more than two missingness indicators, where the same permutation-style independence restrictions would yield recursive ratio formulas; the paper does not state this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This discussion paper responds to Nabi et al. (2022) in three directions: (i) it surveys causal-inference tools (instrumental variables, shadow variables, negative controls) that could transfer to missing-data problems; (ii) it introduces the idea of m-SWIGs for simultaneous interventions on treatment and missingness, illustrated on a missing-exposure example; and (iii) it derives identification and estimation results for the permutation missingness model of Robins (1997), which is a primary example in Nabi et al. (2022). Specifically, Proposition 1 gives an identifying expression for ψ = E(Y(1)) in terms of observed-data functionals, Lemma 1 claims a von Mises expansion and influence function for θ = E(Y(1) | R1=0) in the general (non-binary) outcome case, and Corollary 2 gives a simplified influence function and a one-step estimator in the binary-outcome case with known odds ratio ρ. The paper argues that the proposed estimator is sqrt(n)-consistent and asymptotically efficient when ρ is known, and that plugging in a sqrt(n)-consistent estimator of ρ introduces only an extra asymptotically linear term.
Significance. If the efficiency results are correct, the paper provides the first practical semiparametric efficient estimator for a mean functional in the permutation missingness model, a non-standard MNAR model where estimation theory is otherwise underdeveloped. The m-SWIG discussion and the survey of transferable causal tools are useful conceptual contributions that will likely stimulate further work. The identification result in Proposition 1 is derived in detail and is a useful simplification of the full-law identification of Nabi et al. (2022). However, the central new technical contribution—the influence function and asymptotic-efficiency claim—is not fully established: Lemma 1's proof is omitted, and the extension to unknown ρ is only sketched. The efficiency claim therefore rests on unverified calculations rather than a complete theoretical argument. This is a support gap rather than a demonstrated error, and the paper is clearly written and well organized.
major comments (3)
- [Section 4, Lemma 1] Lemma 1 is the load-bearing step for the paper's efficiency claim, but its proof is omitted: the manuscript says 'We omit details, but the result follows from calculations similar to those discussed for example in Section 4 of Kennedy (2022).' The influence function φ(O;P) must satisfy the pathwise derivative identity, and the remainder Rθ(P;P̄) must be o_P(n^{-1/2}) under suitable conditions, but neither is verified. As written, a reader cannot check the correctness of the displayed φ. Please provide a full proof (or at least a detailed derivation) of the von Mises expansion, including the mean-zero property of φ and a bound on the remainder that would make the one-step estimator asymptotically linear.
- [Section 4, Corollary 2 and the one-step estimator] The influence function in Corollary 2 is derived under the assumption that ρ is known, but the proposed estimator then plugs in an estimated ρ̂, with the statement that 'the resulting estimator of θ will just have an extra asymptotically linear term.' No theorem is stated that gives conditions under which the one-step estimator with estimated ρ is asymptotically linear with a known influence function, nor is the extra term characterized. To support the efficiency claim, the authors should state a formal theorem that includes the regularity conditions (e.g., consistency and rate conditions on ξ̂ and ϖ̂, Donsker or empirical-process conditions, and the structure of the extra term from estimating ρ).
- [Appendix, proof of Corollary 2] The appendix derivation uses 'the influence function for ξ(x) when X is discrete.' The main results are presented for arbitrary X, including continuous covariates in the HIV example. For continuous X, the influence function for the conditional odds ξ(x) involves nonparametric estimation over a continuum, and the displayed remainder and one-step estimator require additional conditions (e.g., smoothness, rate conditions, or sample-splitting) that are not provided. This gap limits the generality of the efficiency claim to discrete X unless the authors supply the appropriate continuous-X theory.
minor comments (2)
- [Section 4, Proposition 1] In the displayed expression for ψ, the term θ is used before it is formally defined; please define θ = E(Y(1) | R1=0) immediately before the proposition to avoid confusion.
- [Section 4, Corollary 1] The interpretation of the posterior odds result is helpful, but the notation ζ(Y) is reused from Lemma 1 for a different quantity; consider a distinct symbol (e.g., ρ0(Y)) for the odds ratio to avoid ambiguity.
Circularity Check
No significant circularity: the identification and influence-function claims are either inherited from external prior work or newly stated with explicit remainders, not derived from their own conclusions.
full rationale
I walked the derivation chain from the permutation-model assumptions through Proposition 1, Lemma 1, and Corollary 2. The identifying expression for ψ is inherited from the full-law identification of Nabi et al. (2022), not defined in terms of ψ, and Proposition 1 re-derives it by integration and Bayes' rule. The influence function in Lemma 1 is stated as a new mathematical claim with the remainder explicitly displayed; the proof is omitted and deferred to 'calculations similar to' Kennedy (2022), but this is a support gap, not a circular reduction, since the expansion is not true by construction and no fitted parameter is relabeled as a prediction. Corollary 2 likewise derives a specialized influence function from elementary differentiation and a stated discrete-X influence function for ξ(x). The paper's self-citations (Kennedy 2022 for method; Levis et al. 2024, 2025 for context) are not load-bearing uniqueness claims or definitions of the target. I therefore find no circular step; the appropriate score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Permutation missingness model: R1 is independent of Y(1) given X(1), and R2 is independent of (Y(1), X(1)) given Y and R1
- domain assumption Nonparametric structural equation model or FFRCISTG independence semantics, so d-separation in m-DAGs and m-SWIGs implies counterfactual independence
- domain assumption Consistency and deterministic relations, e.g., A = R A(1) + (1 - R) '?'
- domain assumption Regularity conditions for von Mises expansions and second-order remainders, including existence of densities and nuisance estimators converging at appropriate rates
Cite this review
Pith. "Pith review of Discussion of "Causal and counterfactual views of missing data models" by Razieh Nabi, Rohit Bhattacharya, Ilya Shpitser, & James M. Robins." pith.science (2026). https://pith.science/paper/S3ZSCZID
@misc{pith2026250613025,
author = {Pith},
title = {Pith review of: Discussion of "Causal and counterfactual views of missing data models" by Razieh Nabi, Rohit Bhattacharya, Ilya Shpitser, & James M. Robins},
year = {2026},
howpublished = {\url{https://pith.science/paper/S3ZSCZID}},
note = {Machine review of arXiv:2506.13025}
}
read the original abstract
We congratulate Nabi et al. (2022) on their impressive and insightful paper, which illustrates the benefits of using causal/counterfactual perspectives and tools in missing data problems. This paper represents an important approach to missing-not-at-random (MNAR) problems, exploiting nonparametric independence restrictions for identification, as opposed to parametric/semiparametric models, or resorting to sensitivity analysis. Crucially, the authors represent these restrictions with missing data directed acyclic graphs (m-DAGs), which can be useful to determine identification in complex and interesting MNAR models. In this discussion we consider: (i) how/whether other tools from causal inference could be useful in missing data problems, (ii) problems that combine both missing data and causal inference together, and (iii) some work on estimation in one of the authors' example MNAR models.
Figures
Reference graph
Works this paper leans on
-
[1]
Balke, A. and Pearl, J. (1994), Probabilistic evaluation of counterfactual queries, in Proceedings of the 12th Conference on Artificial Intelligence-Volume 1, Menlo Park, CA: MIT Press, pp. 230--237
work page 1994
-
[2]
--- (1997), Bounds on treatment effects from studies with imperfect compliance, Journal of the American Statistical Association, 92, 1171--1176
1997
-
[3]
Cui, Y., Pu, H., Shi, X., Miao, W., and Tchetgen Tchetgen, E. (2024), Semiparametric proximal causal inference, Journal of the American Statistical Association, 119, 1348--1359
work page 2024
-
[4]
Frangakis, C. E. and Rubin, D. B. (2002), Principal stratification in causal inference, Biometrics, 58, 21--29
work page 2002
-
[5]
Heckman, J. J. (1979), Sample selection bias as a specification error, Econometrica: Journal of the econometric society, 153--161
work page 1979
-
[6]
Hern \'a n, M. A. and Robins, J. M. (2020), Causal inference: what if, Boca Raton: Chapman & Hill/CRC
work page 2020
-
[7]
Imbens, G. W. and Angrist, J. D. (1994), Identification and Estimation of Local Average Treatment Effects, Econometrica, 62, 467--475
1994
-
[8]
Kennedy, E. H. (2020), Efficient nonparametric causal inference with missing exposure information, The International Journal of Biostatistics, 16
work page 2020
Show all 28 references
-
[9]
--- (2022), Semiparametric doubly robust targeted double machine learning: a review, arXiv preprint arXiv:2203.06469
2022 arXiv
-
[10]
W., Bonvini, M., Zeng, Z., Keele, L., and Kennedy, E
Levis, A. W., Bonvini, M., Zeng, Z., Keele, L., and Kennedy, E. H. (2025), Covariate-assisted bounds on causal effects with instrumental variables, Journal of the Royal Statistical Society: Series B (Statistical Methodology), qkaf028
2025
-
[11]
W., Kennedy, E
Levis, A. W., Kennedy, E. H., and Keele, L. (2024), Nonparametric identification and efficient estimation of causal effects with instrumental variables, arXiv preprint arXiv:2402.09332
2024 arXiv
-
[12]
Li, W., Miao, W., and Tchetgen Tchetgen, E. (2023), Non-parametric inference about mean functionals of non-ignorable non-response data without identifying the joint distribution, Journal of the Royal Statistical Society Series B: Statistical Methodology, 85, 913--935
2023
-
[13]
Miao, W., Geng, Z., and Tchetgen Tchetgen, E. J. (2018), Identifying causal effects with proxy variables of an unmeasured confounder, Biometrika, 105, 987--993
2018
-
[14]
T., and Geng, Z
Miao, W., Liu, L., Tchetgen, E. T., and Geng, Z. (2015), Identification, doubly robust estimation, and semiparametric efficiency theory of nonignorable missing data with a shadow variable, arXiv preprint arXiv:1509.02556
2015 arXiv
-
[15]
and Tchetgen Tchetgen, E
Miao, W. and Tchetgen Tchetgen, E. J. (2016), On varieties of doubly robust estimators under missingness not at random with a shadow variable, Biometrika, 103, 475--482
2016
-
[16]
(2022), Causal and counterfactual views of missing data models, arXiv preprint arXiv:2210.05558
Nabi, R., Bhattacharya, R., Shpitser, I., and Robins, J. (2022), Causal and counterfactual views of missing data models, arXiv preprint arXiv:2210.05558
2022 arXiv
-
[17]
B., and Tchetgen Tchetgen, E
Park, C., Richardson, D. B., and Tchetgen Tchetgen, E. J. (2024), Single proxy control, Biometrics, 80, ujae027
2024
-
[18]
(2009), Causality, Cambridge U niversity P ress
Pearl, J. (2009), Causality, Cambridge U niversity P ress
2009
-
[19]
Richardson, T. S. and Robins, J. M. (2013), Single world intervention graphs (SWIGs): A unification of the counterfactual and graphical approaches to causality, Center for the Statistics and the Social Sciences, University of Washington Series. Working Paper, 128, 2013
2013
-
[20]
Robins, J. M. (1997), Non-response models for the analysis of non-monotone non-ignorable missing data, Statistics in medicine, 16, 21--37
1997
-
[21]
Robins, J. M. and Richardson, T. S. (2010), Alternative graphical causal models and the identification of direct effects, Tech. rep., Center for Statistics and the Social Sciences, University of Washington
2010
-
[22]
S., and Robins, J
Shpitser, I., Richardson, T. S., and Robins, J. M. (2022), Multivariate counterfactual systems and causal graphical models, in Probabilistic and causal inference: The works of Judea Pearl, New York, NY: Association for Computing Machinery, pp. 813--852
2022
-
[23]
Sun, B., Liu, L., Miao, W., Wirth, K., Robins, J., and Tchetgen, E. J. T. (2018), Semiparametric estimation with data missing not at random using an instrumental variable, Statistica Sinica, 28, 1965
2018
-
[24]
A., Hern \'a n, M
Swanson, S. A., Hern \'a n, M. A., Miller, M., Robins, J. M., and Richardson, T. S. (2018), Partial identification of the average treatment effect using instrumental variables: review of methods for binary instruments, treatments, and outcomes, Journal of the American Statisti...
2018
-
[25]
(2014), The control outcome calibration approach for causal inference with unobserved confounding, American journal of epidemiology, 179, 633--640
Tchetgen Tchetgen, E. (2014), The control outcome calibration approach for causal inference with unobserved confounding, American journal of epidemiology, 179, 633--640
2014
-
[26]
Tchetgen Tchetgen, E. J. and Wirth, K. E. (2017), A general instrumental variable framework for regression analysis with outcome missing not at random, Biometrics, 73, 1123--1131
2017
-
[27]
and Tchetgen Tchetgen, E
Wang, L. and Tchetgen Tchetgen, E. (2018), Bounded, efficient and multiply robust estimation of average treatment effects using instrumental variables, Journal of the Royal Statistical Society Series B: Statistical Methodology, 80, 531--550
2018
-
[28]
(2012), Doubly robust estimators of causal exposure effects with missing data in the outcome, exposure or a confounder, Statistics in Medicine, 31, 4382--4400
Williamson, E., Forbes, A., and Wolfe, R. (2012), Doubly robust estimators of causal exposure effects with missing data in the outcome, exposure or a confounder, Statistics in Medicine, 31, 4382--4400
2012
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.