REVIEW 2 major objections 4 minor 41 references
Doubly Robust Estimation of Causal Excursion Effects in Micro-Randomized Trials with Missing Longitudinal Outcomes
T0 review · 2 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper proposes the first doubly robust estimator for causal excursion effects in micro-randomized trials when longitudinal outcomes are missing at random, consistent when either the missingness model or the outcome regression is…
desk verdict Solid doubly robust estimator for Δ=1 CEE with missing outcomes, but the general-Δ claim overreaches and needs a corrected outcome regression. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key identity is the augmented missing-data score $\tilde U_t$ in (3.3), which combines the stabilized treatment weight $W_t$, the excursion weight $W_{t,\Delta}$ (a product of likelihood-ratio terms for future treatments under policy $\pi$ versus the randomization policy $p$), and the augmentation term $E\{U_t(\beta,\mu_t,\tilde p_t)\mid H_t,A_t\}$. The augmentation is what converts inverse probability weighting into double robustness: it has mean zero under a correct missingness model and exactly cancels the bias from a wrong missingness model when the outcome regression is correct. Theorems 1 and 2 then provide the asymptotic distribution, with Theorem 2 relying on the product-rate condition $\|\hat e-e^\star\|\,\|\hat\mu-\mu^\star\|=o_p(n^{-1/2})$ for data-adaptive nuisance estimators.
What would settle it
Evaluate the estimating equation at $\Delta=2$ with a future-treatment policy $\pi\ne p$, generating data where $E(W_{t,2}Y_{t,2}\mid H_t,A_t)\ne E(Y_{t,2}\mid H_t,A_t)$, then misspecify the outcome regression while keeping the missingness model correct; if $\hat\beta$ shows bias or confidence intervals under-cover, the stated double-robustness property does not hold as written for delayed effects.
Extended reading notes
Core claim
At the center is the doubling of protection provided by equation (3.3). The proposed estimator weights the CEE score by the inverse probability of observing the outcome, then subtracts the conditional expectation of that weighted score given history and treatment. Because the subtracted term corrects the misspecified nuisance in one direction, the score stays unbiased when the missingness model is correct even if the outcome regression is wrong, and stays unbiased when the outcome regression is correct even if the missingness model is wrong. Theorems 1 and 2 state this as consistency and asymptotic normality with parametric and nonparametric nuisance estimation, respectively, under missing at random and positivity. Simulations with continuous outcomes confirm low bias and near-nominal coverage when either one nuisance model is misspecified, and the HeartSteps analysis shows the estimator in a real 9.4% missingness setting.
Load-bearing premise
The central guarantee rests on missingness being explainable by history and treatment, and on the outcome regression targeting the plain conditional mean even when estimating delayed effects under a different future policy; if either fails, the double-protection property can collapse.
Editorial extensions
If this is right
- Analyses of MRTs with missing-at-random outcomes can replace complete-case analysis or ad hoc imputation with an estimator that is consistent when either the missingness model or the outcome regression is correct.
- Because the same estimating equations handle identity and log links, the method covers continuous, binary, and count outcomes without separate developments.
- Stage-1 nuisance fits can be flexible: with nonparametric estimators, the estimator remains $\sqrt{n}$-consistent and asymptotically normal under the product-rate condition $\|\hat e-e^\star\|\,\|\hat\mu-\mu^\star\|=o_p(n^{-1/2})$.
- Sandwich variance estimators from Theorems 1 and 2 provide Wald confidence intervals with near-nominal coverage in the simulation settings, including when exactly one nuisance model is misspecified.
Reading between the lines
- When the horizon $\Delta>1$ and the future policy $\pi$ differs from the randomization probabilities, the double-robustness argument as written appears to require the excursion-weighted regression $E(W_{t,\Delta}Y_{t,\Delta}\mid H_t,A_t)$ inside the augmentation, while the paper defines the outcome regression as $E(Y_{t,\Delta}\mid H_t,A_t)$; the two coincide only when $W_{t,\Delta}=1$, so the del
- It would be informative to stress-test the estimator in a $\Delta=2$, $\pi\ne p$ simulation where the outcome regression is misspecified and the missingness model is correct: visible bias would confirm that the excursion-weighted target is the necessary augmentation.
- The augmentation term subtracts the conditional mean of the weighted score, so even when missingness is low the method can reduce variance relative to plain inverse-probability weighting; this gain should grow as the missingness rate increases.
- Because the estimator requires missing at random, a natural next step would be to replace the augmentation with a shadow-variable construction to allow missing-not-at-random mechanisms, preserving the same two-stage structure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a two-stage, doubly robust estimator for the causal excursion effect (CEE) in micro-randomized trials when longitudinal outcomes are missing at random. It defines a stabilized inverse-probability-weighted estimating function and augments it with an outcome-regression term in the style of Robins, Rotnitzky, and Zhao. The authors claim that the estimator is consistent and asymptotically normal if either the missingness model or the outcome regression model is correctly specified, for identity or log links, for both parametric and nonparametric nuisance estimation. Simulation results for Δ=1 with continuous outcomes support the claim under correctly specified or one-sided misspecification models, and an application to HeartSteps illustrates the method. The central claim is double robustness for general excursion windows Δ≥1 and arbitrary future treatment regimes π, but the formal development in the main text only verifies the required algebra for the fully observed, unweighted outcome regression, which agrees with the excursion-weighted regression only when Δ=1 or π=p.
Significance. If the general-Δ double-robustness claim is correct, this would be a useful contribution: it would provide a principled alternative to ad hoc imputation for missing outcomes in MRTs, with a simple two-stage construction and explicit asymptotic variance estimators. The paper's Delta=1 results appear internally consistent, the estimating equation algebra for the augmented missing-data term is standard, and the simulation study is carefully designed with multiple generating functions and implementations. The availability of replication code and the inclusion of a real-data comparison against common imputation approaches are strengths. However, the manuscript's headline claim is broader than what is actually proved, because the outcome regression is not defined as the excursion-weighted regression required by the identification result for Δ>1 and non-π=p regimes.
major comments (2)
- [Section 3.1 and Eq. (3.2)-(3.5)] The double-robustness claim for general Δ>1 is not supported by the definitions given. Identification (3.2) expresses the CEE as a contrast of q_t(h,a)=E(W_{t,Δ}Y_{t,Δ}|H_t=h,A_t=a), but Section 3.1 defines the outcome regression as μ*_t(h,a)=E(Y_{t,Δ}|H_t=h,A_t=a), and this unweighted μ is the object estimated in Algorithm 1 and used in the augmenting term of (3.4)-(3.5). These two regressions coincide only when W_{t,Δ}=1, i.e. Δ=1 or π=p. For Δ>1 with a regime such as π=0, the augmentation algebra requires the conditional moment E{Ũ_t(β*, μ*)|H_t,A_t}=0 to hold at the true weighted regression q_t; with μ=m the displayed estimating function contains terms of the form q_a − p f^Tβ − (1−p)m_1 − p m_0, which do not generally vanish. As written, the estimator is therefore likely inconsistent for Δ>1 and π≠p even when the unweighted outcome regression is correctly specified. The manuscript should either restrict all claims to Δ=1, or redefine μ as E(W_{t,Δ}Y_{t,Δ}|H_t,A_t) (or the π-regime conditional mean) and reprove Theorem 1 and Theorem 2 under that definition.
- [Section 5 and Section 6] The numerical evidence does not exercise the general-Δ claim. Section 5 explicitly states that all simulations focus on Δ=1, and the HeartSteps application uses the 30-minute proximal outcome, which also corresponds to Δ=1. There is therefore no simulation or data analysis that checks consistency, coverage, or double robustness for Δ>1 with π≠p. Given the definitional issue in the previous comment, the empirical work cannot distinguish between the claimed general theorem and a Δ=1-only result. At minimum, the abstract and theorems should be restated to the Δ=1 setting, or the manuscript should add a simulation with Δ>1 and a non-π=p regime.
minor comments (4)
- [Abstract] The phrase 'identify or log link' should be 'identity or log link'.
- [Theorem 1] In the definition of Φ(θ, p̃), the second and third blocks are both written as U_e(γ_e); the third block should presumably be the estimating function for the outcome regression nuisance parameters, U_μ(γ_μ). This typo makes the displayed sandwich variance formula ambiguous.
- [Algorithm 1, Stage 1] Stage 1 says to fit Pr(A_t=1|S_t, I_t=1), while Section 3.1 denotes the corresponding nuisance function as p̃_t(S_t); the notation is consistent in substance but the algorithm should also state that this is the stabilized numerator probability, not a model for the true randomization probability.
- [Section 5, Figures 1 and 2] The four-panel figures compress three generating functions and four sample sizes into small panels; labeling each panel with the implementation (A-D) directly in the plot, rather than only in Table 2 and the figure legend, would improve readability.
Circularity Check
No construction-level circularity: the DR estimator is derived from standard semiparametric theory and validated against simulated truth and external HeartSteps data; minor self-citations are building blocks, and the only substantive concern (a Delta>1 definitional mismatch) is a correctness gap, not a circular reduction.
full rationale
The derivation chain is not circular. The target beta_star is defined by the CEE model in (3.1), the fully observed estimating function U_t is an existing CEE score from Boruvka/Qian/Cheng, and the missing-data version U_tilde_t is obtained by the standard augmented inverse-probability-weighting construction of Robins et al. (1994), as shown in (3.3). The double-robustness claim (Assumption 5, Theorems 1 and 2) is a theorem about a moment condition, not a statement that beta is a renamed nuisance fit. Stage 1 nuisance fits are inputs, not the target, and the estimator is evaluated against a simulated truth and an external HeartSteps dataset, so it is not a fitted-input-called-prediction setup. The citations to Cheng et al. (2023) and Qian et al. (2021) are self-citations, but they supply building-block estimating functions for the fully observed case; the present paper's contribution, the missing-data augmentation and rate double-robustness theory, is independent of those citations and is not justified by them alone. I therefore find no circular step. One non-circular correctness risk should be flagged: Section 3.1 defines mu_star_t(h,a) := E(Y_t,Delta | H_t = h, A_t = a), while identification (3.2) requires E(W_t,Delta Y_t,Delta | H_t, A_t). These coincide only when W_t,Delta = 1, i.e., Delta = 1 or pi = p. For Delta > 1 with pi != p, the unweighted outcome regression is not the regression needed by the weighted estimating function, so the general-Delta double-robustness claim is unsubstantiated. This is a definitional mismatch and a correctness gap, not a circular reduction; the paper's own simulations and application use Delta = 1 only.
Assumptions & free parameters
assumptions (5)
- domain assumption Assumption 1: SUTVA, positivity of treatment randomization, and sequential ignorability.
- domain assumption Assumption 2: positivity for missingness and missing at random, Y_t,Delta independent of R_t,Delta given H_t,A_t.
- domain assumption Assumptions 3 and 5: nuisance estimators converge in L2 and at least one of e or mu is correctly specified.
- domain assumption Theorem 2 product-rate condition: ||e_hat-e*|| ||mu_hat-mu*|| = op(n^{-1/2}).
- standard math Assumption 6: compact parameter space, bounded support, invertible Jacobian, Donsker classes, propensity bounded away from 0 and 1.
Cite this review
Pith. "Pith review of Doubly Robust Estimation of Causal Excursion Effects in Micro-Randomized Trials with Missing Longitudinal Outcomes." pith.science (2026). https://pith.science/paper/UERHYNFK
@misc{pith2026241110620,
author = {Pith},
title = {Pith review of: Doubly Robust Estimation of Causal Excursion Effects in Micro-Randomized Trials with Missing Longitudinal Outcomes},
year = {2026},
howpublished = {\url{https://pith.science/paper/UERHYNFK}},
note = {Machine review of arXiv:2411.10620}
}
read the original abstract
Micro-randomized trials (MRTs) are increasingly utilized for optimizing mobile health interventions, with the causal excursion effect (CEE) as a central quantity for evaluating interventions under policies that deviate from the experimental policy. However, MRT often contains missing data due to reasons such as missed self-reports or participants not wearing sensors, which can bias CEE estimation. In this paper, we propose a two-stage, doubly robust estimator for CEE in MRTs when longitudinal outcomes are missing at random, accommodating continuous, binary, and count outcomes. Our two-stage approach allows for both parametric and nonparametric modeling options for two nuisance parameters: the missingness model and the outcome regression. We demonstrate that our estimator is doubly robust, achieving consistency and asymptotic normality if either the missingness or the outcome regression model is correctly specified. Simulation studies further validate the estimator's desirable finite-sample performance. We apply the method to HeartSteps, an MRT for developing mobile health interventions that promote physical activity.
Figures
Reference graph
Works this paper leans on
-
[1]
and Robins, J
Bang, H. and Robins, J. M. (2005). Doubly robust estimation in missing data and causal inference models. Biometrics , 61(4):962--973
2005
-
[2]
Bell, L., Garnett, C., Bao, Y., Cheng, Z., Qian, T., Perski, O., Potts, H. W., Williamson, E., et al. (2023). How notifications affect engagement with a behavior change app: Results from a micro-randomized trial. JMIR mHealth and uHealth , 11(1):e38342
work page 2023
-
[3]
Boruvka, A., Almirall, D., Witkiewitz, K., and Murphy, S. A. (2018). Assessing time-varying causal effect moderation in mobile health. Journal of the American Statistical Association , 113(523):1112--1121
2018
-
[4]
A., and Davidian, M
Cao, W., Tsiatis, A. A., and Davidian, M. (2009). Improving efficiency and robustness of the doubly robust estimator for a population mean with incomplete data. Biometrika , 96(3):723--734
2009
-
[5]
Cheng, Z., Bell, L., and Qian, T. (2023). Efficient and globally robust causal excursion effect estimation. arXiv preprint arXiv:2311.16529
arXiv 2023
-
[6]
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). Double/debiased machine learning for treatment and structural parameters
2018
-
[7]
Dempsey, W., Liao, P., Klasnja, P., Nahum-Shani, I., and Murphy, S. A. (2015). Randomised trials for the fitbit generation. Significance , 12(6):20--23
work page 2015
-
[8]
Dempsey, W., Liao, P., Kumar, S., and Murphy, S. A. (2020). The stratified micro-randomized trial design: sample size considerations for testing nested causal effects of time-varying treatments. The annals of applied statistics , 14(2):661
2020
Show all 41 references
-
[9]
D \' az, I., Carone, M., and van der Laan, M. J. (2016). Second-order inference for the mean of a variable missing at random. The international journal of biostatistics , 12(1):333--349
2016
-
[10]
and Gijbels, I
Fan, J. and Gijbels, I. (1996). Local polynomial modelling and its applications: monographs on statistics and applied probability 66 , volume 66. CRC Press
1996
-
[11]
Horowitz, J. L. (2009). Semiparametric and nonparametric methods in econometrics , volume 12. Springer
2009
-
[12]
Hudgens, M. G. and Halloran, M. E. (2008). Toward causal inference with interference. Journal of the American Statistical Association , 103(482):832--842
2008
-
[13]
Kennedy, E. H. (2016). Semiparametric theory and empirical processes in causal inference. Statistical causal inferences and their applications in public health research , pages 141--167
2016
-
[14]
Kennedy, E. H. (2022). Semiparametric doubly robust targeted double machine learning: a review. arXiv preprint arXiv:2203.06469
2022 arXiv
-
[15]
B., Shiffman, S., Boruvka, A., Almirall, D., Tewari, A., and Murphy, S
Klasnja, P., Hekler, E. B., Shiffman, S., Boruvka, A., Almirall, D., Tewari, A., and Murphy, S. A. (2015). Microrandomized trials: An experimental design for developing just-in-time adaptive interventions. Health Psychology , 34(S):1220
2015
-
[16]
J., Lee, A., Hall, K., Luers, B., Hekler, E
Klasnja, P., Smith, S., Seewald, N. J., Lee, A., Hall, K., Luers, B., Hekler, E. B., and Murphy, S. A. (2019). Efficacy of contextually tailored suggestions for physical activity: a micro-randomized optimization trial of heartsteps. Annals of Behavioral Medicine , 53(6):573--582
2019
-
[17]
K., Lessler, J., and Stuart, E
Lee, B. K., Lessler, J., and Stuart, E. A. (2011). Weight trimming and propensity score weighting. PloS one , 6(3):e18174
2011
-
[18]
Lok, J. J. (2024). How estimating nuisance parameters can reduce the variance (with consistent variance estimation). Statistics in Medicine , 43(23):4456--4480
2024
-
[19]
J., and Geng, Z
Miao, W., Liu, L., Li, Y., Tchetgen Tchetgen, E. J., and Geng, Z. (2024). Identification and semiparametric efficiency theory of nonignorable missing data with a shadow variable. ACM/JMS Journal of Data Science , 1(2):1--23
2024
-
[20]
Murphy, S. A. (2003). Optimal dynamic treatment regimes. Journal of the Royal Statistical Society Series B: Statistical Methodology , 65(2):331--355
2003
-
[21]
E., Collins, L
Qian, T., Walton, A. E., Collins, L. M., Klasnja, P., Lanza, S. T., Nahum-Shani, I., Rabbi, M., Russell, M. A., Walton, M. A., Yoo, H., et al. (2022). The microrandomized trial for developing digital interventions: Experimental design and data analysis considerations. Psycholo...
2022
-
[22]
Qian, T., Xiang, S., and Cheng, Z. (2023). MRTAnalysis: Primary and Secondary Analyses for Micro-Randomized Trials . R package version 0.1.2
2023
-
[23]
Qian, T., Yoo, H., Klasnja, P., Almirall, D., and Murphy, S. A. (2021). Estimating time-varying causal excursion effects in mobile health with binary outcomes. Biometrika , 108(3):507--527
2021
-
[24]
Robins, J. (1986). A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect. Mathematical modelling , 7(9-12):1393--1512
1986
-
[25]
Robins, J. M. (2004). Optimal structural nested models for optimal sequential decisions. In Proceedings of the Second Seattle Symposium in Biostatistics: analysis of correlated data , pages 189--326. Springer
2004
-
[26]
M., Rotnitzky, A., and Zhao, L
Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1994). Estimation of regression coefficients when some regressors are not always observed. Journal of the American statistical Association , 89(427):846--866
1994
-
[27]
Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of educational Psychology , 66(5):688
1974
-
[28]
P., and Carroll, R
Ruppert, D., Wand, M. P., and Carroll, R. J. (2003). Semiparametric regression . Cambridge University Press
2003
-
[29]
O., Rotnitzky, A., and Robins, J
Scharfstein, D. O., Rotnitzky, A., and Robins, J. M. (1999). Adjusting for nonignorable drop-out using semiparametric nonresponse models. Journal of the American Statistical Association , 94(448):1096--1120
1999
-
[30]
J., Smith, S
Seewald, N. J., Smith, S. N., Lee, A. J., Klasnja, P., and Murphy, S. A. (2019). Practical considerations for data collection and management in mobile health micro-randomized trials. Statistics in biosciences , 11:355--370
2019
-
[31]
and Dempsey, W
Shi, J. and Dempsey, W. (2023). A meta-learning method for estimation of causal excursion effects to assess time-varying moderation. arXiv preprint arXiv:2306.16297
2023 arXiv
-
[32]
Shi, J., Wu, Z., and Dempsey, W. (2023). Assessing time-varying causal effect moderation in the presence of cluster-level treatment effect heterogeneity and interference. Biometrika , 110(3):645--662
2023
-
[33]
Smucler, E., Rotnitzky, A., and Robins, J. M. (2019). A unifying approach for doubly-robust l_1 regularized estimation of causal contrasts. arXiv preprint arXiv:1904.03737
2019 arXiv
-
[34]
Tan, Z. (2010). Bounded, efficient and doubly robust estimation with inverse weighting. Biometrika , 97(3):661--682
2010
-
[35]
Tsiatis, A. A. (2006). Semiparametric theory and missing data , volume 4. Springer
2006
-
[36]
van der Laan, M. (2017). A generally efficient targeted minimum loss based estimator based on the highly adaptive lasso. The international journal of biostatistics , 13(2)
2017
-
[37]
J., Rose, S., Zheng, W., and van der Laan, M
van der Laan, M. J., Rose, S., Zheng, W., and van der Laan, M. J. (2011). Cross-validated targeted minimum-loss-based estimation. Targeted learning: causal inference for observational and experimental data , pages 459--474
2011
-
[38]
and Vansteelandt, S
Vermeulen, K. and Vansteelandt, S. (2015). Bias-reduced doubly robust estimation. Journal of the American Statistical Association , 110(511):1024--1036
2015
-
[39]
E., and Thall, P
Wang, L., Rotnitzky, A., Lin, X., Millikan, R. E., and Thall, P. F. (2012). Evaluation of viable dynamic treatment regimes in a sequentially randomized trial of advanced prostate cancer. Journal of the American Statistical Association , 107(498):493--508
2012
-
[40]
Wood, S. N. (2017). Generalized additive models: an introduction with R . CRC press
2017
-
[41]
Yang, S., Wang, L., and Ding, P. (2019). Causal inference with confounders missing not at random. Biometrika , 106(4):875--888
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.