REVIEW 3 major objections 5 minor 37 references
Efficient estimation of optimal regimes under a no direct effect assumption
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Under the no-direct-effect-of-testing assumption, projecting g-estimating functions onto mean-zero testing-treatment residuals yields more efficient, doubly robust estimators of optimal testing and treatment regimes.
desk verdict The core projection-based variance reduction is real and well proven, but the paper's headline near-optimality claim for continuous outcomes rests on an unproved convergence step that should be flagged in the review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the ortho-complement $\Lambda_{\mathrm{NDE}}^{\perp}$ of the tangent space of the NDE model. The paper represents it as $\{T_b = \sum_{t=0}^{K} T_{b,t}\}$, where $T_{b,t}$ is the residual from projecting $D_{b,t} = b_t(\bar H_t,\bar S_t,Y^d)\,W_{t+1}^{-1}(A_t - \mathrm{E}[A_t\mid \bar H_t,S_t])$ onto the space of testing and treatment scores. Here $W_{t+1}$ is the product of treatment probabilities, and the NDE restriction makes every $T_b$ have mean zero. Subtracting the population least-squares projection $c_{\mathrm{OLS}}T_b$ of an influence function $U$ onto this space yields a new influence function, and the Pythagorean theorem guarantees the variance reduction. For feasibility, the paper projects onto a large subspace $\Omega$ built from $b_t^\ast = (\phi_1(Y^d)I_t^{\top},\ldots,\phi_\xi(Y^d)I_t^{\top})^{\top}$, with $I_t$ the vector of indicators of full treatment histories; Theorem 4 supplies recursive least-squares coefficients for the closed-form projection.
What would settle it
Estimate $\mathrm{E}[D_{b,t}]$ from data with variation in testing and treatment, where $D_{b,t} = b_t(\bar H_t,\bar S_t,Y^d)\,W_{t+1}^{-1}(A_t - \mathrm{E}[A_t\mid \bar H_t,S_t])$; a nonzero mean for some $b_t$, or a difference in mean outcome between randomized testing arms with treatment held fixed, refutes the NDE assumption and implies the efficiency gains rest on a false premise.
Extended reading notes
Core claim
Under the NDE assumption, the paper constructs estimators $\tilde\Psi(q,b)$ that solve $0 = \hat U(q,b,\Psi)$, where $\hat U(q,b,\Psi)$ is the residual from projecting the influence function of the usual opt-SNMM estimating function $\hat U(q,\Psi)$ onto the space $T_b$ of mean-zero random variables implied by NDE. Theorem 3 shows each such estimator is regular asymptotically linear with asymptotic variance $V^{\mathrm{oracle}}(q,b) = J^{-1}\{\operatorname{var}[U(q,\Psi)] - c_{\mathrm{OLS}}\mathrm{E}[T_b T_b^\top]c_{\mathrm{OLS}}^\top\}J^{-\top}$, which is no larger, and strictly smaller whenever $\mathrm{E}[U(q,\Psi^\ast)T_b] \neq 0$, than the variance of the standard g-estimator. The projected estimators remain doubly robust, and the paper provides a closed-form feasible construction based on a large subspace spanned by basis functions of the health outcome and treatment-history indicators, whose efficiency approaches the intractable optimal projection as the basis grows. Simulations and an HIV monitoring application indicate gains that can amount to a roughly 50-fold reduction in variance.
Load-bearing premise
The load-bearing premise is that testing has no direct effect on the health outcome, meaning a test changes outcomes only through the treatment it triggers; if testing affects outcomes directly, the mean-zero property of the residual variables fails and the projected estimators become biased.
Editorial extensions
If this is right
- The projected estimators are never less efficient than standard opt-SNMM g-estimators, and strictly more efficient whenever the initial estimating function correlates with the NDE-implied residuals.
- The efficiency gain translates into a sample-size reduction: in the HIV monitoring application cited by the paper, exploiting NDE was associated with a roughly 50-fold variance reduction.
- The same projection recipe applies to other doubly robust RAL estimators, including estimators of dynamic marginal structural models, so the improvement is not tied to opt-SNMMs.
- A feasible, closed-form implementation exists by projecting onto a basis-expanded subspace; its efficiency approaches the semiparametrically optimal projected estimator as the basis dimension grows.
- Under stronger NDE variants, such as no direct effect on covariates or on latent test results, larger projection spaces are available, so the estimators are at least as efficient and typically more efficient than under NDE on the health outcome alone.
Reading between the lines
- An immediate extension is to use the same projected quantities as a specification test: the NDE assumption implies infinitely many mean-zero restrictions, so a data-driven check of whether empirical projections are near zero could precede the efficiency-gaining analysis.
- The construction suggests a general principle for causal inference: any domain assumption that generates extra mean-zero variables can be converted into precision by projecting influence functions onto the ortho-complement of the implied tangent space, not just the NDE assumption considered here.
- In cost-benefit decisions near a value-of-information threshold, the variance reduction could change conclusions: the paper's simulation shows the projected estimator raising the empirical rejection rate of 'screening is not cost-effective' from around 25 percent to 100 percent, so policy conclusions may flip.
- At the boundary where everyone is tested, the NDE assumption identifies parameters that are otherwise unidentified; whether the optimal regime can still be computed there is left open, so a natural extension is to study computation and estimation under such positivity failures.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops semiparametric estimators of the parameters of an optimal regime structural nested mean model (opt-SNMM) under the assumption that a diagnostic test has no direct effect on the health outcome except through treatment choice (the NDE assumption). The main construction subtracts from a standard g-estimating function its projection onto the ortho-complement of the tangent space of the NDE model, thereby producing estimating equations with smaller asymptotic variance. The paper characterizes that ortho-complement (Theorem 1), proves a variance-reduction result for any user-specified projection subspace (Theorem 3), and provides a closed-form projection formula for a finite-dimensional subspace spanned by treatment-history indicators and basis functions of the outcome (Theorem 4 and Corollary 1). For discrete outcomes the subspace can be taken to be the full ortho-complement, while for continuous outcomes the authors propose a finite-basis subspace and assert in Remark 4 that the resulting estimator is asymptotically equivalent to the optimal estimator as the basis dimension grows. The paper also develops a cross-fitted doubly robust feasible estimator and studies the NDE-IPW estimator as a special case, with simulations illustrating large efficiency gains in an HIV monitoring setting.
Significance. If the central results stand, the paper makes a substantive contribution to causal inference for dynamic treatment and testing regimes. Theorem 3 gives a clean projection argument that, for fixed b, any estimator solving the projected estimating equation is RAL and dominates the standard opt-SNMM g-estimator whenever the projection is nonzero, while retaining double robustness. Theorem 1 and Theorem 4 are proved in sufficient detail and provide a workable route to constructing the projection, including a closed form for discrete outcomes. The explicit connection to NDE-IPW estimation and the simulation studies, which show large efficiency gains and improved regime selection, are valuable and are based on reproducible code in the supplementary materials. The main weakness is that the recommended near-optimal estimator for continuous outcomes relies on an unproved convergence assertion in Remark 4; as written, the paper proves improved efficiency over g-estimation for any finite basis but does not prove the advertised approach to the efficiency bound.
major comments (3)
- [Sec. 5.2, Remark 4] The claim that \hat\Psi(q,bsub) is asymptotically equivalent to \hat\Psi(q,bopt) for continuous outcomes is supported only by the sentence that the projection 'should converge' to \Pi[U|\Lambda^\perp_NDE]. No theorem establishes L2(P)-density of \cup_\xi \Omega_\xi in \Lambda^\perp_NDE, no rates for \xi(n) are given, and no conditions are stated under which the map b \mapsto T_b preserves L2 density. This statement is load-bearing: it is the basis for recommending \hat\Psi(q,bsub) as a near-optimal estimator. If the closure of \cup_\xi \Omega_\xi is a proper subspace of \Lambda^\perp_NDE, the asymptotic variance of \hat\Psi(q,bsub) has a strictly positive gap from V_oracle(q,bopt), and the paper's central practical message that relative efficiency can be made arbitrarily close to optimal would be unsupported. The fixed-b variance reduction of Theorem 3 would remain valid, but the near-optimality claim requires a proof or explicit sufficient conditions.
- [Sec. 4, Eq. (8) and Sec. 7] The entire projection construction uses that E[T_b]=0, which follows from Eq. (8), a consequence of the NDE(Yd) assumption. The manuscript acknowledges in Section 7 that NDE(Yd) can fail in realistic settings (for example, through ancillary care), but it does not quantify the bias of the adjusted estimators under such violations or provide a sensitivity analysis. This is not an internal inconsistency, but it is a substantive limitation of the practical recommendation: the efficiency gains are conditional on an untestable assumption, and the paper would be strengthened by an explicit statement of the resulting bias-variance trade-off or a small sensitivity analysis.
- [Sec. 6, Theorem 5] The feasible estimator is shown to be RAL under high-level conditions E[\hat U(q,\Psi^*)|Nu] = op(n^{-1/2}) and \sum_t E[\hat T_{b,t}|Nu] = op(n^{-1/2}). These conditions are stated as sufficient and are standard in the double/debiased machine learning literature, but for the recommended near-optimal estimator with estimated bsub they are combined with the unproved density claim of Remark 4. The paper should make explicit that the practical guarantee for \hat\Psi(q,\hat bsub) is therefore only the fixed-b variance reduction of Theorem 3 unless Remark 4 is upgraded to a theorem.
minor comments (5)
- [Sec. 2, paragraph after notation] The sentence 'We let \bar H_m be the sample space of the random vector \bar H_m' uses the same symbol for the random vector and its sample space; this is confusing and should be rephrased, for example by using a script or calligraphic letter for the sample space.
- [Sec. 5.2, Remark 4] Remark 4 refers to \phi(Y) while Corollary 1 defines b*_t in terms of \phi(Y_d). Since the total utility Y includes the known testing cost and the NDE assumption concerns Y_d, the paper should clarify which variable is used in the basis functions and why the cost-adjusted version is appropriate or not.
- [Introduction, Section 1] The 50-fold efficiency gain attributed to Caniglia et al. [3] concerns an NDE-IPW estimator in a dyn-MSM analysis, not the opt-SNMM estimators developed here; the text should make this distinction explicit to avoid overstating the simulation evidence for the proposed estimators.
- [Sec. 3.2, paragraph after Eq. (4)] The paper excludes exceptional laws but does not define them in the main text; a one-sentence definition or a more precise pointer to Robins [27] would make the exclusion self-contained.
- [Sec. 5.1, Theorem 3] Theorem 3 assumes \hat\Psi(q,b) is RAL rather than stating conditions under which the estimator solving 0 = \hat U(q,b,\Psi) is RAL; the regularity conditions in Section 6 and Appendix A.5 are stated for the cross-fitted version, and it would help the reader if the theorem explicitly noted that these conditions are being assumed.
Circularity Check
No significant circularity: efficiency gain is a direct projection variance identity, NDE is an input, and the Remark 4 gap is a correctness risk, not circularity.
full rationale
The paper's central claim is that projecting the influence function of a baseline RAL estimator onto the space T_b of mean-zero variables implied by the NDE assumption yields an estimator with variance Voracle(q,b) = J^-1(var[U] - c_OLS E[T_b T_b^T] c_OLS^T)J^-T. This is a mathematical identity following from the definition of c_OLS as the population least-squares coefficient; it is not a fitted quantity renamed as a prediction. The target Psi* and the baseline estimating function U(q,Psi) are defined from the opt-SNMM and ID assumptions independently of the projection, so the improvement is not self-definitional. The NDE assumption (eq. 7) is an explicit substantive input, and its observed-data consequence (eq. 8) is used to give every T_b mean zero; the paper proves the key characterization Lambda_perp_NDE = T_0+...+T_K as Theorem 1 in Appendix A.1 rather than importing it as an unverified postulate. Citations to Robins [25,27] supply the baseline opt-SNMM and the M1 nuisance tangent space; these are parameter-free published characterizations with stated assumptions that do not include the paper's efficiency result, so they do not constitute load-bearing self-citation. The one passage requiring explicit flagging is Remark 4: for continuous Y, the paper asserts that Pi[U|Omega] 'should converge' to Pi[U|Lambda_perp_NDE] as xi->infinity and that choosing xi slowly with n makes the estimators asymptotically equivalent. No proof or rate condition is given, so the near-optimality of the recommended finite-basis estimator is not established; however, this is an omitted convergence proof, not a circularity, because the target projection is defined independently of the approximating subspace. The simulation studies compare the proposed estimators with standard g-estimation and IPW benchmarks and do not redefine those benchmarks in terms of the proposed method. Overall, I find no step in which a claimed derivation reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (1)
- basis dimension ξ of the subspace Ω =
6 in simulations; grows to infinity in theory
assumptions (6)
- domain assumption ID assumptions: consistency, positivity, and sequential exchangeability (Section 2, assumptions 1-3).
- domain assumption NDE(Yd) assumption, equation (7): testing history has no direct effect on Yd except through treatment history.
- domain assumption Exclusion of exceptional laws (Robins 2004, page 219).
- domain assumption Correct specification of the opt-SNMM in equation (5).
- standard math Regularity conditions for RAL and cross-fitted DR-ML estimators (Appendix A.5).
- ad hoc to paper For the near-optimal continuous-outcome estimator, convergence of the basis projection as ξ and n grow (Remark 4).
invented entities (1)
-
Latent underlying test result R*_t
Cite this review
Pith. "Pith review of Efficient estimation of optimal regimes under a no direct effect assumption." pith.science (2026). https://pith.science/paper/BRSQVAX3
@misc{pith2026190810448,
author = {Pith},
title = {Pith review of: Efficient estimation of optimal regimes under a no direct effect assumption},
year = {2026},
howpublished = {\url{https://pith.science/paper/BRSQVAX3}},
note = {Machine review of arXiv:1908.10448}
}
read the original abstract
We derive new estimators of an optimal joint testing and treatment regime under the no direct effect (NDE) assumption that a given laboratory, diagnostic, or screening test has no effect on a patient's clinical outcomes except through the effect of the test results on the choice of treatment. We model the optimal joint strategy using an optimal regime structural nested mean model (opt-SNMM). The proposed estimators are more efficient than previous estimators of the parameters of an opt-SNMM because they efficiently leverage the `no direct effect (NDE) of testing' assumption. Our methods will be of importance to decision scientists who either perform cost-benefit analyses or are tasked with the estimation of the `value of information' supplied by an expensive diagnostic test (such as an MRI to screen for lung cancer).
Figures
Reference graph
Works this paper leans on
-
[1]
Bang, H. and J. M. Robins (2005). Doubly robust estimation in missing data and causal inference models. Biometrics 61 (4), 962–973
work page 2005
-
[2]
Bellman, R. (1952). On the theory of dynamic programming. Proceedings of the National Academy of Sciences of the United States of America 38 (8), 716
work page 1952
-
[3]
Caniglia, E. C., J. M. Robins, L. E. Cain, C. Sabin, R. Logan, S. Abgrall, M. J. Mugavero, S. Hern´ andez-D´ ıaz, L. Meyer, R. Seng, et al. (2019). Emulating a trial of joint dynamic strategies: An application to monitoring and treatment of HIV-positive individuals. Statistics in Medicine 38 (13), 2428–2446
work page 2019
-
[4]
Chernozhukov, V., D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins (2018). Double/debiased machine learning for treatment and structural parameters. The Econometrics Journal 21 (1), C1–C68
work page 2018
-
[5]
DART Trial Team (2010). Routine versus clinically driven laboratory monitoring of hiv antiretroviral therapy in Africa (DART): a randomised non-inferiority trial. The Lancet 375 (9709), 123–131
work page 2010
-
[6]
Ford, D., J. M. Robins, M. L. Petersen, D. M. Gibb, C. F. Gilks, P. Mugyenyi, H. Grosskurth, J. Hakim, E. Katabira, A. G. Babiker, et al. (2015). The impact of different cd4 cell-count monitoring and switching strategies on mortality in hiv-infected african adults on antiretroviral therapy: an application of dynamic marginal structural models. American Jou...
work page 2015
-
[7]
Gould, J. P. (1974). Risk, stochastic preference, and the value of information. Journal of Economic Theory 8 (1), 64–84
work page 1974
-
[8]
statistical issues arising in the Women’s Health Initiative
Hern´ an, M. A., J. M. Robins, and L. A. Garc´ ıa Rodr´ ıguez (2005). Discussion on 48 “statistical issues arising in the Women’s Health Initiative”. Biometrics 61 (4), 922– 930
work page 2005
Show all 37 references
-
[9]
Hilton, R. W. (1981). The determinants of information value: Synthesizing some general results. Management Science 27 (1), 57–64
1981
-
[10]
Mao, and M
Kallus, N., X. Mao, and M. Uehara (2019). Localized debiased machine learning: Efficient estimation of quantile treatment effects, conditional value at risk, and beyond. arXiv preprint arXiv:1912.12945
2019 arXiv
-
[11]
Krahn, M. D., J. E. Mahoney, M. H. Eckman, J. Trachtenberg, S. G. Pauker, and A. S. Detsky (1994). Screening for prostate cancer: a decision analytic view. Jama 272 (10), 773–780
1994
-
[12]
Sofrygin, J
Kreif, N., O. Sofrygin, J. A. Schmittdiel, A. S. Adams, R. W. Grant, Z. Zhu, M. J. van der Laan, and R. Neugebauer (2020). Exploiting nonsystematic covariate moni- toring to broaden the scope of evidence about the causal effects of adaptive treatment strategies. Biometrics
2020
-
[13]
Lara, A. M., J. Kigozi, J. Amurwon, L. Muchabaiwa, B. N. Wakaholi, R. E. M. Mota, A. S. Walker, R. Kasirye, F. Ssali, A. Reid, et al. (2012). Cost effectiveness analysis of clinically driven versus routine laboratory monitoring of antiretroviral therapy in Uganda and Zimbabwe. ...
2012
-
[14]
LaValle, I. H. (1968a). On cash equivalents and information evaluation in decisions under uncertainty Part I: Basic theory. Journal of the American Statistical Associa- tion 63 (321), 252–276
1968
-
[15]
LaValle, I. H. (1968b). On cash equivalents and information evaluation in decisions under uncertainty Part II: Incremental information decisions. Journal of the American Statistical Association 63 (321), 277–284. 49
1968
-
[16]
Luedtke, A. R., O. Sofrygin, M. J. van der Laan, and M. Carone (2017). Se- quential double robustness in right-censored longitudinal models. arXiv preprint arXiv:1705.02459
2017 arXiv
-
[17]
Murphy, S. A. (2003). Optimal dynamic treatment regimes. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 65 (2), 331–355
2003
-
[18]
Mushlin, A. I. and L. Fintor (1992). Is screening for breast cancer cost-effective? Cancer 69 (S7), 1957–1962
1992
-
[19]
Neugebauer, R., J. A. Schmittdiel, A. S. Adams, R. W. Grant, and M. J. van der Laan (2017). Identification of the joint effect of a dynamic treatment intervention and a stochastic monitoring intervention under the no direct effect assumption. Journal of Causal Inference 5 (1)
2017
-
[20]
Rotnitzky, and J
Orellana, L., A. Rotnitzky, and J. M. Robins (2010a). Dynamic regime marginal structural mean models for estimation of optimal dynamic treatment regimes, Part I: main content. The International Journal of Biostatistics 6 (2)
2010
-
[21]
Rotnitzky, and J
Orellana, L., A. Rotnitzky, and J. M. Robins (2010b). Dynamic regime marginal structural mean models for estimation of optimal dynamic treatment regimes, Part II: proofs of results. The International Journal of Biostatistics 6 (2)
2010
-
[22]
Robins, J. (1986). A new approach to causal inference in mortality studies with a sustained exposure period -— application to control of the healthy worker survivor effect. Mathematical Modelling 7 (9-12), 1393–1512
1986
-
[23]
Orellana, and A
Robins, J., L. Orellana, and A. Rotnitzky (2008). Estimation and extrapolation of optimal treatment and testing strategies. Statistics in Medicine 27 (23), 4678–4721
2008
-
[24]
A new approach to causal inference in mortality 50 studies with a sustained exposure period – application to control of the healthy worker survivor effect
Robins, J. M. (1987). Addendum to “A new approach to causal inference in mortality 50 studies with a sustained exposure period – application to control of the healthy worker survivor effect”. Computers & Mathematics with Applications 14 (9-12), 923–945
1987
-
[25]
Robins, J. M. (1999). Testing and estimation of direct effects by reparameterizing directed acyclic graphs with structural nested models. In Computation, Causation, and Discovery (C. Glymour and G. Cooper, eds.) , pp. 349–405. AAAI Press, Menlo Park, CA
1999
-
[26]
Robins, J. M. (2000). Marginal structural models versus structural nested models as tools for causal inference. In Statistical models in epidemiology, the environment, and clinical trials, pp. 95–133. Springer
2000
-
[27]
Robins, J. M. (2004). Optimal structural nested models for optimal sequential de- cisions. In Proceedings of the Second Seattle Symposium in Biostatistics , pp. 189–326. Springer
2004
-
[28]
Robins, J. M. and A. Rotnitzky (1995). Semiparametric efficiency in multivariate regression models with missing data. Journal of the American Statistical Associa- tion 90 (429), 122–129
1995
-
[29]
Rotnitzky, A., J. M. Robins, and L. Babino (2017). On the multiply robust estimation of the mean of the g-functional. arXiv preprint arXiv:1705.08582
2017 arXiv
-
[30]
Rotnitzky, and J
Smucler, E., A. Rotnitzky, and J. M. Robins (2019). A unifying approach for doubly- robust 𝓁1 regularized estimation of causal contrasts. arXiv preprint arXiv:1904.03737
2019 arXiv
-
[31]
van der Laan, M. J. and M. L. Petersen (2007). Causal effect models for realistic individualized treatment and intention to treat rules. The International Journal of Biostatistics 3 (1)
2007
-
[32]
van der Vaart, A. and J. Wellner (1996). Weak Convergence and Empirical Processes: with Applications to Statistics . Springer Science & Business Media. 51
1996
-
[33]
van der Vaart, A. W. (1998). Asymptotic statistics, Volume 3. Cambridge University Press
1998
-
[34]
Vansteelandt, S. and M. Joffe (2014). Structural nested models and G-estimation: The partially realized promise. Statistical Science 29 (4), 707–731
2014
-
[35]
Report of the commission on macroeconomics and health
World Health Organization (2001). Report of the commission on macroeconomics and health. 52 A Appendix The online supplementary materials include technical proofs and more details on the simulation studies omitted in the main text. A.1 Proof of Theorem 1 In this section, we pr...
2001
-
[36]
Notice thatyk+1,η† k+1 ( ¯Hk+1) coincides with E [ η† k+1( ¯Hk+1,Sk+1) p(Sk+1| ¯Hk+1) ⏐⏐⏐⏐ ¯Hk+1 ]
Given arbitrary functions η† k( ¯Hk,Sk), for k = 0,...,K , define yk+1,η† k+1 ( ¯Hk+1) ≡ ∑ sk+1 η† k+1( ¯Hk+1,sk+1). Notice thatyk+1,η† k+1 ( ¯Hk+1) coincides with E [ η† k+1( ¯Hk+1,Sk+1) p(Sk+1| ¯Hk+1) ⏐⏐⏐⏐ ¯Hk+1 ] . Given a fixed timet and a functionbt( ¯Ht, ¯St,Y d) defineηK( ...
-
[37]
always treat
So, we can write, ~b1,0(R∗ 1, ¯H1,C 2,R∗ 2, (S1,S 2),Y d) =~b1,0( ¯H1,C 2, (S1,S 2),Y d). Next, evaluating at A0 = 1,A 1 = 0 we conclude that − ~b1,1(R∗ 1, ¯H1,C 2,R∗ 2, (S1,S 2),Y d) p(S2| ¯H2) 1 Π0 − ~b1,0( ¯H1,C 2, (S1,S 2),Y d) p(S2| ¯H2) ( 1 Π0 − 1 ) cannot depend on R∗ 2...
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.