REVIEW 3 major objections 4 minor 98 references
Doubly Robust Inference on Causal Derivative Effects for Continuous Treatments
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper constructs kernel-smoothed, doubly robust estimators for the derivative of the causal dose-response curve and proves they are asymptotically normal at the nonparametric rate $\sqrt{nh^3}$, with and without the positivity…
desk verdict Solid DR inference for the derivative of the dose-response curve under positivity; the no-positivity half is honest about its additive-model assumption in the body but the abstract overstates its causal scope. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The workhorse is the doubly robust kernel estimator (8), which writes the derivative estimate as an average of a local-polynomial residual: each observation contributes $\frac{(T_i - t)/h \cdot K((T_i - t)/h)}{h \kappa_2 \widehat{p}_{T|S}(T_i|S_i)}\left(Y_i - \widehat{\mu}(t,S_i) - (T_i - t)\widehat{\beta}(t,S_i)\right)$ plus a regression adjustment $h\widehat{\beta}(t,S_i)$, with $\kappa_2 = \int u^2 K(u)\,du$. The local-polynomial correction pushes the inverse-probability residual to second order so that the Neyman orthogonality underlying double robustness holds as $h \to 0$. Without positivity, the central object is the modified IPW quantity (21), where the weight $1/p_{T|S}(T|S)$ is replaced by $p_\zeta(S|t)/p(T,S)$ using the $\zeta$-interior conditional density $p_\zeta$: the conditional density of $S$ given $T = t$ restricted to either the shrunken support $\mathcal{S}(t)\ominus\zeta$ or the level set $L_\zeta(t) = \{s : p_{S|T}(s|t) \geq \zeta\}$. Assumption A6 ensures the interior sits inside $\mathcal{S}(t+\delta)$ for small $\delta$, so the boundary discrepancy disappears; nuisance functions are estimated on independent folds via cross-fitting, and inference uses the sample variance of the influence function (9) or a multiplier bootstrap for uniform bands.
What would settle it
Generate data from design (28) with an added interaction, $Y = T^3 + T^2 + 10S + \gamma\,T\,S + \varepsilon$, so that model (13) fails while positivity still fails; if the bias-corrected DR estimator (26) develops a bias proportional to $\gamma$ that does not vanish as $n$ grows, the additive structural assumption is shown to be indispensable. A check of the positivity-half asymptotics: with the density model correct and the outcome model misspecified, the empirical bias of $\widehat{\theta}_{\mathrm{DR}}(t)$ should equal $h^2 B_\theta(t)$ with $B_\theta(t) = (\kappa_4/6\kappa_2)\cdot\mathbb{E}_S[\partial^3/\partial t^3\,\mu(t,S)]$ to leading order, so a detectable mismatch at $h \asymp n^{-1/5}$ would refute the claimed rate and bias expansion.
Extended reading notes
Core claim
On the paper's own terms the central discovery is an estimator together with a theorem about it. Under the identification conditions A1, smoothness conditions A3–A5, and positivity A2, the doubly robust estimator (8) satisfies, for fixed $t$ and bandwidth $h$ with $nh^7 \to c_3 \geq 0$, $\sqrt{nh^3}\big(\widehat{\theta}_{\mathrm{DR}}(t) - \theta(t) - h^2 B_\theta(t)\big) \to N(0, V_\theta(t))$, where $B_\theta(t)$ is an explicit bias term and $V_\theta(t)$ is the variance of the influence-function contribution $\phi_{h,t}$. The double robustness means the limit holds if either the outcome model $(\mu, \beta)$ is correct or the conditional density model $p_{T|S}$ is correct. When positivity fails, the paper proves that even oracle IPW estimators do not converge to $\theta(t)$: they converge instead to $\bar{m}'(t)\cdot \rho(t)$, where $\rho(t) = P(S \in \mathcal{S}(t))$ is the probability that the covariates fall in the conditional support of $S$ given $T = t$, because of the support discrepancy between $\mathcal{S}(t+uh)$ and $\mathcal{S}(t)$. The bias-corrected IPW and DR estimators (25)–(26) multiply by the density ratio $p_{S|T}(S|t)/p_S(S)$ and restrict to a $\zeta$-interior of the conditional support, which removes the boundary discrepancy and restores $\sqrt{nh^3}$-asymptotic normality with the same double robustness (Theorem 6), under the additive confounding model (13). The paper also shows the estimator attains the nonparametric efficiency bound of a smoothly approximated functional, a necessary detour because $\theta(t)$ itself lacks a pathwise derivative.
Load-bearing premise
For the no-positivity half, identification of the derivative effect rests on the additive confounding model $Y = \bar{m}(T) + \eta(S) + \varepsilon$ (equation 13), which forbids any interaction between treatment and covariates; if the true outcome contains a $T$-by-$S$ interaction, the bias-corrected IPW and DR estimators converge to a different quantity than the causal derivative $\theta(t)$.
Editorial extensions
If this is right
- Choosing the bandwidth $h$ of order $n^{-1/5}$ makes the leading bias $h^2 B_\theta(t)$ asymptotically negligible, so the Wald intervals $\widehat{\theta}_{\mathrm{DR}}(t) \pm q_{1-\tau/2}\sqrt{\widehat{V}_\theta(t)/(nh^3)}$ have valid coverage at the standard nonparametric rate.
- The rate conditions in Theorem 1 are satisfied by flexible machine-learned nuisance estimates, so the method can be used with neural-network fits of $\mu$, $\beta$, and $p_{T|S}$ under cross-fitting without changing the asymptotic theory.
- Integrating the bias-corrected derivative estimators yields IPW and DR estimators of the dose-response curve $m(t)$ that remain consistent when positivity fails, extending the regression-adjustment construction to the full doubly robust family.
- The multiplier bootstrap version gives uniform confidence bands over the whole treatment range, allowing simultaneous statements about where the derivative curve is significantly nonzero.
Reading between the lines
- An automatic choice of the trimming level $\zeta$ is left open; the level-set threshold $0.5\cdot\max \widehat{p}$ used in the paper suggests a tunable bias-variance trade-off, and cross-validating $\zeta$ against the estimated variance rather than fixing the multiplier is a natural extension.
- Proposition 3's bias formula implies a practical diagnostic: when positivity is doubtful, reporting $\rho(t) = P(S \in \mathcal{S}(t))$ alongside the estimate tells the user how far a naive IPW analysis is from the true derivative effect.
- The boundary-correction idea transfers to other non-regular targets with partial overlap, such as the continuous-instrument DR estimator of Zeng et al. (2025), which the paper notes does not address positivity violations, or average derivative effects under trimmed support.
- A misspecification test of the additive model (13), for instance checking whether the residual $\mathbb{E}[Y - \bar{m}(T) - \eta(S) \mid T, S]$ carries a $T\cdot S$ interaction, would delineate exactly where the no-positivity guarantee ends.
Formalized claims in Lean
-
Claim #1: The proposed DR estimator is asymptotically normal at the rate sqrt(n h^3) under positivity and either correct outcome or density model.
/-- @claim 1 The proposed DR estimator is asymptotically normal at the rate sqrt(n h^3) under positivity and either correct outcome or density model. -/ noncomputable def main_doubly_robust_asymptotic_normality : Prop :=
-
Claim #2: Without positivity, conventional IPW estimators converge to the support-discrepancy limit instead of theta(t).
/-- @claim 2 Without positivity, conventional IPW estimators converge to the support-discrepancy limit instead of theta(t). -/ noncomputable def conventional_ipw_inconsistent_without_positivity : Prop :=
-
Claim #3: Under additive confounding and without positivity, the bias-corrected DR estimator recovers sqrt(n h^3) asymptotic normality.
/-- @claim 3 Under additive confounding and without positivity, the bias-corrected DR estimator recovers sqrt(n h^3) asymptotic normality. -/ noncomputable def bias_corrected_dr_asymptotic_normality : Prop :=
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops kernel-smoothing estimators for the causal derivative effect θ(t)=d/dt E[Y(t)] with a continuous treatment. Under a positivity assumption, it proposes RA, IPW, and DR estimators of θ(t); Theorem 1 gives sqrt(nh^3)-asymptotic normality for the DR estimator, with double robustness in either the outcome model (μ and β) or the conditional density model. Without positivity, the paper shows that conventional IPW and DR estimators are inconsistent, imposes an additive confounding model of the form Y=mbar(T)+η(S)+ε, and proposes bias-corrected IPW and DR estimators based on an interior conditional density; Theorem 6 establishes analogous asymptotic normality under further rate conditions. Simulations and a Job Corps case study illustrate the methods.
Significance. Direct inference on the derivative of the dose-response curve is a useful and underexplored target, and the positivity-half of the paper is technically solid: the DR estimator is derived from the influence function of a smoothed functional, the proofs contain explicit remainder analyses, and the double-robustness claim is supported under the stated conditions. The no-positivity half is a valuable first step that connects causal inference to support and level-set estimation, but its causal scope is conditional on the additive confounding model (13), and the rate conditions for the interior-density estimator are not fully resolved for general covariate dimension. The paper also ships reproducible code and a detailed appendix, which strengthens its value.
major comments (3)
- [Abstract; Section 4.1.1; Section 5.1] The abstract and the contribution list advertise bias-corrected IPW and DR estimators 'without positivity' without qualification. The identification in Section 4.1.1 requires the additive confounding model (13), and Proposition 5's limiting target is mbar'(t) under that model. For a non-additive outcome such as μ(t,s)=mbar(t)+η(s)+τ t s with a positivity violation, the modified IPW quantity (21) converges to ∫∂_t μ(t,s) pbarζ(s|t) ds, which is not the causal derivative θ(t). Please rephrase the abstract and the contribution list to state explicitly that the no-positivity results require the additive structural model.
- [Theorem 6; Remark 8] Theorem 6(a) requires sqrt(nh^3) ||bpζ(S|t)-pbarζ(S|t)||_{L2}=o(1). With the recommended bandwidth h ≍ n^{-1/5}, this is an L2 convergence rate of o(n^{-1/5}) for the interior conditional density estimator. The typical nonparametric rate O(n^{-2/(5+d)}) quoted in Remark 8 satisfies this requirement only when d<5. The remark's assertion that pbarζ can be viewed as a 'smooth surrogate' and that bpζ can be constructed with a smaller bandwidth is not supported by a concrete estimator or proof. As stated, Theorem 6's applicability to general fixed d is not established. Please provide an explicit construction of pbarζ and bpζ satisfying the rate, or restrict the theorem (and the advertised claims) to dimensions where such rates are available.
- [Section 6.1] The high-dimensional simulation (d=20) may not lie in the theoretical regime of Theorem 1. With h ≍ n^{-1/5}, condition (c) of Theorem 1 requires sqrt(nh) = n^{2/5} times the product of nuisance errors to be o(1). Using the neural-network rates O(n^{-2/(4+d)}) cited in Remark 8, the product does not vanish for d=20. The d=20 experiment can still be reported as an empirical illustration, but the paper should either verify the rate conditions for the implemented estimators or explicitly state that this simulation is outside the scope of the theorem.
minor comments (4)
- [Section 4.1.1] There is a duplicated word in the sentence defining β(t,s): 'estimators of of β(t, s)' should read 'estimators of β(t, s)'.
- [Theorem 1] Condition (b)(i) in the main-text statement says 'with only h ||βbar(t,S)-β(t,S)||_{L2} → 0'; since the same condition already assumes βbar=β, this is vacuous. The appendix version, which instead requires h times the estimation error of bβ to βbar to vanish, is the intended condition and should be stated consistently in the main text.
- [Figure 1] The vertical axis labels in the right-hand panel appear as '(t)' without the Greek symbol; please ensure the derivative-effect curve is labeled consistently with θ(t).
- [Section 6.2] The no-positivity simulation uses S ∈ R (d=1), which is consistent with the rate discussion in Remark 8. It would be helpful to note this explicitly when summarizing the simulation evidence for Theorem 6.
Circularity Check
No significant circularity: the DR estimators are derived from influence-function/local-polynomial expansions and verified against oracle IPW/RA benchmarks; the no-positivity results rest on an explicitly stated additive structural model whose identification is re-proved in Proposition C.1.
full rationale
The paper's derivation chain is self-contained rather than circular. Under positivity, the DR estimator (8) is constructed from a local-polynomial-approximated IPW term plus an RA term, and Theorem 1's asymptotic linear form, variance Vθ(t), and bias Bθ(t) are obtained by explicit Taylor expansions of the oracle quantities (Section F.3), not by assuming the conclusion. The double robustness property follows from a standard decomposition in which the first-order bias contains a product of the density error and the outcome-model error (Term X and Term XI of the proof), which vanishes under either correct specification; this is a derived property, not a relabeled fit. The efficiency claim (Theorem 2) is explicitly relative to the smoothed functional ϖ_{h,t}(P), and the paper states that the target θ(t) lacks a pathwise derivative, so the efficiency statement is honestly scoped rather than tautological. In the no-positivity half, the identification of θ(t)=m̄'(t) is imported from the authors' own prior work (Zhang et al., 2024) via the additive confounding model (13), but the paper states (13) as an explicit structural assumption, re-proves the identification theory in Proposition C.1, and grounds the assumption in external spatial-statistics literature (Paciorek, 2010; Schnell and Papadogeorgou, 2020). The bias-corrected estimators (25)-(26) are shown consistent for m̄'(t) by direct expectation calculations (Propositions 4 and 5) whose integrals use μ(t,s)=m̄(t)+η(s), so the conclusion tracks the stated assumption rather than being forced by construction. The liberty taken is in the abstract, which advertises bias-corrected IPW/DR estimators 'when the positivity condition is violated' without mentioning that their causal interpretation requires (13); the body (Sections 4.1.1 and 5) and Remark 8 qualify this clearly. That is a scope/overclaim concern ('correctness risk'), not circularity, because the paper never defines θ(t) in terms of the estimator and never relabels a fitted parameter as a prediction. The remaining self-citations (Zhang et al. 2024 for the additive identification; Fan et al. 2022 for uniform bootstrap validity) are provenance for externally checkable assumptions and re-derived results, so they are not load-bearing reductions. Score 1 reflects the self-citation provenance for the no-positivity identification while confirming the central DR derivations are independent.
Assumptions & free parameters
free parameters (2)
- Kernel bandwidth h =
h = C_h sigma_T n^{-1/5} in simulations; theory allows h ~ n^{-1/5} for inference
- Interior trimming parameter zeta =
zeta = 0.5 max_i bp_{S|T}(S_i|t) in simulations
assumptions (8)
- domain assumption Assumption A1(a,b): consistency and unconfoundedness Y(t) independent of T given S
- domain assumption Assumption A1(c): positive conditional variance of T given S
- domain assumption Assumption A1(d): interchangeability of expectation and differentiation
- domain assumption Assumption A2: positivity of the conditional density p_{T|S}
- standard math Assumptions A3-A5: smoothness of μ, densities, and regular kernel properties
- domain assumption Additive confounding model Y = mbar(T) + eta(S) + epsilon
- domain assumption Assumption A6: Lipschitz-type support smoothness for S(t)
- domain assumption Nuisance rate conditions in Theorem 1(c) and Theorem 6(c)
Cite this review
Pith. "Pith review of Doubly Robust Inference on Causal Derivative Effects for Continuous Treatments." pith.science (2026). https://pith.science/paper/IX45YXAA
@misc{pith2026250106969,
author = {Pith},
title = {Pith review of: Doubly Robust Inference on Causal Derivative Effects for Continuous Treatments},
year = {2026},
howpublished = {\url{https://pith.science/paper/IX45YXAA}},
note = {Machine review of arXiv:2501.06969}
}
read the original abstract
Statistical methods for causal inference with continuous treatments mainly focus on estimating the mean potential outcome function, commonly known as the dose-response curve. However, it is often not the dose-response curve but its derivative function that signals the treatment effect. In this paper, we investigate nonparametric inference on the derivative of the dose-response curve with and without the positivity condition. Under the positivity and other regularity conditions, we propose a doubly robust (DR) inference method for estimating the derivative of the dose-response curve using kernel smoothing. When the positivity condition is violated, we demonstrate the inconsistency of conventional inverse probability weighting (IPW) and DR estimators, and introduce novel bias-corrected IPW and DR estimators. In all settings, our DR estimator achieves asymptotic normality at the standard nonparametric rate of convergence with nonparametric efficiency guarantees. Additionally, our approach reveals an interesting connection to nonparametric support and level set estimation problems. Finally, we demonstrate the applicability of our proposed estimators through simulations and a case study of evaluating a job training program.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Bang and J
H. Bang and J. M. Robins. Doubly robust estimation in missing data and causal inference models. Biometrics, 61 0 (4): 0 962--973, 2005
2005
-
[2]
A. G. Baydin, B. A. Pearlmutter, A. A. Radul, and J. M. Siskind. Automatic differentiation in machine learning: a survey. Journal of Machine Learning Research, 18 0 (153): 0 1--43, 2018
2018
-
[3]
Bickel, C
P. Bickel, C. Klaassen, Y. Ritov, and J. Wellner. Efficient and Adaptive Estimation for Semiparametric Models. Springer New York, 1998
1998
-
[4]
M. Blondel and V. Roulet. The elements of differentiable programming. arXiv preprint arXiv:2403.14606, 2024
arXiv 2024
-
[5]
S. Bong and K. Lee. Local causal effects with continuous exposures: A matching estimator for the average causal derivative effect. arXiv preprint arXiv:2311.18532, 2023
work page Pith review arXiv 2023
-
[6]
M. Bonvini and E. H. Kennedy. Fast convergence rates for dose-response estimation. arXiv preprint arXiv:2207.11825, 2022
arXiv 2022
-
[7]
Bonvini, A
M. Bonvini, A. McClean, Z. Branson, and E. H. Kennedy. Incremental causal effects: an introduction and review. In Handbook of matching and weighting adjustments for causal inference, pages 349--372. Chapman and Hall/CRC, 2023
2023
-
[8]
Z. Branson, E. H. Kennedy, S. Balakrishnan, and L. Wasserman. Causal effect estimation after propensity score trimming with continuous treatments. arXiv preprint arXiv:2309.00706, 2023
arXiv 2023
Show all 98 references
-
[9]
B. Cadre. Kernel estimation of density level sets. Journal of Multivariate Analysis, 97 0 (4): 0 999--1023, 2006
2006
-
[10]
Calonico, M
S. Calonico, M. D. Cattaneo, and M. H. Farrell. On the effect of bias estimation on coverage accuracy in nonparametric inference. Journal of the American Statistical Association, 113 0 (522): 0 767--779, 2018
2018
-
[11]
Carone, A
M. Carone, A. R. Luedtke, and M. J. van der Laan. Toward computerized efficient estimation in infinite-dimensional models. Journal of the American Statistical Association, 114 0 (527): 0 1174--1190, 2019
2019
-
[12]
M. D. Cattaneo, R. K. Crump, and M. Jansson. Robust data-driven inference for density-weighted average derivatives. Journal of the American Statistical Association, 105 0 (491): 0 1070--1083, 2010
2010
-
[13]
Chen and Z
X. Chen and Z. Liao. Sieve m inference on irregular parameters. Journal of Econometrics, 182 0 (1): 0 70--86, 2014
2014
-
[14]
X. Chen, Z. Liao, and Y. Sun. Sieve inference on possibly misspecified semi-nonparametric time series models. Journal of Econometrics, 178: 0 639--658, 2014
2014
-
[15]
Cheng and Y.-C
G. Cheng and Y.-C. Chen. Nonparametric inference via bootstrapping the debiased estimator . Electronic Journal of Statistics, 13 0 (1): 0 2194 -- 2256, 2019
2019
-
[16]
Chernozhukov, D
V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins. Double/debiased machine learning for treatment and structural parameters . The Econometrics Journal, 21 0 (1): 0 C1--C68, 01 2018
2018
-
[17]
Chernozhukov, W
V. Chernozhukov, W. K. Newey, and R. Singh. Automatic debiased machine learning of causal and structural effects. Econometrica, 90 0 (3): 0 967--1027, 2022
2022
-
[18]
Colangelo and Y.-Y
K. Colangelo and Y.-Y. Lee. Double debiased machine learning nonparametric inference with continuous treatments. arXiv preprint arXiv:2004.03036, 2020
2004 arXiv
-
[19]
S. R. Cole and M. A. Hern \'a n. Constructing inverse probability weights for marginal structural models. American Journal of Epidemiology, 168 0 (6): 0 656--664, 2008
2008
-
[20]
A. Cuevas. Set estimation: Another bridge between statistics and geometry. Bolet \' n de Estad \' stica e Investigaci \'o n Operativa , 25 0 (2): 0 71--85, 2009
2009
-
[21]
Cuevas and R
A. Cuevas and R. Fraiman. A plug-in approach to support estimation . The Annals of Statistics, 25 0 (6): 0 2300 -- 2312, 1997
1997
-
[22]
Devroye and G
L. Devroye and G. L. Wise. Detection of abnormal behavior via nonparametric estimation of the support. SIAM Journal on Applied Mathematics, 38 0 (3): 0 480--488, 1980
1980
-
[23]
D \' az and N
I. D \' az and N. S. Hejazi. Causal mediation analysis for stochastic interventions. Journal of the Royal Statistical Society Series B: Statistical Methodology, 82 0 (3): 0 661--683, 2020
2020
-
[24]
D \' az and M
I. D \' az and M. J. van der Laan. Targeted data adaptive estimation of the causal dose--response curve. Journal of Causal Inference, 1 0 (2): 0 171--192, 2013
2013
-
[25]
Einmahl and D
U. Einmahl and D. M. Mason. Uniform in bandwidth consistency of kernel-type function estimators . The Annals of Statistics, 33 0 (3): 0 1380 -- 1403, 2005
2005
-
[26]
Fan and I
J. Fan and I. Gijbels. Local polynomial modelling and its applications, volume 66. Chapman & Hall/CRC, 1996
1996
-
[27]
Fan, Y.-C
Q. Fan, Y.-C. Hsu, R. P. Lieli, and Y. Zhang. Estimation of conditional average treatment effects with high-dimensional data. Journal of Business & Economic Statistics, 40 0 (1): 0 313--327, 2022
2022
-
[28]
M. H. Farrell, T. Liang, and S. Misra. Deep neural networks for estimation and inference. Econometrica, 89 0 (1): 0 181--213, 2021
2021
-
[29]
C. Flores. Estimation of dose-response functions and optimal doses with a continuous treatment. Technical report, Department of Economics, University of Miami, 2007. URL https://core.ac.uk/download/pdf/7169663.pdf
2007
-
[30]
C. A. Flores and A. Flores-Lagunes. Identification and estimation of causal mechanisms and net effects of a treatment under unconfoundedness. IZA Discussion Papers 4237, Institute of Labor Economics (IZA), 2009
2009
-
[31]
C. A. Flores, A. Flores-Lagunes, A. Gonzalez, and T. C. Neumann. Estimating the effects of length of exposure to instruction in a training program: The case of job corps. Review of Economics and Statistics, 94 0 (1): 0 153--171, 2012
2012
-
[32]
A. F. Galvao and L. Wang. Uniformly semiparametric efficient estimation of treatment effects with a continuous treatment. Journal of the American Statistical Association, 110 0 (512): 0 1528--1542, 2015
2015
-
[33]
Gasser and H.-G
T. Gasser and H.-G. M \"u ller. Estimating regression functions and their derivatives by the kernel method. Scandinavian Journal of Statistics, pages 171--185, 1984
1984
-
[34]
R. D. Gill and J. M. Robins. Causal inference for complex longitudinal data: the continuous case. Annals of Statistics, 29 0 (6): 0 1785--1811, 2001
2001
-
[35]
Godambe and V
V. Godambe and V. Joshi. Admissibility and bayes estimation in sampling finite populations. i. The Annals of Mathematical Statistics, 36 0 (6): 0 1707--1722, 1965
1965
-
[36]
Z. Guo, W. Yuan, and C.-H. Zhang. Decorrelated local linear estimator: Inference for non-linear effects in high-dimensional additive models. arXiv preprint arXiv:1907.12732, 2019
1907 arXiv
-
[37]
H \"a rdle and T
W. H \"a rdle and T. M. Stoker. Investigating smooth multiple regression by the method of average derivatives. Journal of the American Statistical Association, 84 0 (408): 0 986--995, 1989
1989
-
[38]
J. D. Hart and P. Vieu. Data-driven bandwidth choice for density estimation based on dependent data. The Annals of Statistics, pages 873--890, 1990
1990
-
[39]
Hines, K
O. Hines, K. Diaz-Ordaz, and S. Vansteelandt. Optimally weighted average derivative effects. arXiv preprint arXiv:2308.05456, 2023
2023 arXiv
-
[40]
Hirano and G
K. Hirano and G. W. Imbens. The Propensity Score with Continuous Treatments, chapter 7, pages 73--84. John Wiley & Sons, Ltd, 2004
2004
-
[41]
D. A. Hirshberg and S. Wager. Debiased inference of average partial effects in single-index models: Comment on wooldridge and zhu. Journal of Business & Economic Statistics, 38 0 (1): 0 19--24, 2020
2020
-
[42]
M. Huber. Identifying causal mechanisms (primarily) based on inverse probability weighting. Journal of Applied Econometrics, 29 0 (6): 0 920--943, 2014
2014
-
[43]
Huber, Y.-C
M. Huber, Y.-C. Hsu, Y.-Y. Lee, and L. Lettry. Direct and indirect effects of continuous treatments based on generalized propensity score weighting. Journal of Applied Econometrics, 35 0 (7): 0 814--840, 2020
2020
-
[44]
Ichimura and W
H. Ichimura and W. K. Newey. The influence function of semiparametric estimators. Quantitative Economics, 13 0 (1): 0 29--61, 2022
2022
-
[45]
Imai and D
K. Imai and D. A. van Dyk. Causal inference with general treatment regimes: Generalizing the propensity score. Journal of the American Statistical Association, 99 0 (467): 0 854--866, 2004
2004
-
[46]
Kallus and A
N. Kallus and A. Zhou. Policy evaluation and optimization with continuous treatments. In International Conference on Artificial Intelligence and Statistics, pages 1243--1251. PMLR, 2018
2018
-
[47]
E. H. Kennedy. Semiparametric doubly robust targeted double machine learning: a review. Handbook of Statistical Methods for Precision Medicine, pages 207--236, 2024
2024
-
[48]
E. H. Kennedy, Z. Ma, M. D. McHugh, and D. S. Small. Nonparametric methods for doubly robust estimation of continuous treatment effects. Journal of the Royal Statistical Society Series B: Statistical Methodology, 79 0 (4): 0 1229--1245, 2017
2017
-
[49]
S. Klosin. Automatic double machine learning for continuous treatment effects. arXiv preprint arXiv:2104.10334, 2021
2021 arXiv
-
[50]
D. S. Lee. Training, wages, and sample selection: Estimating sharp bounds on treatment effects. The Review of Economic Studies, 76 0 (3): 0 1071--1102, 2009
2009
-
[51]
Y.-Y. Lee. Partial mean processes with generated regressors: Continuous treatment effects and nonseparable models. arXiv preprint arXiv:1811.00157, 2018
2018 arXiv
-
[52]
Lee and C.-A
Y.-Y. Lee and C.-A. Liu. Lee bounds with a continuous treatment in sample selection. arXiv preprint arXiv:2411.04312, 2024
2024
-
[53]
E. L. Lehmann. Elements of large-sample theory. Springer, 1999
1999
-
[54]
Li and J
Q. Li and J. Racine. Cross-validated local linear nonparametric regression. Statistica Sinica, 14: 0 485--512, 2004
2004
-
[55]
A. Luedtke. Simplifying debiased inference via automatic differentiation and probabilistic programming. arXiv preprint arXiv:2405.08675, 2024
2024 arXiv
-
[56]
Luedtke and I
A. Luedtke and I. Chung. One-step estimation of differentiable hilbert-valued parameters. The Annals of Statistics, 52 0 (4): 0 1534--1563, 2024
2024
-
[57]
Mack and H.-G
Y. Mack and H.-G. M \"u ller. Derivative estimation in nonparametric regression with random predictor variable. Sankhy \=a : The Indian Journal of Statistics, Series A , pages 59--72, 1989
1989
-
[58]
McClean, Y
A. McClean, Y. Li, S. Bae, M. A. McAdams-DeMarco, I. D \' az, and W. Wu. Fair comparisons of causal parameters with many treatments and positivity violations. arXiv preprint arXiv:2410.13522, 2024
2024
-
[59]
Meier, S
L. Meier, S. van de Geer, and P. B \"u hlmann. High-dimensional additive modeling . The Annals of Statistics, 37 0 (6B): 0 3779 -- 3821, 2009
2009
-
[60]
J. Meloche. Asymptotic behaviour of the mean integrated squared error of kernel density estimators for dependent observations. The Canadian Journal of Statistics/La Revue Canadienne de Statistique, pages 205--211, 1990
1990
-
[61]
Neugebauer and M
R. Neugebauer and M. van der Laan. Nonparametric causal effects based on marginal structural models. Journal of Statistical Planning and Inference, 137 0 (2): 0 419--434, 2007
2007
-
[62]
W. K. Newey and J. R. Robins. Cross-fitting and fast remainder rates for semiparametric estimation. arXiv preprint arXiv:1801.09138, 2018
2018 arXiv
-
[63]
W. K. Newey and T. M. Stoker. Efficiency of weighted average derivative estimators and index models. Econometrica, 61 0 (5): 0 1199--1223, 1993
1993
-
[64]
J. Neyman. Optimal asymptotic tests of composite hypotheses. Probability and Statsitics, pages 213--234, 1959
1959
-
[65]
J. Neyman. C( ) tests and their use. Sankhy \=a : The Indian Journal of Statistics, Series A , 41 0 (1/2): 0 1--21, 1979
1979
-
[66]
C. J. Paciorek. The importance of scale for spatial-confounding bias and precision of spatial regression estimators. Statistical Science, 25 0 (1): 0 107--125, 2010
2010
-
[67]
Paszke, S
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Automatic differentiation in pytorch. In NIPS 2017 Workshop on Autodiff, 2017
2017
-
[68]
Paszke, S
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. Pytorch: An imperative style, high-per...
2019
-
[69]
J. L. Powell, J. H. Stock, and T. M. Stoker. Semiparametric estimation of index coefficients. Econometrica, 57 0 (6): 0 1403--1430, 1989
1989
-
[70]
Ratkovic and D
M. Ratkovic and D. Tingley. Estimation and inference on nonlinear and heterogeneous effects. The Journal of Politics, 85 0 (2): 0 421--435, 2023
2023
-
[71]
J. Robins. A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect. Mathematical Modelling, 7 0 (9-12): 0 1393--1512, 1986
1986
-
[72]
J. M. Robins, M. A. Hernan, and B. Brumback. Marginal structural models and causal inference in epidemiology. Epidemiology, 11 0 (5): 0 550--560, 2000
2000
-
[73]
D. B. Rubin. Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66 0 (5): 0 688--701, 1974
1974
-
[74]
A. Schick. On Asymptotically Efficient Estimation in Semiparametric Models . The Annals of Statistics, 14 0 (3): 0 1139 -- 1151, 1986
1986
-
[75]
Schindl, S
K. Schindl, S. Shen, and E. H. Kennedy. Incremental effects for continuous exposures. arXiv preprint arXiv:2409.11967, 2024
2024
-
[76]
Schnell and G
P. Schnell and G. Papadogeorgou. Mitigating unobserved spatial confounding when estimating the effect of supermarket access on cardiovascular disease deaths. Annals of Applied Statistics, 14: 0 2069--2095, 12 2020
2020
-
[77]
P. Z. Schochet, J. Burghardt, and S. Glazerman. National job corps study: The impacts of job corps on participants' employment and related outcomes. Mathematica policy research reports, Mathematica Policy Research, 2001
2001
-
[78]
P. Z. Schochet, J. Burghardt, and S. McConnell. Does job corps work? impact findings from the national job corps study. American Economic Review, 98 0 (5): 0 1864--1886, 2008
2008
-
[79]
J. Shao. Mathematical Statistics. Springer Science & Business Media, 2003
2003
-
[80]
R. M. Stolzenberg. The measurement and decomposition of causal effects in nonlinear and nonadditive models. Sociological Methodology, 11: 0 459--488, 1980
1980
-
[81]
C. J. Stone. Additive regression and other nonparametric models. The Annals of Statistics, 13 0 (2): 0 689--705, 1985
1985
-
[82]
L. Su, T. Ura, and Y. Zhang. Non-separable models with high-dimensional data. Journal of Econometrics, 212 0 (2): 0 646--677, 2019
2019
-
[83]
Swaminathan and T
A. Swaminathan and T. Joachims. The self-normalized estimator for counterfactual learning. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 28, 2015
2015
-
[84]
Takatsu and T
K. Takatsu and T. Westling. Debiased inference for a covariate-adjusted regression function. Journal of the Royal Statistical Society Series B: Statistical Methodology, 87 0 (1): 0 33--55, 2025
2025
-
[85]
H. F. Trotter and J. W. Tukey. Conditional monte carlo for normal samples. In Symposium on Monte Carlo Methods, pages 64--79. John Wiley and Sons, 1956
1956
-
[86]
A. B. Tsybakov. On nonparametric estimation of density level sets. The Annals of Statistics, 25 0 (3): 0 948--969, 1997
1997
-
[87]
M. J. van der Laan and J. M. Robins. Unified methods for censored longitudinal data and causality. Springer, 2003
2003
-
[88]
M. J. van der Laan, A. Bibaut, and A. R. Luedtke. Cv-tmle for nonpathwise differentiable target parameters. In M. J. van der Laan and S. Rose, editors, Targeted Learning in Data Science: Causal Inference for Complex Longitudinal Studies, pages 455--481. Springer, 2018
2018
-
[89]
van der Vaart
A. van der Vaart. On differentiable functionals. The Annals of Statistics, 19 0 (1): 0 178--204, 1991
1991
-
[90]
A. W. van der Vaart. Asymptotic Statistics. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 1998
1998
-
[91]
M. P. Wand and M. C. Jones. Kernel Smoothing. CRC press, 1994
1994
-
[92]
Wasserman
L. Wasserman. All of nonparametric statistics. Springer Science & Business Media, 2006
2006
-
[93]
Westreich and S
D. Westreich and S. R. Cole. Invited commentary: positivity in practice. American Journal of Epidemiology, 171 0 (6): 0 674--677, 2010
2010
-
[94]
X. Wu, F. Mealli, M.-A. Kioumourtzoglou, F. Dominici, and D. Braun. Matching on generalized propensity scores with continuous exposures. Journal of the American Statistical Association, 119 0 (545): 0 757--772, 2024
2024
-
[95]
Y. Xu, N. Sani, A. Ghassami, and I. Shpitser. Multiply robust causal mediation analysis with continuous treatments. arXiv preprint arXiv:2105.09254, 2021
2021 arXiv
-
[96]
Z. Zeng, A. W. Levis, J. Lee, E. H. Kennedy, and L. Keele. Nonparametric estimation of local treatment effects with continuous instruments. arXiv preprint arXiv:2504.03063, 2025
2025 arXiv
-
[97]
Zhang, Y.-C
Y. Zhang, Y.-C. Chen, and A. Giessing. Nonparametric inference on dose-response curves without the positivity condition. arXiv preprint arXiv:2405.09003, 2024
2024 arXiv
-
[98]
Zhou and D
S. Zhou and D. A. Wolfe. On derivative estimation in spline regression. Statistica Sinica, 10 0 (1): 0 93--108, 2000
2000
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.