REVIEW 3 major objections 5 minor 30 references
A class of nonparametric methods for evaluating the effect of continuous treatments on survival outcomes
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper develops a class of nonparametric tests for whether the counterfactual survival probability through a fixed time is constant across all levels of a continuous exposure, under right censoring and with machine-learning nuisance…
desk verdict Useful extension of the linear-contrast testing framework to right-censored survival, but the efficient influence function in the main text and the supplementary proof are inconsistent, so the theoretical core needs repair before this is citable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the linear contrast functional $\psi_{P,t}(h)$ and its efficient influence function $D_P(O; h)$, which combines the conditional survival curve, the conditional censoring function, and the conditional exposure density. The one-step bias correction adds the empirical mean of the estimated influence function to the plug-in estimator, removing first-order nuisance bias. Uniform convergence over $h\in H$ is secured by condition (B1), which requires the influence functions to fall in a P-Donsker class (a function class whose empirical process is stochastically bounded) with probability tending to one, together with rate conditions (B2) and (B3) on the nuisance estimators, permitting flexible machine-learning nuisance estimation at slower than $n^{1/2}$ rates. The test statistic is the supremum of the one-step estimator over $H$; under the null it converges to the supremum of a mean-zero Gaussian process whose covariance is the inner product of the influence functions, approximated by Monte Carlo.
What would settle it
Run the proposed test in a null simulation where the conditional survival and censoring functions are rough and are estimated by survival random forests with limited data; if the empirical distribution of $\sup_{h\in H} n^{1/2}\psi^\dagger_{n,t}(h)$ across many datasets does not match the Monte Carlo Gaussian supremum distribution, or if type-1 error stays above the nominal level as the sample size grows, then conditions (B1)-(B3) are failing and the test's validity claim is refuted.
Extended reading notes
Core claim
The central claim is that the causal null $\theta_P^a(t) = E_P[\theta_P^A(t)]$ for all $a$ can be tested nonparametrically by targeting the function-valued contrast $\psi_{P,t}(h) = E_P[(\theta_P^A(t) - E_P[\theta_P^A(t)])h(A)]$ over a suitably constrained class $H$ of functions $h$. Theorem 2 gives the efficient influence function $D_P(O; h)$ and shows $\psi_{P,t}(h)$ is pathwise differentiable even though $\theta_P^a(t)$ is not. The one-step estimator adds the empirical mean of the estimated influence function to a plug-in estimator, and Theorem 3 states that, under Donsker and rate conditions, the estimator is uniformly asymptotically linear: $\psi^\dagger_{n,t}(h) - \psi_{P,t}(h) = n^{-1}\sum_i D_P(O_i; h) + r_n(h)$ with $\sup_{h\in H}|r_n(h)| = o_p(n^{-1/2})$; consequently the process $n^{1/2}(\psi^\dagger_{n,t} - \psi_{P,t})$ over $H$ converges weakly to a mean-zero Gaussian process. The null is tested with $\Psi^\dagger_{n,t}(H) = \sup_{h\in H}|\psi^\dagger_{n,t}(h)|$ compared to Monte Carlo draws from the limiting Gaussian supremum. The paper argues this yields asymptotically valid type-1 error control, and power is determined by whether $\Psi_{P,t}(H)$ is a norm of the centered curve, for example an $\ell^1$ or variance norm under monotonicity or bounded-variation constraints.
Load-bearing premise
The whole inference rests on an unverified regularity premise: the flexible machine-learning estimators used for survival, censoring, and exposure density must be well-behaved enough (belonging to a suitably regular function class with adequate convergence rates) that the one-step estimator's remainder vanishes uniformly; if those estimators are too erratic, the Gaussian null distribution and Monte Carlo p-values are not justified.
Editorial extensions
If this is right
- Under the relaxed null $\sup_{h\in H}|\psi_{P,t}(h)|=0$, the proposed supremum test controls type-1 error asymptotically at the nominal level, for any contrast class $H$ satisfying the Donsker and rate conditions.
- For monotone dose-response alternatives, choosing sign-function contrasts makes the power target the probability-weighted $\ell^1$ norm of the centered counterfactual survival curve; under a variance constraint, the target is its $\ell^2$ norm.
- The method accommodates right censoring and uses flexible machine-learning nuisance estimators for survival, censoring, and exposure density, so it can be applied to observational studies and randomized trials with continuous biomarkers.
- Applied to the AMP HIV monoclonal antibody trials, the test yields non-significant p-values across tuning parameters, consistent with a nearly flat estimated counterfactual survival curve across Day-61 VRC01 concentration.
- As a corollary of Theorem 3, uniform asymptotic linearity makes possible simultaneous inference on the whole contrast process, not only the supremum statistic.
Reading between the lines
- The same linear-contrast construction could be extended to test equality of entire survival curves across exposure levels by taking the supremum over time $t$ as well as $h$, yielding a functional test of stochastic dominance; the paper only treats one fixed time point.
- The Donsker condition is the practical bottleneck: an empirical robustness check would compare p-values from provably regular contrast classes (for example, bounded-variation sieves) against those from survival random forest nuisances in a null simulation, to see whether inference is sensitive to nuisance irregularity.
- Because the causal assumptions (A2) are unverifiable in observational data, the test is also interpretable as a nonparametric test of the conditional exposure-survival association, which remains meaningful even if causal identification fails.
- One could attempt to choose the contrast class $H$ adaptively, for example by maximizing estimated power subject to a Donsker constraint; the paper leaves tuning selection as future work.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a nonparametric hypothesis-testing framework for whether the counterfactual survival probability at a fixed time point varies with a continuous exposure when the outcome is right-censored. The approach reformulates the null in terms of linear contrasts ψ_P,t(h) of the centered dose–response curve, establishes pathwise differentiability of these contrasts, constructs a one-step estimator with a uniform asymptotic linear representation over a function class H, and derives a Gaussian-process null distribution used for Monte Carlo p-values. The methodology is illustrated with simulations and an application to the AMP HIV prevention trials. The paper is an extension of the authors' earlier work on continuous exposures with uncensored outcomes to the censored-data setting.
Significance. If the central theorem is correct, the paper offers a principled way to test for exposure effects on survival across a continuum without parametric assumptions, which is valuable for infectious-disease biomarker studies. The paper provides a clear target estimand, an explicit influence-function construction, a careful discussion of function classes H, and simulation evidence. The proof strategy follows the one-step estimation route, and the application to AMP data is relevant. However, the manuscript currently has a serious inconsistency in the influence-function derivation and relies on unverified regularity conditions for machine-learning nuisance estimators; these must be resolved before the theoretical claims can be accepted.
major comments (3)
- [Theorem 2 and Supplementary Materials, Proof of Theorem 2] The efficient influence function D_P stated in Theorem 2 differs from the expression derived in the supplementary proof by a constant term: Theorem 2 contains −2E_P{h(A)θ̄_A^P(t)}, whereas the supplementary derivation ends with −E_P[h(A)]E_P[θ_A^P(t)]. These two quantities are not generally equal. Moreover, the supplementary expression is not mean-zero; its expectation is 2Cov_P(h(A), θ_A^P(t)) − E_P[h(A)]E_P[θ_A^P(t)] in general, which is impossible for an influence function. Since the one-step estimator in Eq. (5), the covariance estimator in §5.3, and the proof of Theorem 3 are all stated in terms of D_P, the proof of Theorem 3 is internally inconsistent until this discrepancy is resolved.
- [Conditions (B1)-(B3) in §4.1 and nuisance estimators in §5.1] Theorem 3 rests on conditions (B1)-(B3), which require the estimated efficient influence functions to fall in a P-Donsker class and the nuisance estimators to satisfy L2 convergence and a product-rate condition. Section 5.1 recommends survival random forests, survival super learner, and kernel density estimation via the highly adaptive lasso. For these data-adaptive estimators, the paper does not verify or cite results establishing (B1)-(B3); indeed, for such methods the Donsker condition is known to be fragile and may fail. As a consequence, the theoretical guarantee of uniform asymptotic linearity and the resulting Gaussian-process null distribution is not established for the implementations used in the simulations and application. The authors should either prove or cite conditions under which these specific nuisance estimators satisfy (B1)-(B3), or present the theoretical results as conditional on unverified assumptions.
- [Abstract, Introduction, and Eq. (3) in §3] The abstract and introduction claim the method tests whether the counterfactual survival probability is constant across the full continuous exposure range (H0 in Eq. (1)). However, the test is formally for the relaxed null H̄0 in Eq. (3): sup_{h∈H} |ψ_P,t(h)| = 0 for a chosen class H. As the paper notes, H̄0 may hold while H0 is false. The contribution as stated should be carefully qualified: the proposed test controls type-1 error for the relaxed null and has power only against alternatives detectable by H. The abstract and introduction should be amended to avoid overclaiming the scope of the null hypothesis being tested.
minor comments (5)
- [§6, Figure 1] The label "knots Max" in Figure 1 is unclear; it should be explained that "Max" corresponds to κ = n for the class in Eq. (9).
- [§4.1, Eq. (5)] D_n is called a plug-in estimator for the efficient influence function, but the displayed expression only contains the censoring-adjustment term. The relationship between D_n and the full D_P of Theorem 2 should be clarified, in particular how the remaining terms of D_P are handled by the plug-in component ψ_n,t(h).
- [Supplementary Materials, proof of Theorem 3, rII_n] In the rate calculation for rII_n, the term (n^{-1}Σ_i h(A_i) − E[h(A)])^2 is OP(n^{-1}), not OP(n^{-1/2}); the stated conclusion that the product is oP(n^{-1/2}) still follows, but the displayed rate should be corrected.
- [Data availability] The paper states that the supporting code is available on GitHub, but no repository URL is provided; the authors should include the link.
- [§4.2, Eq. (8)] The claim that for a function θ̄_a^P(t) that changes sign K times, the sign function lies in a class with variation norm bounded by 2K should be justified or accompanied by a reference, as it is used to motivate the choice of H.
Circularity Check
No significant circularity: the EIF, one-step estimator, and Gaussian-process null distribution are derived from the model rather than assumed, and the self-citations are non-load-bearing.
full rationale
The derivation chain is not circular. The target contrast ψ_P,t(h) is defined directly as E[θ̄_A(t) h(A)], and the reformulation of the null as ψ_P,t(h)=0 for all bounded h is a mathematical equivalence, not an assumption of the conclusion. Theorem 2 derives the efficient influence function by pathwise-differentiation calculations, and Theorem 3 establishes uniform asymptotic linearity under explicit conditions (B1)-(B3) on nuisance estimators and the function class H; the one-step estimator is the plug-in estimator plus an empirical mean of the estimated EIF, with the remainder terms bounded in the supplementary proof. The paper does not fit any parameter to the target quantity and then relabel it as a prediction; the relaxed null and the approximating classes H are explicitly presented as approximations, with the loss of power against alternatives outside H acknowledged. Reuse of the authors' earlier linear-contrast framework and of Westling et al.'s censored-survival results is real external support rather than a circular self-citation chain: those prior works are published, and the present paper's extension to continuous exposures with right-censored outcomes and its EIF theorem are derived independently. The discrepancy noted by the skeptic between the EIF displayed in Theorem 2 and the EIF expression in the supplementary proof is an internal consistency and correctness issue, not a circular reduction of the paper's claims to its inputs; it does not raise the circularity score.
Assumptions & free parameters
free parameters (2)
- number of knots or basis functions κ =
10, 20, 50, 100 in application and simulations
- variation norm bound λ =
4, 6, or infinity
assumptions (4)
- domain assumption Causal identification assumptions A1-A4: consistency, conditional ignorability for T(a) and C(a), positivity of exposure density and uncensored mass, and no interference.
- ad hoc to paper Condition (B1): existence of a P-Donsker class containing D_P(·;h) and D_n(·;h) for each h∈H.
- ad hoc to paper Conditions (B2) and (B3): L2 convergence of nuisance estimators and a product-rate condition on survival, censoring, and density estimators.
- standard math Standard product-integral representation of conditional survival functions and standard empirical process theory for weak convergence.
Cite this review
Pith. "Pith review of A class of nonparametric methods for evaluating the effect of continuous treatments on survival outcomes." pith.science (2026). https://pith.science/paper/R6RRIU56
@misc{pith2026241209786,
author = {Pith},
title = {Pith review of: A class of nonparametric methods for evaluating the effect of continuous treatments on survival outcomes},
year = {2026},
howpublished = {\url{https://pith.science/paper/R6RRIU56}},
note = {Machine review of arXiv:2412.09786}
}
read the original abstract
In randomized trials and observational studies, it is often necessary to evaluate the extent to which an intervention affects a time-to-event outcome, which is only partially observed due to right censoring. For instance, in infectious disease studies, it is frequently of interest to characterize the relationship between risk of acquisition of infection with a pathogen and a biomarker previously measuring for an immune response against that pathogen induced by prior infection and/or vaccination. It is common to conduct inference within a causal framework, wherein we desire to make inferences about the counterfactual probability of survival through a given time point, at any given exposure level. To determine whether a causal effect is present, one can assess if this quantity differs by exposure level. Recent work shows that, under typical causal assumptions, summaries of the counterfactual survival distribution are identifiable. Moreover, when the treatment is multi-level, these summaries are also pathwise differentiable in a nonparametric probability model, making it possible to construct estimators thereof that are unbiased and approximately normal. In cases where the treatment is continuous, the target estimand is no longer pathwise differentiable, rendering it difficult to construct well-behaved estimators without strong parametric assumptions. In this work, we extend beyond the traditional setting with multilevel interventions to develop approaches to nonparametric inference with a continuous exposure. We introduce methods for testing whether the counterfactual probability of survival time by a given time-point remains constant across the range of the continuous exposure levels. The performance of our proposed methods is evaluated via numerical studies, and we apply our method to data from a recent pair of efficacy trials of an HIV monoclonal antibody.
Figures
Reference graph
Works this paper leans on
-
[1]
Peter C Austin. The use of propensity score methods with survival or time-to-event outcomes: reporting measures of effect similar to those used in randomized experiments. Statistics in medicine, 33 0 (7): 0 1242--1258, 2014
work page 2014
-
[2]
The highly adaptive lasso estimator
David Benkeser and Mark van der Laan. The highly adaptive lasso estimator. In 2016 IEEE international conference on data science and advanced analytics (DSAA), pages 689--696. IEEE, 2016
work page 2016
-
[3]
Efficient and adaptive estimation for semiparametric models
Peter J Bickel, Chris AJ Klaassen, Ya’acov Ritov, and Jon A Wellner. Efficient and adaptive estimation for semiparametric models. Springer, 1998
work page 1998
-
[4]
One-step targeted maximum likelihood estimation for time-to-event outcomes
Weixin Cai and Mark J van der Laan. One-step targeted maximum likelihood estimation for time-to-event outcomes. Biometrics, 76 0 (3): 0 722--733, 2020
work page 2020
-
[5]
Jacob Cohen. The cost of dichotomization. Applied psychological measurement, 7 0 (3): 0 249--253, 1983
work page 1983
-
[6]
Adjusted survival curves with inverse probability weights
Stephen R Cole and Miguel A Hern \'a n. Adjusted survival curves with inverse probability weights. Computer methods and programs in biomedicine, 75 0 (1): 0 45--49, 2004
work page 2004
-
[7]
Two randomized trials of neutralizing antibodies to prevent hiv-1 acquisition
Lawrence Corey, Peter B Gilbert, Michal Juraska, David C Montefiori, Lynn Morris, Shelly T Karuna, Srilatha Edupuganti, Nyaradzo M Mgodi, Allan C Decamp, Erika Rudnicki, et al. Two randomized trials of neutralizing antibodies to prevent hiv-1 acquisition. New England Journal of Medicine, 384 0 (11): 0 1003--1014, 2021
work page 2021
-
[8]
Regression models and life-tables
David R Cox. Regression models and life-tables. Journal of the Royal Statistical Society: Series B (Methodological), 34 0 (2): 0 187--202, 1972
1972
Show all 30 references
-
[9]
CVXR: An R package for disciplined convex optimization
Anqi Fu, Balasubramanian Narasimhan, and Stephen Boyd. CVXR: An R package for disciplined convex optimization . Journal of Statistical Software, 94 0 (14): 0 1–34, 2020. doi:10.18637/jss.v094.i14. URL https://www.jstatsoft.org/index.php/jss/article/view/v094i14
2020 doi
-
[10]
Gill and James M
Richard D. Gill and James M. Robins. Causal inference for complex longitudinal data: The continuous case. The Annals of Statistics, 29 0 (6): 0 1785--1811, 2001
2001
-
[11]
Generalized additive models
Trevor J Hastie. Generalized additive models. In Statistical models in S, pages 249--307. Routledge, 2017
2017
-
[12]
Demystifying statistical learning based on efficient influence functions
Oliver Hines, Oliver Dukes, Karla Diaz-Ordaz, and Stijn Vansteelandt. Demystifying statistical learning based on efficient influence functions. The American Statistician, 76 0 (3): 0 292--304, 2022
2022
-
[13]
Nonparametric locally efficient estimation of the treatment specific survival distribution with right censored data and covariates in observational studies
Alan E Hubbard, Mark J Van Der Laan, and James M Robins. Nonparametric locally efficient estimation of the treatment specific survival distribution with right censored data and covariates in observational studies. In Statistical Models in Epidemiology, the Environment, and Cli...
2000
-
[14]
Inference on function-valued parameters using a restricted score test
Aaron Hudson, Marco Carone, and Ali Shojaie. Inference on function-valued parameters using a restricted score test. arXiv preprint arXiv:2105.06646, 2021
2021 arXiv
-
[15]
An approach to nonparametric inference on the causal dose response function
Aaron Hudson, Elvin H Geng, Thomas A Odeny, Elizabeth A Bukusi, Maya L Petersen, and Mark J van der Laan. An approach to nonparametric inference on the causal dose response function. arXiv preprint arXiv:2306.07736, 2023
2023 arXiv
-
[16]
Kogalur, Eugene H
Hemant Ishwaran, Udaya B. Kogalur, Eugene H. Blackstone, and Michael S. Lauer. Random survival forests . The Annals of Applied Statistics, 2 0 (3): 0 841 -- 860, 2008. doi:10.1214/08-AOAS169. URL https://doi.org/10.1214/08-AOAS169
2008 doi
-
[17]
Covariate-adjusted non-parametric survival curve estimation
Honghua Jiang, James Symanowski, Yongming Qu, Xiao Ni, and Yanping Wang. Covariate-adjusted non-parametric survival curve estimation. Statistics in Medicine, 30 0 (11): 0 1243--1253, 2011
2011
-
[18]
Adjusted survival curve estimation using covariates
Robert W Makuch. Adjusted survival curve estimation using covariates. Journal of Chronic Diseases, 35 0 (6): 0 437--443, 1982
1982
-
[19]
Bivariate median splits and spurious statistical significance
Scott E Maxwell and Harold D Delaney. Bivariate median splits and spurious statistical significance. Psychological Bulletin, 113 0 (1): 0 181, 1993
1993
-
[20]
On the estimation of average treatment effects with right-censored time to event outcome and competing risks
Brice Maxime Hugues Ozenne, Thomas Harder Scheike, Laila St rk, and Thomas Alexander Gerds. On the estimation of average treatment effects with right-censored time to event outcome and competing risks. Biometrical Journal, 62 0 (3): 0 751--763, 2020
2020
-
[21]
Contributions to a general asymptotic statistical theory
Johann Pfanzagl. Contributions to a general asymptotic statistical theory. Springer, 1982
1982
-
[22]
Model-based direct adjustment
Paul R Rosenbaum. Model-based direct adjustment. Journal of the American Statistical Association, 82 0 (398): 0 387--394, 1987
1987
-
[23]
The central role of the propensity score in observational studies for causal effects
Paul R Rosenbaum and Donald B Rubin. The central role of the propensity score in observational studies for causal effects. Biometrika, 70 0 (1): 0 41--55, 1983
1983
-
[24]
Dichotomizing continuous predictors in multiple regression: a bad idea
Patrick Royston, Douglas G Altman, and Willi Sauerbrei. Dichotomizing continuous predictors in multiple regression: a bad idea. Statistics in Medicine, 25 0 (1): 0 127--141, 2006
2006
-
[25]
Pharmacokinetic serum concentrations of VRC01 correlate with prevention of HIV -1 acquisition
Kelly E Seaton, Yunda Huang, Shelly Karuna, Jack R Heptinstall, Caroline Brackett, Kelvin Chiong, Lily Zhang, Nicole L Yates, Mark Sampson, Erika Rudnicki, et al. Pharmacokinetic serum concentrations of VRC01 correlate with prevention of HIV -1 acquisition. EBioMedicine, 93: 0...
2023
-
[26]
Asymptotic statistics, volume 3
Aad W van der Vaart. Asymptotic statistics, volume 3. Cambridge university press, 2000
2000
-
[27]
Weak convergence
Aad W van Der Vaart, Jon A Wellner, Aad W van der Vaart, and Jon A Wellner. Weak convergence. Springer, 1996
1996
-
[28]
Nonparametric tests of the causal null with nondiscrete exposures
Ted Westling. Nonparametric tests of the causal null with nondiscrete exposures. Journal of the American Statistical Association, 117 0 (539): 0 1551--1562, 2022
2022
-
[29]
Inference for treatment-specific survival curves using machine learning
Ted Westling, Alex Luedtke, Peter B Gilbert, and Marco Carone. Inference for treatment-specific survival curves using machine learning. Journal of the American Statistical Association, 119 0 (546): 0 1541--1553, 2024. doi:10.1080/01621459.2023.2205060
2024
-
[30]
Contrasting treatment-specific survival using double-robust estimators
Min Zhang and Douglas E Schaubel. Contrasting treatment-specific survival using double-robust estimators. Statistics in Medicine, 31 0 (30): 0 4255--4268, 2012
2012
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.