{"id":"c2269a7d-ebc9-48a1-bbeb-77abb3be7f48","arxiv_id":"2502.00120","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"New efficient estimators for the causal effect of treatment on years of life lost due to a specific competing event, plus a variable importance test for effect heterogeneity.","lead":"This paper introduces a way to measure how much a treatment changes the number of healthy life-years lost before a fixed time due to a specific competing event, using causal inference and machine learning. It also provides a statistical test for which patient characteristics drive treatment effect differences, demonstrated on antidepressant registry data.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Asymptotic normality for the RF-based estimators rests on Assumption B3 and Theorem 2(ii), rate conditions that are neither proved nor empirically verified for random survival forests; the paper's simulation does not close this gap.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing condition: Assumption B3 for the ATE estimator and the n^{-1/4} rate for the meta-learned τ_l in Theorem 2. I agree with that assessment. The paper's theorems are conditional and logically coherent, so there is no internal contradiction. The gap is that the central advertised contribution—valid inference under flexible ML nuisance estimation—depends on rate conditions that the paper does not verify for the recommended random survival forests, and the simulation study cannot substitute for a rate proof or a direct check of the remainder. The simulation's good coverage in one DGP is suggestive but not decisive, especially because the Discussion overstates it as 'confirmed' asymptotic behavior. A conditional verdict is therefore appropriate: the paper should be accepted only if the missing verification is supplied or the claims are scaled back to the conditional theorems. My read does not change the reader's verdict, so I mark it UNCHANGED.","tokens_in":26397,"tokens_out":6596,"duration_ms":71708,"concrete_test":"Compute the B3 remainder directly in the paper's simulation DGP: for each Monte Carlo replicate and each n in {250, 500, 1000, 2000}, evaluate R_n = |P{φ_1(hat ν) - τ_1}| using the known true Λ1, Λ2, Λc, π and the RSF-based hat ν with the same default settings, then regress log R_n on log n. If the slope is significantly greater than -0.5, or if R_n fails to decay at n^{-1/2}, Assumption B3 is violated for RSF even in this favorable setting, and the asserted asymptotic linearity of the RFCF estimator is not supported. A complementary check is to extend Table 1's RFCF coverage to n = 2000 and 4000; if coverage does not approach 0.95, the finite-sample agreement is not evidence of the claimed asymptotic theory.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The main theorems are conditional on high-level nuisance rate conditions. The most load-bearing is Assumption B3 (Section 4.1, Lemma C.1): the remainder term must be o_p(n^{-1/2}). As the authors note, for cumulative hazard estimators that are not absolutely continuous (e.g., Breslow-type or random survival forest estimators), the usual n^{-1/4}-type product-rate argument does not apply, and the text admits it is unclear whether RSF satisfy B3. Theorem 2 additionally assumes (ii) an n^{-1/4} rate for the meta-learned regression of CATE estimates onto X_{-l}, a rate the authors state is generally unknown. Nevertheless, Section 7 claims the simulations 'confirmed' the asymptotic results when using random forests. Finite-sample coverage in a single DGP cannot establish a rate condition; the statistical guarantee for the recommended ML implementation is therefore not settled. The theorems themselves are logically valid conditionally, so the issue is not internal inconsistency but an unsupported bridge from abstract assumptions to the practical 'model-agnostic' claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces causal estimands for the expected number of life-years lost due to a specific competing event within a fixed horizon, defines the average treatment effect (ATE) and a variable importance measure based on the best partially linear projection of the CATE, derives their efficient influence functions in the nonparametric model, and constructs cross-fitted one-step estimators with machine-learned nuisance parameters. Theorem 1 and Theorem 2 establish asymptotic linearity under high-level nuisance conditions, and a simulation study plus an application to Danish register data on antidepressants illustrate the methods.","tokens_in":11,"tokens_out":7388,"duration_ms":197388,"significance":"The estimand is interpretable on the original time scale and complements cumulative-incidence-based effect measures, which is a genuine practical advantage. The EIF derivations are detailed and the remainder-term representation in Lemma C.1 is explicit, making the theoretical framework reproducible. The variable importance extension with a formal heterogeneity test is valuable for applied work. The principal weakness is that the advertised model-agnostic guarantees depend on high-level rate conditions, especially Assumption B3 and Theorem 2(ii), that are not shown to hold for the random forest implementations recommended in the simulation and application.","major_comments":[{"comment":"The paper's central claim that the estimators are 'model-agnostic, asymptotically normal, and efficient' for machine-learned nuisance parameters is not established for the recommended implementations. Assumption B3 requires an o_p(n^{-1/2}) remainder term, and the text itself states (Section 4.1 and Section 7) that this is difficult to prove without absolute continuity of the cumulative hazard estimators and is unclear for random survival forests. The simulation results in Section 5 cannot verify an asymptotic rate condition; finite-sample coverage in a single DGP is not evidence that B3 holds. To support the abstract's claim, the authors should either prove B3 for a concrete class of data-adaptive estimators or provide verifiable sufficient conditions, or restrict the theoretical claims to estimators satisfying the high-level assumptions and label the random forest results as an empirical illustration rather than a confirmation.","section":"§4.1 (Assumption B3), §5, §7"},{"comment":"Theorem 2's asymptotic normality for the variable importance estimator assumes that the meta-learned regression of the CATE estimates onto X_{-l} achieves ||\\hat τ_l - τ_l|| = o_p(n^{-1/4}). The authors acknowledge in Section 4.2 that convergence rates for this step are 'generally less known.' Since the p-values in Section 6 are obtained from Theorem 2, the inference guarantee for the recommended procedure is incomplete. Please provide sufficient conditions for the meta-learner that yield condition (ii), or explicitly state that the inference is conditional on an unverified rate.","section":"§4.2, Theorem 2, condition (ii)"}],"minor_comments":[{"comment":"The abstract cites 'Andersen et al. (2013)' but the reference list entry is a single-author paper (Andersen, 2013); please correct the citation.","section":"Abstract and references"},{"comment":"The drug name is misspelled as 'Setraline' in Sections 1 and 6; the correct spelling is 'Sertraline'. Also fix 'its standard' to 'it is standard' (Section 1), 'abbreveations' to 'abbreviations' (Table 1 caption), and 'on the marked' to 'on the market' and 'we constrict ourselves' to 'we restrict ourselves' (Section 6).","section":"Throughout"},{"comment":"The displayed formula for \\hat σ^{2,CF}_{Ω^l_j} is difficult to parse because of the nested parentheses; add explicit brackets or define an intermediate quantity for clarity.","section":"§4.2"},{"comment":"The notation for the counterfactual event time T^a_j is introduced but the construction Y^a_j(t*) = t* - T^a_j ∧ t* could be stated more explicitly so that the relation to the observed years-lost quantity is immediate.","section":"§3"},{"comment":"Lemma C.1 uses the notation φ^a_j(\\hat ν), which is not defined in the main text; please define it or replace it with the explicit expression.","section":"Appendix C, Lemma C.1"}],"recommendation":"major_revision","confidential_remarks":"This is a competent but largely incremental extension of the authors' prior work (Ziersen and Martinussen 2024) to the life-years-lost estimand. The main selling point is the clinically interpretable estimand, not the semiparametric machinery. The consistency of the cross-fitted variance estimator is delegated to Lemma 1 of an unpublished arXiv preprint; this should be made self-contained or supported by a published reference before publication. The fit with the journal depends on whether the editors consider the incremental methodological extension sufficiently novel."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a competent methods paper. The estimand itself has been around since Andersen (2013), but the efficient cross-fitted estimator with flexible nuisance estimation is new, and so is the extension of the best partially linear projection variable importance to competing risks. The influence function derivations are careful and standard, and the explicit remainder representation in Lemma C.1 is a real strength. The simulation study is honest: it compares correctly specified parametric nuisance estimators with random forests, with and without cross-fitting, and the RFCF results show good coverage in that setup. That is genuine evidence of practical behavior, and the paper deserves credit for reporting both the failures of non-cross-fitted RF and the overcoverage in the cross-fitted version.\n\nThe soft spot is exactly what the stress-test note says. Assumption B3 is the load-bearing condition for the ATE theorem, and the authors themselves admit it is not established for Breslow-type or random survival forest estimators. Theorem 2 adds an n^{-1/4} rate for the meta-learned regression of CATE estimates onto X_{-l}, a rate they say is generally unknown. The simulations show finite-sample coverage around 0.95-0.97 for RFCF in one DGP, but that cannot establish a rate condition. So the claims of being 'model-agnostic' and 'efficient' are stronger than what is actually proved. The theorems are logically valid conditionally; the issue is an unsupported bridge between abstract assumptions and the recommended implementation, not an internal contradiction. A careful referee should ask the authors to either verify B3 for their RSF implementation under additional structural assumptions, or temper the practical claims accordingly.\n\nMinor issues: no code or data are provided, and the consistency of the variance estimator relies on a self-cited lemma from their earlier paper. That is acceptable if the cited result is correct, but it would strengthen the paper to state or prove it in the appendix.\n\nWho it is for: biostatisticians and methodological statisticians working on causal inference with competing risks, especially those interested in time-scale estimands and heterogeneous treatment effects. It deserves a serious referee. My recommendation: send it to peer review, and push for a revision that confronts the gap between the rate conditions and the recommended machine learning procedures.","headline":"Solid, conditional semiparametric theory for a useful life-years-lost estimand, but the machine-learning bridge rests on unverified rate conditions and the simulation does not close that gap.","tokens_in":27100,"tokens_out":2139,"would_cite":true,"duration_ms":24059,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D20","62N01","62G20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper derives efficient, machine-learning-based estimators for the causal effect on the number of life years lost due to a specific event, with valid inference under flexible nuisance estimation.","keywords":["Causal inference","number of life years lost","debiased learning","heterogeneity","nonparametric inference","survival data","competing risks","variable importance measure"],"falsifier":"Simulate data with a known true ATE and generate it with the cause-specific hazards as in Section 5, then estimate the ATE with the proposed cross-fitted estimator using random survival forests as nuisance learners and compute the empirical coverage of the confidence interval at increasing sample sizes; if coverage does not approach 0.95 and bias does not decay at $n^{-1/2}$ as $n$ grows, the remainder condition B3 is violated for that learner and the theorem's conclusion fails for that data process.","tokens_in":26162,"feed_emoji":"🩺","tokens_out":9701,"duration_ms":82734,"temperature":0.7,"pith_summary":"This paper introduces a causal version of the number of life years lost due to a specific event, an estimand that reports how much time before a given horizon is lost because a particular cause of death occurs. The authors show that the average treatment effect on this estimand, as well as a variable importance measure for whether the treatment effect changes with a covariate, is identified under standard causal assumptions with competing risks. They derive efficient influence functions and build cross-fitted one-step estimators that remain asymptotically normal when the nuisance functions are estimated with flexible machine learning. The result matters because it puts a clinically interpretable, time-scale treatment effect on the same footing as modern debiased causal inference, and the accompanying heterogeneity test gives a scale-free way to identify which covariates drive differential treatment response.","feed_headline":"Life-years-lost causal effects get valid ML-based inference","feed_subtitle":"Cross-fitted influence-function estimators stay valid when hazards are learned by random forests.","key_machinery":"The engine is the efficient influence function (EIF) of each target parameter combined with K-fold cross-fitting of a one-step estimator. For the ATE, the uncentered EIF is $\\phi_j(\\nu)(O)=\\tau_j(X)+\\{1(A=1)/\\pi(1|X)-1(A=0)/\\pi(0|X)\\}\\sum_{i=1}^2\\int_0^{t^*}H_{ij}(s,t^*|A,X)/S_C(s|A,X)\\,dM_i(s|A,X)$, where $H_{ij}$ is an integral of cumulative-incidence differences and $M_i$ are the cause-specific martingales; for the importance parameter, the EIF is built from the product of this centered ATE influence function with the covariate residual $X_l-E(X_l|X_{-l})$, normalized by the variance denominator. Cross-fitting—splitting the data into $K$ folds and evaluating each fold's score at nuisance estimates from the other folds—controls the empirical process term, so the remaining work is to bound a second-order remainder term and impose $n^{-1/4}$ convergence rates for the meta-learned regressions.","core_discovery":"The central claim is that the average treatment effect on the number of life years lost due to a specific event, defined as $\\psi_j(P)=E\\{L_j(0,t^*|1,X)-L_j(0,t^*|0,X)\\}$ with $L_j$ the integrated cumulative incidence, is estimable with a cross-fitted one-step estimator that is asymptotically linear with influence function $\\tilde{\\psi}_{\\psi_j}$ under Assumption B. For the variable importance measure $\\Omega_l^j$, defined as the best partially linear projection of the conditional average treatment effect onto covariate $X_l$, the paper proves an analogous result with influence function $\\tilde{\\psi}_{\\Omega_l^j}$, and gives an asymptotically standard normal test statistic for $H_0:\\Omega_l^j=0$. The proofs identify the efficient influence functions explicitly, including the inverse-propensity-weighted martingale representation, and show that cross-fitting removes the Donsker-class requirement, leaving only high-level convergence-rate conditions on the nuisance estimators. In simulations, random-forest nuisance estimation with cross-fitting achieves near-nominal coverage, and the application to antidepressant registry data illustrates how the new estimand changes the interpretation from probability differences to days of healthy life lost.","pith_inferences":["Beyond the paper's claims, the explicit form of the remainder term suggests the ATE estimator may be doubly robust—consistent if either the hazard pair or the propensity/censoring pair is correctly specified—which would be worth testing under model misspecification.","Beyond the paper's claims, the method could be turned into a functional estimator by letting the horizon $t^*$ vary and reporting $\\hat{\\psi}_j^{CF}(t^*)$ as a curve, providing a treatment-effect trajectory rather than a single summary.","Beyond the paper's claims, because the importance parameter is scale-dependent, an explicit recommendation would be to standardize covariates before computing $\\Omega_l^j$ when comparing importance across variables measured in different units, or to rely solely on the p-value ranking as the application does.","Beyond the paper's claims, pairing the EIF-based estimator with a stacking ensemble that explicitly targets the $n^{-1/4}$ rate for the meta-learned CATE regression might make the asymptotic guarantees more robust in small samples than a single black-box learner."],"forward_implications":["Applied researchers can use random survival forests or other flexible learners for the cause-specific hazards, censoring, and propensity score without giving up valid confidence intervals, provided the high-level rate conditions hold; the simulation study shows near-nominal coverage for the cross-fitted version.","The test based on $T_{ST}^l$ is asymptotically standard normal under the null $\\Omega_l^j=0$, so it provides a scale-free, p-value-based ranking of covariates by their contribution to treatment-effect heterogeneity in the presence of competing risks.","Because the estimand is defined directly on the time scale of the data, the ATE is reported in units such as days of healthy life lost, which is more directly interpretable to clinicians than differences in cumulative incidence probabilities.","The estimator avoids the need for correctly specified parametric models for the conditional average treatment effect or the censoring distribution, unlike pseudo-observation-based competitors.","The framework applies to any specified competing event, and the paper notes that recomputing the importance ranking with event and competing event swapped can reveal whether an apparent effect on one cause is driven by the other."],"supporting_citations":[{"why":"Defines the decomposition of years of life lost by cause and the estimand $L_j(0,t^*|a,x)$ that this paper makes causal.","marker":"Andersen (2013)"},{"why":"Introduces the best partially linear projection variable importance measure and its semiparametric estimators, the template extended here to the life-years-lost CATE.","marker":"Ziersen and Martinussen (2024)"},{"why":"Supplies the double/debiased machine-learning framework and cross-fitted one-step estimation that underlies the estimators' construction and asymptotic justification.","marker":"Chernozhukov et al. (2018)"},{"why":"Provides the cross-fitting decomposition and the empirical-process/remainder propositions used in the proofs of Theorems 1 and 2.","marker":"Kennedy (2022)"},{"why":"Gives the influence-function and functional delta-method theory that turns the derived EIFs into asymptotic normality results.","marker":"van der Vaart (2000)"},{"why":"Supplies the random survival forests used as the data-adaptive nuisance estimators in simulations and the application.","marker":"Ishwaran et al. (2008)"},{"why":"Provides the Danish registry study of antidepressant response that the application section analyzes.","marker":"Kessing et al. (2024)"},{"why":"Derives asymptotics for pseudo-observation-based causal survival estimators, the parametric-model-based competitor that the proposed estimator improves upon.","marker":"Overgaard et al. (2019)"}],"fun_headline_variants":["Causal life-years lost: efficient ML-based inference","Life-years lost: causal effects with efficient ML estimators","Efficient causal estimation of life-years lost and variable importance","Causal life-years lost: new estimand and ML-based inference","Life-years lost: causal estimands and variable importance via ML"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All the stated guarantees depend on unverifiable high-level rate conditions: the ATE needs the nuisance remainder to vanish at root-n speed, and the heterogeneity test additionally needs the meta-learned regression of treatment-effect estimates onto other covariates to be accurate to within $n^{-1/4}$; if the chosen machine learners do not meet these rates in a given data-generating process, the asymptotic normality and standard errors do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Causal life-years lost: efficient ML-based inference","Life-years lost: causal effects with efficient ML estimators","Efficient causal estimation of life-years lost and variable importance","Causal life-years lost: new estimand and ML-based inference","Life-years lost: causal estimands and variable importance via ML"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000647,"raw_usage":{"total_tokens":2990,"prompt_tokens":979,"completion_tokens":2011,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":1929}},"tokens_in":595,"tokens_out":2011,"duration_ms":12767,"temperature":1.0,"reasoning_tokens":1929,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T20:07:15.340182+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate data with a known true ATE and generate it with the cause-specific hazards as in Section 5, then estimate the ATE with the proposed cross-fitted estimator using random survival forests as nuisance learners and compute the empirical coverage of the confidence interval at increasing sample sizes; if coverage does not approach 0.95 and bias does not decay at $n^{-1/2}$ as $n$ grows, the remainder condition B3 is violated for that learner and the theorem's conclusion fails for that data process.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the best partially linear projection variable importance measure and its semiparametric estimators, the template extended here to the life-years-lost CATE."}],"review_version":1}