{"id":"38397e42-8f8f-4aef-bb3d-9c15518b893e","arxiv_id":"2506.07910","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A structural nested rate model is proposed to estimate immediate and lagged causal effects of time-varying exposures on recurrent events while accounting for a correlated terminal event.","lead":"This paper develops a statistical method for estimating the short-term and delayed causal effects of a time-varying exposure on repeated health events when some people die during follow-up. The method is applied to 299,661 Medicare patients to estimate how monthly fine particulate matter (PM2.5) affects recurrent cardiovascular hospitalizations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Delayed-effect estimates (m≥1) hinge on the parametric additive-hazard death model encoded in weight (7); the simulation never misspecifies it, and a wrong death model biases every β_m, m≥1, at root-n, including the headline cumulative PM2.5 effect.","rationale":"I read the paper as claiming: under the structural nested rate model (1), no unmeasured confounding, positivity, and the survival-weight model (7), the two estimators are asymptotically linear with the stated influence functions, and the Medicare analysis estimates a cumulative 27.3 per 100,000 additional CVD hospitalizations for a 10 µg/m3 increase in PM2.5 over lags 0-6. The least secure condition for that claim is the weight model (7): it is parametric, needed for every delayed effect, untested against misspecification in the simulations, and untouched by the robust estimator. This is exactly the reader's weakest_assumption, hence agreement is 'agree'. I stress-tested possible stronger objections and do not find them decisive: the Remark 4.2 algebraic error (multiplication instead of division by p(B|C)) is real but repairable, since after correction Assumption 1b still renders the quantity independent of A_k; Assumption 3's rate conditions are high-level but standard for cross-fitted estimators, and the appendix provides a verification lemma (Lemma S1); the survival-conditional interpretation of the target parameter is stated in Section 6.1, though the abstract could mislead. The paper's strengths should be credited: the sequential influence-function derivation with propagation of prior β_j terms, the honest simulation reporting, the R package sncure, and the first application of delayed-effect recurrent-event estimation to PM2.5 and CVD hospitalizations. The verdict stays conditional acceptance with required revisions: add a misspecified death-process sensitivity analysis (or qualify the causal claims), correct Remark 4.2, supply primitive rate conditions for Assumption 3, and clarify in the abstract that the estimand is a survival-conditional rate effect rather than a population marginal effect.","tokens_in":33417,"tokens_out":18017,"duration_ms":185150,"concrete_test":"Re-run the Section 5 simulation (n=5000, simple exposure scenario) with the death process changed to a proportional hazards model λ_D(t) = λ_0(t) exp(0.02A_k + 0.01A_{k-1}), calibrating the baseline λ_0 to keep cumulative mortality comparable, while leaving the recurrent event, exposure, and covariate processes as in the paper. Compute √n-bias and coverage for β_0 through β_4 under the parametric and robust estimators. If m=0 bias stays small while |√n-bias| for m≥1 grows materially relative to Table 1, the weight model is load-bearing and the causal claims must be qualified; if bias remains flat, the concern is mitigated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"For m≥1, the unbiasedness of estimating equations (5)-(6) requires weight w_km(t) in (7) to equal the true ratio of counterfactual to observed survival, p{D^(A_{k-m},0) ≥ t | A_{k-m}, L_{k-m}, D ≥ k} / p{D ≥ t | A_{k-m}, L_{k-m}, D ≥ k}. Equation (7) imposes an additive-hazard structural nested cumulative survival model on the death process (Seaman et al., 2020): the survival ratio is the exponential of a linear combination of lagged exposures with ν_jk(t) of the form (1,...,1,t-k), forcing the already occurred effect of A_{k-j} and the current effect of A_k to share parameters. This is the least secure condition because it is a nuisance assumption rather than the target of inference, because neither estimator relaxes it (the robust estimator in Theorem 4.4 nonparametrics only μ and dρ, keeping the parametric weight), and because the Section 5 simulation always generates death from λ̃(t) = 0.02A_k + 0.01A_{k-1} + Q exp{L_1k + L_2k - 1}, precisely the model class that makes (7) correct, with α always fit under the true model. The headline application result, 27.3 additional CVD hospitalizations per 100,000 (95% CI 2.7, 51.9) over lags 0-6, accumulates β_1 through β_6, all of which require the weight. If death instead followed a proportional hazards model λ_0(t) exp(γ'A), a common mortality model in this literature, the moment conditions E[Y(t)w_km(t)Δ_km(t)dN_0^{(m)}(t)] = 0 would fail and every delayed estimator would carry root-n bias. Secondary but related: Remark 4.2 writes E[Z|B] = E[Z|C] p(B|C), whereas the correct identity divides by p(B|C); the conclusion survives the fix, but the proof as printed is unsound. Assumption 3 also states rate conditions with no primitive conditions supplied, leaving the applicability of Theorem 4.4 to the implemented SuperLearner procedure unverifiable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a class of semiparametric structural nested rate models (SNCURE) for estimating short-term and delayed causal effects of time-varying exposures on recurrent event outcomes when a correlated terminal event (death) is present. Two estimators are developed: one with parametric exposure models and one that allows nonparametric estimation of the exposure model and baseline rate nuisance functions via crossfitting. The main theoretical claims are asymptotic linearity of both estimators, with influence functions stated in Theorems 4.3 and 4.4. The methods are evaluated in simulations and applied to estimate effects of monthly PM2.5 exposure on recurrent cardiovascular hospitalizations among Medicare beneficiaries, reporting a cumulative effect over six months of 27.3 additional hospitalizations per 100,000 (95% CI 2.7, 51.9). An R package, sncure, is provided.","tokens_in":33829,"tokens_out":12481,"duration_ms":143311,"significance":"If the theoretical claims are correct, this is a useful and novel contribution: it extends structural nested cumulative survival models to recurrent event processes in the presence of a terminal event, and it is, to the authors' knowledge, the first approach to estimate both short-term and delayed marginal causal effects in this setting. The explicit influence-function derivations go beyond prior work on structural nested survival models, and the availability of software and the realistic Medicare application increase the paper's practical value. The asymptotic results are supported by supplement proofs and the simulation study covers both linear and nonlinear exposure mechanisms. However, the delayed-effect estimators (m≥1) rest on a parametric additive-hazard model for the death process whose misspecification is never examined, and a key displayed argument for unbiasedness in Remark 4.2 contains an algebraic error. These issues are load-bearing for the paper's central claim and need to be addressed before the results can be fully trusted.","major_comments":[{"comment":"The displayed factorization justifying unbiasedness for m≥1 is algebraically incorrect. For a counting-process increment X=dN^{(A_{k-1},0)}(t), which is zero whenever D^{(A_{k-1},0)}<t, the correct identity is E[X | A_k,L_k,D^{(A_{k-1},0)}≥t] = E[X | A_k,L_k,D^{(A_{k-1},0)}≥k] / P(D^{(A_{k-1},0)}≥t | A_k,L_k,D^{(A_{k-1},0)}≥k), not the product shown in the remark. The same error appears in the second equality. Because this remark is the stated justification for the unbiasedness of estimating equations (5) and (6), the proof of Theorems 4.3 and 4.4 for m≥1 is not rigorous as written; the authors should provide a correct derivation of the moment condition P0[Y(t)w_{km}(t)Δ_{km}(t)dN_0^{(m)}(t)]=0, for example by showing how the weight in (7) converts the observed risk-set indicator into the counterfactual survival probability.","section":"Remark 4.2"},{"comment":"For every m≥1, the estimating equations require the weight w_{km}(t) in (7) to equal the true ratio of counterfactual to observed survival, which is equivalent to assuming an additive-hazard structural nested cumulative survival model for the death process. This is a nuisance assumption, not the target of inference, and it is not probed by the simulation: the death intensity in Section 5 is λ̃(t)=0.02A_k+0.01A_{k-1}+Q exp{L_{1k}+L_{2k}-1}, which is exactly the model class that makes (7) correct, and α_m is always fit under this true model. Consequently, the simulation provides no evidence on finite-sample behavior when death follows, say, a proportional hazards model, under which the moment conditions for β_m (m≥1) would generally fail. The headline cumulative PM2.5 estimate in Section 6 accumulates β_1 through β_6 and therefore inherits this untested sensitivity. The authors should add a misspecification scenario for the death model or clearly state this limitation and temper the corresponding conclusions.","section":"Eq. (7) and Section 5"},{"comment":"The description of the second estimator as 'robust' is more limited than the text suggests: Theorem 4.4 relaxes parametric assumptions on μ_{km} and dρ_{km}, but the weight function w_{km}(t) remains a parametric model estimated under Assumption 2.b. Since the unbiasedness of the m≥1 estimating equations depends on correct specification of this weight, the robust estimator is not robust to misspecification of the death process. The abstract and Section 7 should be revised to state this qualification, and the practical implication for the Medicare analysis should be acknowledged.","section":"Theorem 4.4"},{"comment":"The simulation includes non-administrative censoring generated from λ̃_C(t)=0.2 exp{L_{1k}+L_{2k}-1}, where L_k is time-varying and is used to generate exposure A_k. Section 4.5 states that, under such censoring, the original model holds only under two conditions, and that the second condition can be relaxed by multiplying w_{km}(t) by a censoring weight w^C_{km}(t). The simulation description does not state whether this censoring weight was used. If it was not used, the censoring is informative (because L_k is correlated with exposure), and this could bias the reported results; if it was used, the authors should say so explicitly and describe how w^C was estimated.","section":"Section 5"}],"minor_comments":[{"comment":"The sentence 'they consistent and asymptotically normal' is missing a verb; it should read 'they are consistent and asymptotically normal.'","section":"Section 7"},{"comment":"The reference 'World Heath Organization' should be 'World Health Organization'.","section":"Reference list"},{"comment":"The notation in (13) is confusing: the counterfactual exposure history a* is defined as A*_{k'}=A_{k'}∧a, but the summation index k and the counterfactual level a are not clearly distinguished from the scalar exposure threshold used in the intervention description; please clarify the notation.","section":"Equation (13)"},{"comment":"In the complex-exposure scenario, the parametric estimator's √n-bias for β_0 increases from -0.69 at n=2000 to -0.85 at n=5000, which is consistent with misspecification-induced inconsistency; this is not discussed in the text and should be interpreted explicitly.","section":"Table 1"},{"comment":"The simulation description states that exposures are generated using two scenarios and then 'normalized such that A_k∈[0,1]', but it is not specified whether the normalization is applied per subject, per time, or globally; this affects interpretation of the scaling factor c in the event rate and should be stated.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and, if the proof issue and death-model sensitivity are addressed, could be a solid contribution. The main concern for the editor is that the key Remark 4.2 contains an algebraic error that directly affects the proof of the central asymptotic results for m≥1, and the simulation never examines misspecification of the death model on which all delayed-effect estimates rely. I do not see evidence of questionable research practices; the issues are technical and scientific rather than ethical. I recommend major revision rather than rejection because the errors appear fixable and the proposed framework is potentially valuable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a real extension of Seaman et al.'s structural nested cumulative survival model to recurrent events with a competing death process. It adds influence functions, a cross-fitted nonparametric estimator, an R package, and a large Medicare application. That is a solid contribution and worth referee time.\n\nWhat it does well: the model is clearly motivated, the two estimators are sensible (parametric exposure model, plus a robust version that nonparametrics both the exposure and the baseline rate), and the simulation is thorough enough to show root-n behavior and reasonable coverage. The Medicare analysis is the first to estimate delayed PM2.5 effects on recurrent CVD hospitalizations, and the cumulative effect estimate (27.3 per 100,000 over six months) is directly policy-relevant. The authors also ship code and data-analysis details, which makes the work reproducible.\n\nWhere it's soft: the delayed-effect estimates for m≥1 all rest on the weight function w_km(t) in Eq (7), which imposes a specific parametric additive-hazard structure on the death process. This is a nuisance assumption, not the target of inference, and the simulation never misspecifies it—death is always generated from exactly the same model class. If death followed, say, a proportional hazards model, the moment conditions would fail and every β_m for m≥1 would carry root-n bias, including the headline cumulative effect. That is a load-bearing soft spot, not cosmetic.\n\nThere's also a concrete proof error: Remark 4.2 writes E[Z|B] = E[Z|C] p(B|C), but the correct identity divides by p(B|C). The conclusion likely survives the fix, but the printed argument is unsound. And Assumption 3 states rate conditions with no primitive conditions supplied, so the nonparametric theorem's applicability to the implemented SuperLearner is unverified.\n\nNone of these are fatal—the central idea is sound and the paper is honestly written. But the derivation needs correcting, the death-model misspecification needs a simulation or sensitivity analysis, and Assumption 3 needs either primitive conditions or a clear statement that this is an open condition.\n\nFor you: if you work on recurrent events or causal inference with time-varying exposures, this is worth reading and probably citing after revision. I'd bring it to reading group and I'd send it to a serious referee, with the expectation of major revision.","headline":"A genuinely useful extension of structural nested models to recurrent events, but the delayed-effect causal estimates lean on an untested parametric death model and a proof slip that needs fixing.","tokens_in":34457,"tokens_out":2651,"would_cite":true,"duration_ms":31518,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G20","62D20","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A new structural nested rate model isolates short-term and delayed causal effects of time-varying exposure on recurrent events while accounting for the competing risk of death.","keywords":["causal inference","recurrent events","structural nested rate models","time-varying exposure","terminal event","semiparametric estimation","asymptotic linearity","PM2.5 cardiovascular hospitalizations"],"falsifier":"Simulate recurrent-event data under a death process that violates the multiplicative weight model of equation (7), for example an additive hazard with exposure-squared terms or non-proportional effects, and check whether the delayed $\\tilde{\\beta}_m$ estimates remain centered on their true values. The paper's own simulations always fit the weight model in the correct family, so a misspecification comparison would directly test the robustness of the delayed-effect estimators.","tokens_in":33205,"feed_emoji":"🌫️","tokens_out":8985,"duration_ms":103268,"temperature":0.7,"pith_summary":"The paper aims to estimate the causal effect of a time-varying exposure on a recurrent event outcome such as repeated hospitalizations when some people die before follow-up ends and death is correlated with both exposure and events. Existing inverse-probability-weighted and marginal structural estimators cannot cleanly deliver delayed exposure effects or handle the competing risk of death, so the authors build a class of structural nested rate models in which each lag's effect is blipped out of the observed event rate sequentially. They prove that two estimators, one with a parametric exposure model and one that lets the exposure and nuisance models be fitted by flexible machine-learning tools, are asymptotically linear, meaning their bias shrinks at the usual $\\sqrt{n}$ rate and bootstrap confidence intervals are valid. Simulations support the theory, and the method is applied to 299,661 Medicare beneficiaries, estimating that a $10\\,\\mu$g/m$^3$ increase in PM2.5 over the current and six prior months corresponds to 27.3 additional cardiovascular hospitalizations per 100,000 people, with 95% CI 2.7 to 51.9.","feed_headline":"PM2.5 over seven months adds 27.3 CVD hospitalizations per 100,000","feed_subtitle":"A new structural nested rate model separates lagged exposure effects from the competing risk of death in Medicare data.","key_machinery":"The load-bearing object is the structural nested cumulative number of recurrent events (SNCURE) model, equation (1): for each lag $m$, the counterfactual event rate under an exposure history with the most recent exposure fixed and later exposures set to zero equals an exposure-free baseline rate minus $A_{k-m}\\tilde{\\beta}_m\\,dt$. Estimation works through nested blip-down estimating equations (4)-(6) that remove estimated effects of more recent exposures before extracting the next lag, using the residual deviation $A_{k-m}-\\mu_{km}(t)$ as an instrument. For $m\\ge 1$ the equations are weighted by $w_{km}(t)=\\prod_{j=0}^{m-1}\\exp\\{A_{k-j}\\nu_{jk}(t)'\\alpha_m\\}$, a parametric model for the ratio of counterfactual to observed survival that reweights the at-risk set to match the counterfactual population. The robust estimator adds a transformation that subtracts the conditional mean of the blipped process, so both the exposure model and the nuisance baseline rate can be fitted nonparametrically; crossfitting removes the Donsker conditions usually needed for such proofs. The asymptotic linearity results in Theorems 4.3 and 4.4 are obtained by empirical-process expansions and yield explicit influence functions.","core_discovery":"The central claim is that the parameters $\\tilde{\\beta}_m$ in the structural nested rate model (1) are identifiable from observational recurrent-event data and can be estimated at root-$n$ rate in the presence of a possibly dependent terminal event. For $m=0$ the model isolates the short-term effect of the current exposure; for $m>0$ it nests models so that the effect of exposure $m$ months earlier is recovered after blipping down the effects of more recent exposures. The paper's first estimator requires a correctly specified parametric exposure model but leaves the event-rate model unspecified, while the second estimator also allows the exposure and baseline rate functions to be estimated nonparametrically and uses crossfitting. Theorems 4.3 and 4.4 give the influence functions and establish asymptotic linearity under regularity conditions; simulations show $\\sqrt{n}$-bias near zero and coverage near the nominal level. The Medicare application reports cumulative lagged effects, with the headline estimate of 27.3 additional CVD hospitalizations per 100,000 beneficiaries for a $10\\,\\mu$g/m$^3$ increase in PM2.5 over seven months.","pith_inferences":["If the death-process weight model is misspecified in real data, delayed-effect estimates will likely be biased; a natural extension is a sensitivity analysis that fits the death hazard flexibly and compares the resulting $\\beta$ estimates.","The at-risk process counts hospital days as time at risk; in populations with frequent or long hospitalizations, estimates would be attenuated, so adjusting $Y(t)$ for hospital-stay gaps would be an important refinement.","Because the target parameters are marginal effects, the framework maps naturally onto cost and regulatory analyses that compare counterfactual exposure limits, not just single-endpoint studies."],"forward_implications":["Short-term and delayed exposure effects on recurrent events can be estimated separately even when the terminal event is correlated with exposure, which previous recurrent-event estimators did not provide.","The robust estimator allows flexible machine-learning estimation of exposure and nuisance models while retaining root-$n$ inference, so the method is applicable when parametric exposure models are hard to specify.","The influence-function results extend to the structural nested survival-time setting as a special case, supplying variance formulas where none were available.","The Medicare results give a concrete regulatory quantity: limiting monthly average PM2.5 to 9 $\\mu$g/m$^3$ would correspond to roughly 528 fewer CVD hospitalizations, a 6.8% reduction, in the study cohort.","The provided R package sncure makes the two estimators and bootstrap inference available for other recurrent-event studies."],"supporting_citations":[{"why":"Supplies the structural nested cumulative survival time framework and the parametric additive-hazard weight model $w_{km}(t)$ that the estimators use to adjust for the terminal event.","marker":"Seaman et al. (2020)"},{"why":"Establishes point-exposure g-estimation on an additive hazard scale, the basis for the rate-model blip-down equations.","marker":"Martinussen et al. (2011)"},{"why":"Introduces structural nested models and the g-estimation rationale used to define the estimating equations.","marker":"Robins and Greenland (1994)"},{"why":"Provides the crossfitting lemma used to control the remainder terms in the nonparametric asymptotic theory.","marker":"Ertefaie et al. (2021)"},{"why":"Gives the conditions under which non-administrative censoring can be treated as ignorable in the proposed estimating equations.","marker":"Vansteelandt and Sjolander (2016)"},{"why":"Supplies the validated ensemble PM2.5 exposure predictions used in the Medicare application.","marker":"Di et al. (2019)"}],"fun_headline_variants":["New model isolates PM2.5 effect on repeat heart hospitalizations","Method separates pollution effect from death risk in recurrent events","Structural nested model recovers delayed PM2.5 impact on CVD readmissions","Medicare data: PM2.5 lifts CVD hospitalizations after accounting for death","Tool estimates causal effect of time-varying exposure on recurrent events"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The delayed-effect estimates are unbiased only if the parametric formula for how past exposures affect the probability of surviving to each time point is exactly correct, and only if no unmeasured confounders affect both exposure and the recurrent event process; the paper assumes both rather than testing them.","fun_headline_variants_meta":{"raw":{"variants":["New model isolates PM2.5 effect on repeat heart hospitalizations","Method separates pollution effect from death risk in recurrent events","Structural nested model recovers delayed PM2.5 impact on CVD readmissions","Medicare data: PM2.5 lifts CVD hospitalizations after accounting for death","Tool estimates causal effect of time-varying exposure on recurrent events"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000662,"raw_usage":{"total_tokens":3047,"prompt_tokens":990,"completion_tokens":2057,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":1966}},"tokens_in":606,"tokens_out":2057,"duration_ms":19594,"temperature":1.0,"reasoning_tokens":1966,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:25:02.738886+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate recurrent-event data under a death process that violates the multiplicative weight model of equation (7), for example an additive hazard with exposure-squared terms or non-proportional effects, and check whether the delayed $\\tilde{\\beta}_m$ estimates remain centered on their true values. The paper's own simulations always fit the weight model in the correct family, so a misspecification comparison would directly test the robustness of the delayed-effect estimators.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the structural nested cumulative survival time framework and the parametric additive-hazard weight model $w_{km}(t)$ that the estimators use to adjust for the terminal event."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes point-exposure g-estimation on an additive hazard scale, the basis for the rate-model blip-down equations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces structural nested models and the g-estimation rationale used to define the estimating equations."},{"cited_title":"R., Oslin, D., and Strawderman, R","cited_arxiv_id":null,"evidence_quote":"Provides the crossfitting lemma used to control the remainder terms in the nonparametric asymptotic theory."},{"cited_title":"and Sjolander, A","cited_arxiv_id":null,"evidence_quote":"Gives the conditions under which non-administrative censoring can be treated as ignorable in the proposed estimating equations."}],"review_version":1}