{"id":"b4c2aacb-bd88-41aa-a7b5-3343634d6268","arxiv_id":"1908.03646","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Sharp and fuzzy regression discontinuity estimators for censored survival outcomes are constructed using IPCW and doubly robust censoring unbiased transformations, with asymptotic normality and efficiency results.","lead":"This statistics paper extends regression discontinuity designs, a causal inference tool, to survival outcomes that are only partially observed due to censoring. It provides estimators, asymptotic theory, and an application to prostate cancer screening data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Data-driven bandwidth is used and inferred on without asymptotic justification, so Theorem 1 does not cover the implemented estimator; Section 7 admits this gap.","rationale":"In good faith, the paper has real content: it extends censoring-unbiased transformations to RD, states a plausible fixed-bandwidth asymptotic theory (Theorem 1 and Corollary 1), gives explicit variance formulas, and supports the estimators with simulations and a real-data example. The most load-bearing gap is not the censoring transformation itself but the inference procedure actually implemented: the data-driven bandwidth is used in the estimator, the variance estimator, and the simulations, yet no theorem covers it. The authors' own Section 7 limitation statement confirms this. The reader identified exactly this weakest assumption; I agree. The concern is addressable—either develop the bandwidth asymptotics or verify empirically that inference is insensitive—so it does not warrant rejection. Since the reader's verdict is already CONDITIONAL with moderate confidence, I recommend no change to the verdict.","tokens_in":28526,"tokens_out":9527,"duration_ms":105418,"concrete_test":"Run 1,000 replications of the sharp RD setting in Section 5 at n=400 and compare empirical coverage of nominal 95% plug-in CIs using (a) fixed h=C n^{-1/5} for several constants C and (b) the LM ĥ from Section 4.3. If coverage for (b) deviates from the fixed-h range by more than the Monte Carlo standard error, the missing bandwidth theory is consequential. In parallel, verify analytically whether ĥ/h→1 in probability under conditions (C1)–(C7) and (R1)–(R9); if this fails, derive the distributional effect of bandwidth selection before relying on the variance estimator.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1 and Corollary 1 assume a fixed bandwidth with h∼n^{-1/5} (condition R8), and Lemmas 1–6 treat h as nonrandom. Section 4.3 replaces h with the LM cross-validation minimizer ĥ (and with min{ĥ_DR, ĥ_Z} in the fuzzy case), and Section 4.4 then plugs ĥ into both the estimator and the variance formulas. No theorem establishes that ĥ/h→1 in probability or that the limiting distribution in Theorem 1 survives data-dependent bandwidth selection. The paper itself says in Section 7 that asymptotic theory for the adapted LM bandwidth 'is currently under investigation.' This is load-bearing because every simulation result in Section 5 and the PLCO application in Section 6 use the data-driven ĥ. If ĥ has non-negligible sampling variability, the variance estimator omits its contribution and the nominal coverage in Tables 1–2 is not justified by the stated theory. The fuzzy-RD rule of taking the minimum of two CV bandwidths is especially ad hoc: even if each individual ĥ were asymptotically equivalent to a fixed sequence, the minimum of the two estimators is not automatically covered by the same argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes estimators of sharp and fuzzy regression discontinuity treatment effects when the outcome is a censored survival time. It applies inverse probability censoring weighted and doubly robust censoring-unbiased transformations to the log survival time, then uses local linear RD estimators on the transformed outcomes. The main theoretical results (Theorem 1, Corollary 1) claim n^{2/5}-asymptotic normality with explicit bias and variance after replacing unknown censoring and survival distributions by uniformly consistent estimators. Theorem 2 claims an efficiency ordering in favor of the DR estimator with correctly specified models. Bandwidth selection adapts the Ludwig–Miller cross-validation criterion; inference uses plug-in or nearest-neighbor variance estimators and the bootstrap. The methods are evaluated in simulations and applied to the PLCO prostate cancer screening data.","tokens_in":28802,"tokens_out":9190,"duration_ms":96693,"significance":"If the distributional results were valid for the implemented estimator, the paper would make a useful contribution: it extends RD designs to censored outcomes, brings doubly robust efficiency improvements into the RD setting, and offers a practical implementation route through existing software such as rdrobust. The paper has clear strengths: explicit regularity conditions, an extension of the Hahn, Todd and Van der Klaauw proof strategy, reliance on established doubly robust transformation results, and a real-data illustration. However, the load-bearing inferential claims are not supported by the stated theorems because the data-driven bandwidth is used without asymptotic justification and the bias term in Theorem 1 is not addressed in the confidence-interval construction. In addition, the fuzzy RD efficiency theorem appears to contain a sign error in its covariance condition. These issues currently prevent the manuscript from supporting its inferential claims at the level stated.","major_comments":[{"comment":"The asymptotic theory assumes a fixed bandwidth with h∼n^{−1/5} (condition R8) and treats h as nonrandom in Lemmas 1–6. The implemented procedure in Section 4.3 chooses ĥ by the adapted Ludwig–Miller criterion, and Section 4.4 plugs ĥ into both the estimator and the variance formulas. No theorem shows ĥ/h→1 in probability or that the limiting distribution in Theorem 1 survives data-dependent bandwidth selection; the manuscript’s Section 7 states that this asymptotic theory is 'currently under investigation.' Because every simulation in Section 5 and the PLCO analysis in Section 6 use ĥ, the coverage probabilities in Tables 1–2 and the standard errors in Table 3 are not justified by the stated theory. The fuzzy case is worse: the bandwidth is min{ĥ_DR, ĥ_Z}, and the minimum of two data-dependent bandwidths is not covered by an argument that applies to each separately.","section":"§4.3–4.4, Theorem 1 (condition R8)"},{"comment":"Theorem 1 gives n^{2/5}(τhat−τ−φ) → N(0,Σ) with a nonzero bias φ of order n^{−2/5}. The variance estimator in Section 4.4 estimates only the asymptotic variance; the reported confidence intervals are centered at τhat with no estimation or correction for φ. Since the bandwidth is of the optimal order n^{−1/5}, the bias contributes a first-order term to coverage error, and no undersmoothing or explicit bias-correction argument is supplied. Thus the normal-approximation intervals in Tables 1–2 are not a consequence of Theorem 1 even if the bandwidth were fixed.","section":"§4.4 and Theorem 1"},{"comment":"The variance formulas attached to Theorem 1 in the Supplement have first coefficient 1/τZ in the fuzzy RD variance, whereas the delta-method variance of the ratio τhat_Y/τhat_Z is (1/τZ^2)Var(τhat_Y) − 2τY/τZ^3 Cov + (τY^2/τZ^4)Var(τhat_Z). Section 4.4 itself uses 1/τhat_Z^2, matching the delta method. Consequently the Σ definitions in the Supplement do not correspond either to the delta method or to the implemented variance estimator. If the Supplement definitions were used for inference, the confidence intervals would be wrong.","section":"Supplementary Materials, definitions of Σ_IPCW_FRD and Σ_DR_FRD; §4.4"},{"comment":"The fuzzy RD efficiency claim is not established by the given proof. With the delta-method variance (or even the Supplement’s version), the difference AVar(DR)−AVar(IPCW) contains a term −2τY/τZ^3 times the difference in the boundary covariances η_DR−η_IPCW. The theorem assumes η_DR≤η_IPCW, so that difference is nonpositive, which makes the covariance term positive (for τY,τZ>0). Hence the displayed reasoning 'we have ΣDR≤ΣIPCW' does not follow; smaller covariance with the treatment indicator increases, not decreases, the variance of the ratio estimator. The covariance condition appears to have the wrong sign, and the theorem as stated is unsupported.","section":"Theorem 2 (fuzzy RD) and its proof in the Supplement"}],"minor_comments":[{"comment":"The text says 'Let G0 and S0 be true distributions of failure and censoring times, respectively,' but the Supplement and subsequent use define G0 as censoring and S0 as failure; please correct the wording.","section":"§4.2"},{"comment":"There are several typos, including 'resepctively', 'addtion', and 'in addtion'; the manuscript should be carefully proofread.","section":"§4.4"},{"comment":"The nearest-neighbor variance estimator depends on a number of neighbors K that is never specified in the main text or simulations; please clarify the choice of K.","section":"§4.4 and §5"},{"comment":"The bootstrap in the simulations uses only B=50 resamples for bootstrap standard errors and coverage; this is small for coverage estimation and should be stated with a justification.","section":"§5"},{"comment":"The data-driven RD plots are described as useful even though the theory of Calonico, Cattaneo and Titiunik (2015a) does not directly apply to the transformed responses; the paper should either justify this or present the plots as purely descriptive.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The self-identified gap about bandwidth selection in Section 7 is the most significant issue; the authors should either develop the bandwidth asymptotics or substantially temper the inferential claims in the simulations. The sign error in Theorem 2 should be fixed before any efficiency claims are made. Given the current state, the paper is not publishable as is, but the central idea is promising and a careful revision could address the issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a solid first combination of censoring-unbiased transformations (IPCW and the doubly robust version) with local linear RD estimation, and it backs the method with asymptotic normality results for sharp and fuzzy designs. The authors are honest about the main weakness: Theorem 1 and Corollary 1 assume a fixed bandwidth h~n^{-1/5}, but Section 4.3 computes bandwidth by the modified Ludwig-Miller criterion and Section 4.4 plugs that h-hat into both the estimator and the variance formulas. No theorem shows h-hat behaves like the fixed sequence, and Section 7 says the theory \"is currently under investigation.\" This is not a minor nit: every simulation in Section 5 and the PLCO application use the data-driven bandwidth, so the coverage probabilities in the tables are not actually covered by the stated asymptotics. In a field where bandwidth selection is known to be delicate, this needs a fix or at least a clear redeclaration of what the theory does and does not support.\n\nWhat the paper does well: the combination is new, the nuisance estimation (KM for censoring, parametric AFT for the conditional expectation) is standard and reproducible, and the simulations show the DR transform beating IPCW in bias and efficiency, matching the variance formulas. The efficiency comparison in Theorem 2 is a nice extension of Steingrimsson et al. The PLCO application is a sensible showcase, though I would flag that labeling it \"sharp RD\" is questionable given that a PSA>=4 only triggered a recommendation; there is likely noncompliance, so a fuzzy RD treatment would have been more appropriate. That is a moderate issue, not fatal.\n\nThe proofs in the supplement are sketched and have typos (some brackets and terms), but the structure tracks Hahn-Todd-van der Klaauw and the key unbiasedness lemmas are standard. I believe the central argument is right, but the bandwidth gap keeps me from treating the empirical claims as fully justified.\n\nWho this is for: anyone doing causal inference in survival settings, especially with a clinical cutoff. It deserves a serious referee. I would send it out with the request that the authors either supply asymptotics for the data-driven bandwidth or state the fixed-h caveat in the abstract and conclusions and refrain from claiming nominal coverage in the simulations.","headline":"Useful first bridge between censored-outcome transformations and RD designs, but the data-driven bandwidth used in practice is not covered by the stated theory—a gap the paper itself admits.","tokens_in":29282,"tokens_out":2466,"would_cite":true,"duration_ms":27324,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G20","62N01"],"pacs":[],"model":"deepseek-v4-flash","headline":"Regression discontinuity designs can estimate causal effects from censored survival data.","keywords":["regression discontinuity design","censored data","survival analysis","causal inference","inverse probability censoring weighting","doubly robust estimation","local linear regression","bandwidth selection"],"falsifier":"Simulate the sharp RD design with right censoring, compute the proposed estimator under the paper's data-driven bandwidth, and compare the empirical coverage of nominal 95% intervals with coverage when the bandwidth is fixed at $h = c n^{-1/5}$. If the data-driven version departs materially from 0.95 while the fixed-bandwidth version does not, the point of failure is the assumption that the estimated bandwidth can be ignored in the asymptotics.","tokens_in":28354,"feed_emoji":"🩺","tokens_out":9954,"duration_ms":95548,"temperature":0.7,"pith_summary":"This paper argues that regression discontinuity designs, usually developed for uncensored outcomes, extend naturally to right-censored time-to-event data. The insight is to replace each censored log-survival time with a censoring-unbiased transformation, either inverse probability censoring weighting or a doubly robust version, and then run standard local linear RD estimation on the transformed outcome. The paper proves that, with a bandwidth shrinking as $n^{-1/5}$, both sharp and fuzzy RD estimators are asymptotically normal with an explicit bias and variance, and that the doubly robust estimator is at least as efficient as inverse probability weighting when its nuisance models are correctly specified. If these results hold, medical researchers can estimate causal effects of threshold-based treatment rules, such as a PSA cutoff that triggers biopsy, without discarding or ignoring censoring.","feed_headline":"Regression discontinuity designs can handle censored survival data","feed_subtitle":"Transformed outcomes let local linear RD estimators recover causal effects from right-censored time-to-event data.","key_machinery":"The load-bearing object is the censoring unbiased transformation: a function of the observed data whose conditional expectation given the forcing variable equals the conditional expectation of the true log survival time. The paper uses $Y_{\\mathrm{IPCW}}=\\Delta Y/G(T)$ and $Y_{\\mathrm{DR}}=\\Delta Y/G(T)+\\int_0^{\\widetilde T}\\{Q_Y(u,W)/G(u)\\}\\,dM_G(u)$, where $Q_Y(u,W)$ is the conditional mean of log survival given survival past $u$ and $M_G$ is the censoring martingale. The transformation is what lets censored observations be handed to the ordinary local linear RD algorithm; the martingale term is what recovers information from censored cases and produces the claimed efficiency gain. Around this core sit the usual RD pieces: local linear fits on each side of the cutoff, a modified cross-validation bandwidth selector, and sandwich or nearest-neighbor variance estimators.","core_discovery":"The central claim is that the causal RD estimand is identifiable and estimable from censored survival data by local linear regression on a transformed outcome. For log survival time $Y$, the inverse-probability transform is $Y_{\\mathrm{IPCW}}=\\Delta Y/G(T)$, and the doubly robust transform adds a martingale integral that uses information from censored observations. Theorem 1 and its sharp-design corollary state that, under regularity conditions, $n^{2/5}(\\hat{\\tau}-\\tau-\\phi)$ converges in distribution to a centered normal law, where $\\phi$ is the asymptotic bias and the variance matrix is built from boundary variances and covariances of the transformed outcome and treatment assignment. Theorem 2 states that the doubly robust estimator using the true censoring and failure-time models has asymptotic variance no larger than that of the IPCW estimator or of a doubly robust estimator built from an incorrect failure-time model. The simulations and the PLCO prostate-cancer analysis are presented as illustrations of the method in use.","pith_inferences":["A practical rule the paper leaves implicit: use the doubly robust transform when a trustworthy failure-time model is available, but prefer the simpler IPCW transform when neither nuisance model can be relied on, because IPCW consistency requires only the censoring model.","Because the main theorem fixes $h\\sim n^{-1/5}$ while the implementation uses a data-driven bandwidth, a natural next test is whether the bandwidth selector preserves the claimed coverage; the paper does not settle this.","The same censoring-unbiased transformation trick could be applied to other boundary estimands, such as distributional, quantile, or multi-cutoff RD, wherever a transformation with the right conditional expectation exists."],"forward_implications":["Both sharp and fuzzy RD estimators can be computed by applying existing local linear RD software to the transformed outcomes, so the method does not require a new estimation algorithm.","In the fuzzy design, the ratio-of-jumps estimator identifies the complier average treatment effect on log survival time, giving an instrumental-variable interpretation under censoring.","With correctly specified nuisance models, the doubly robust transform dominates inverse probability weighting in asymptotic variance and shows smaller bias in the paper's simulations.","In the PLCO example, the method finds no statistically significant effect of the PSA $\\geq 4$ ng/ml threshold on mortality or first cancer incidence."],"supporting_citations":[{"why":"supplies the local linear RD asymptotic lemmas that the proof adapts to the censored-transform setting.","marker":"Hahn, Todd and Van der Klaauw (1999)"},{"why":"establishes nonparametric identification and estimation of sharp and fuzzy RD treatment effects.","marker":"Hahn, Todd and Van der Klaauw (2001)"},{"why":"introduces local linear regression for censored responses and the censoring-unbiased transformation idea.","marker":"Fan and Gijbels (1994)"},{"why":"introduces the doubly robust censoring unbiased transformation used to build the DR estimator.","marker":"Rubin and Van der Laan (2007)"},{"why":"extends the DR transformation to general transformations of the failure time and supplies the martingale form and efficiency comparison.","marker":"Steingrimsson, Diao and Strawderman (2019)"},{"why":"provides the MSE cross-validation bandwidth criterion modified in the paper's bandwidth selection.","marker":"Ludwig and Miller (2007)"},{"why":"provides the bias-corrected inference and nearest-neighbor variance estimators adapted for confidence intervals.","marker":"Calonico, Cattaneo and Titiunik (2014)"},{"why":"defines the sharp and fuzzy RD framework and the instrumental-variable interpretation used for the fuzzy estimand.","marker":"Imbens and Lemieux (2008)"}],"fun_headline_variants":["Censored data no barrier for regression discontinuity","Regression discontinuity works even with censored data","Causal effects via RD with censored data","Transformed outcomes enable RD with censored data","Regression discontinuity handles censored survival data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The asymptotic theory treats the data-driven bandwidth from the paper's cross-validation selector as if it were the fixed bandwidth of order $n^{-1/5}$; the paper does not prove that the estimated bandwidth is asymptotically equivalent to that fixed sequence.","fun_headline_variants_meta":{"raw":{"variants":["Censored data no barrier for regression discontinuity","Regression discontinuity works even with censored data","Causal effects via RD with censored data","Transformed outcomes enable RD with censored data","Regression discontinuity handles censored survival data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001021,"raw_usage":{"total_tokens":4259,"prompt_tokens":848,"completion_tokens":3411,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":3343}},"tokens_in":464,"tokens_out":3411,"duration_ms":24528,"temperature":1.0,"reasoning_tokens":3343,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:07:26.714746+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the sharp RD design with right censoring, compute the proposed estimator under the paper's data-driven bandwidth, and compare the empirical coverage of nominal 95% intervals with coverage when the bandwidth is fixed at $h = c n^{-1/5}$. If the data-driven version departs materially from 0.95 while the fixed-bandwidth version does not, the point of failure is the assumption that the estimated bandwidth can be ignored in the asymptotics.","supporting_citations":[],"review_version":1}