Pith. sign in

REVIEW 3 major objections 4 minor 15 references

Efficient Inference for Time-to-Event Outcomes by Integrating Right-Censored and Current Status Data

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Fusing right-censored and current-status samples yields a root-n efficient estimator of survival probabilities.

desk verdict Real contribution with a genuine proof gap: the efficiency theorem for the one-step estimator relies on an unproven assertion about how the nuisance-dependent η* affects the empirical process term. read the letter →

arxiv 2508.10357 v2 pith:VHKIOMC2 submitted 2025-08-14 stat.ME

classification stat.ME MSC 62N0162N02
keywords datafusionright-censoredcurrentstatussurvivalprobabilitysemiparametricefficiencycanonicalgradientone-stepestimationdoublerobustness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks how to estimate a survival probability when two datasets observe the same event process in different ways: one follows patients until a right-censoring time, and the other records only whether the event has occurred by a single inspection time. It claims that, although current-status-only estimators converge at the slow rate $n^{-1/3}$ with non-normal limits, fusing the two data types in a semiparametric model yields a root-n estimator that attains the semiparametric efficiency bound. The construction works by deriving the canonical gradient of the survival functional, which requires solving a Fredholm integral equation whose solution couples the two likelihood contributions. If correct, this means researchers no longer have to discard slow, coarse data when precise follow-up data are available: the coarse data can only help, and no regular estimator can do better asymptotically under the assumptions. A second, doubly-robust estimator stays consistent when either the event-time model or the censoring/inspection-time models are misspecified.

What carries the argument

The central object is the canonical gradient (efficient influence function) of the survival functional in the observed-data tangent space. It is characterized by a Fredholm integral equation of the second kind whose solution $\eta^*$ couples the two data sources: the right-censored score enters through a martingale weighted by the censoring survival $\Gamma$, and the current-status score enters through $\Theta^*(c,w)=\int_0^c \eta^*(u,w)\,d\Lambda_{T|W}(u|w)$. Solving this equation and plugging the solution into a one-step estimator is what converts a coarse, slow current-status sample into a root-n efficient correction to the right-censored estimate.

What would settle it

Simulate data satisfying Assumptions 1--4 with known distributions and with $t^*$ in the support of the inspection time. Compute the efficient one-step estimator on many large samples; if its empirical variance is not close to the variance of the canonical gradient $\tau^{\mathrm{eff}}_P$, or the Wald confidence intervals do not hold nominal coverage, the efficiency claim is false. In a second design, make the inspection time depend on the event time even after conditioning on $W$ (for example, inspect sicker patients later): the estimator should show growing bias, confirming that Assumption 2

Watch

Extended reading notes

Core claim

The paper's central discovery is that the survival probability at a fixed time is pathwise differentiable in the fused data model—even though it is not pathwise differentiable in current-status data alone—and its canonical gradient has an explicit integral-equation characterization. Treating the joint event-time distribution as exchangeable across sources, the authors show that a one-step estimator built from the canonical gradient is asymptotically normal at root-n with variance equal to the semiparametric efficiency bound. A second estimator based on a valid but non-canonical gradient is doubly robust. Simulations show 12--44% shorter confidence intervals than using right-censored data alo

Load-bearing premise

The load-bearing premise is Assumption 2: in the current-status sample, the inspection time is independent of the event time once covariates are held fixed, so an individual's unobserved disease duration does not influence when they happen to be checked; if inspection timing remains informative about the event after conditioning on recorded covariates, the current-status contribution is biased and the efficiency claim collapses.

Editorial extensions

If this is right

  • Any regular estimator of the survival probability in this fused model has asymptotic variance at least that of the canonical gradient; the proposed one-step estimator attains it, so under the assumptions no fusion scheme can be asymptotically more precise.
  • Adding current-status data to a right-censored sample shortens confidence intervals (12--44% in the paper's simulations) while preserving the root-n rate, so coarse surveys are genuinely informative even when their own estimators converge slowly.
  • Adding right-censored data to a current-status sample lifts the convergence rate from $n^{-1/3}$ to $n^{-1/2}$, changing the limiting distribution from non-normal to Gaussian and enabling Wald inference.
  • The doubly robust estimator remains consistent when either the event-time model or the censoring/inspection-time models are correctly specified, so flexible nuisance estimation can be used without sacrificing the fusion gain.
  • Under covariate shift (conditional exchangeability rather than full exchangeability), the same integral-equation approach gives efficient estimators of target-population survival probabilities, with density ratios entering the equation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: the same projection machinery should deliver efficient influence functions for other pathwise-differentiable functionals of the fused distribution, such as restricted mean survival time or survival differences between treatment groups, whenever the analogue of the Fredholm equation has a square-integrable solution; the paper identifies these as open directions.
  • Inference: the efficiency gain should be largest when inspection times cluster near the target time $t^*$, because the current-status observation $\mathbb{1}\{T\le C\}$ is then most strongly correlated with $\mathbb{1}\{T>t^*\}$; this suggests designing current-status surveys with inspection times tuned to the survival horizon of interest.
  • Inference: if Assumption 2 is doubted, the practical remedy is to add a sensitivity analysis that reweights or down-weights the current-status contribution; the paper does not provide this, but the integral-equation structure makes the influence of the current-status term explicit enough to construct one.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper develops a semiparametric framework for estimating the survival probability at a fixed time by fusing right-censored and current status data. Under exchangeability of the event-time/covariate distribution across sources and conditional independence of censoring/inspection times, the authors derive a canonical gradient for the survival functional. This gradient is characterized by the solution of a Fredholm integral equation, and the paper proposes a doubly robust one-step estimator and an efficient one-step estimator. The main theoretical claim is that the efficient estimator is asymptotically linear with influence function equal to the canonical gradient and therefore attains the semiparametric efficiency bound under Donsker conditions and o_p(n^{-1/4}) nuisance convergence rates. Simulation studies illustrate efficiency gains from fusing the two data types, and an appendix extends the approach to covariate shift.

Significance. If the main theorem is correct, this is a valuable contribution to data fusion for mixed censoring types. The paper identifies a setting where current status data, whose nonparametric estimators typically converge at n^{-1/3} and are non-normal, can still improve root-n inference when combined with right-censored data. The derivation of the canonical gradient via a hazard perturbation and coarsening-at-random is nontrivial, and the explicit Fredholm-equation representations are useful for implementation. The double robustness result is coherent and the simulation evidence for efficiency gains is suggestive. The manuscript is generally careful, with detailed appendix proofs, but the proof of the central efficiency theorem has a load-bearing gap.

major comments (3)
  1. [Section 4 / Appendix C.2, Theorem 4.3] The proof of Theorem 4.3 does not establish the asserted asymptotic linearity. In the one-step expansion, the empirical-process term (P_n-P)(τ_bP−τ_P) contains s[δ_R(bη*−η*)−∫_0^Y(bη*bλ−η*λ)dt] plus the current-status score, so bounding Rem(bP,P) alone is not enough. The text asserts that 'bη* does not introduce additional estimation errors', but no lemma controls ‖bη*−η*‖_{L2(P)} or shows that the implicit map bP↦bη* preserves Donsker properties. This same gap affects Theorem 4.1 condition (b), which asserts convergence of bE[bh*] from consistency of bF. Without a stability lemma for the solutions of the integral equations (2) and (4), Theorem 4.3 is not proven.
  2. [Section 4 / Appendix A] The theorems assume that bη* (and bh*) solve the empirical integral equation exactly, but the implementation uses a discretization or basis-expansion approximation. The numerical error from solving these Fredholm equations is not analyzed. A statement that this error is asymptotically negligible under the stated conditions is needed; otherwise there is a mismatch between the estimator analyzed and the estimator implemented.
  3. [Section 5, Table 1] The efficient estimator's coverage is below nominal in several scenarios, for example 0.904 at N=600 and t*=0.9, 0.921 at N=300 and t*=0.9, 0.929 at N=300 and t*=0.7, and 0.932 at N=1500 and t*=0.7. The paper does not discuss these violations. Since the paper's practical conclusion relies on valid Wald-type confidence intervals, the authors should explain possible finite-sample bias, variance underestimation, or nuisance estimation effects, and provide diagnostics.
minor comments (4)
  1. [Section 5, Figure 1] The caption says 'RS Only' but the text and estimator list use 'RC'. Please correct.
  2. [Section 5, paragraph before Table 1] The text references the 'survMLR package' but the reference list gives 'survML'. Also 'Kaplain-Meier' should be 'Kaplan-Meier'.
  3. [Section 2, Eq. (2)] The display of the kernel K(t,s|w) immediately after Eq. (2) is garbled; the factor involving π should be written unambiguously, e.g., (1−π)/π as in Appendix A.
  4. [Section 4, Theorem 4.2/4.3 discussion] The remark that the Donsker condition 'can be removed by cross-fitting' is not accompanied by a theorem or proof. This is fine as a claim, but it should be labeled as a conjecture or supported by a formal result.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation: gradients and efficiency bound are self-contained; only a minor non-load-bearing self-citation and an omitted stability proof for bη* (a correctness gap, not circularity).

full rationale

The derivation chain is not circular. The canonical gradients in Lemmas 2.2 and 3.2 are obtained by projecting valid gradients onto explicitly characterized tangent spaces; h* and η* are defined as unique solutions of Fredholm integral equations (2) and (4) that depend on the nuisance distributions, not on the target ψ except through the standard centering terms μ(w)−ψ. No fitted parameter is renamed as a prediction, and no target-dependent normalization is hidden in the construction. The one-step estimators are standard: asymptotic linearity is claimed under explicit nuisance convergence rates and Donsker/cross-fitting conditions, and the proof gives a detailed second-order remainder decomposition. The only self-citation is to Li and Luedtke (2023), one of whose authors is Xiudi Li, for the general data-fusion framework and for the overlap condition in Appendix B; this is contextual and not load-bearing—the gradient derivations here are self-contained. One flagged issue is non-circular but real: in the proof of Theorem 4.3, the paper asserts 'Importantly, estimating bη* does not introduce additional asymptotic variance' and later 'Since bη* does not introduce additional estimation errors' without proving a stability lemma that bounds ||bη*−η*|| in terms of the nuisance error rates. This is an omitted proof / potential gap in the efficiency argument, but it is not a circular reduction: the theorem does not define bη* in terms of the estimand or fit the influence function to the outcome. It should be weighed as a correctness risk, not as circularity. Overall circularity burden is low.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

The central theoretical claim is parameter-free: no numbers are fitted to data or chosen to force the efficiency result. Implementation choices such as polynomial basis degree in the simulations are not part of the asymptotic claim. The main assumptions are the four statistical conditions in Section 2-3 plus standard semiparametric theory and rate conditions for nuisance estimation.

assumptions (7)
  • domain assumption Assumption 1: (T,W) ⊥ S (exchangeability)
    Requires identical joint distribution of event time and covariates across data sources; used in Section 2 to derive the tangent space and gradients.
  • domain assumption Assumption 2: T ⊥ C | W, S=0 (conditional independent inspection times)
    Identifies the survival probability from current status data and underpins the CS component of the gradient. The paper notes dependence is a known failure mode.
  • domain assumption Assumption 3: inspection window positivity, F_T|W(c|w) ∈ [ζ, 1-ζ]
    Technical condition ensuring denominators are bounded in the gradient and integral equations.
  • domain assumption Assumption 4: T ⊥ R | W, S=1 (conditional independent censoring)
    Standard coarsening-at-random condition for the right-censored component.
  • domain assumption Nuisance estimators achieve n^{-1/4} rates (Theorem 4.3) and belong to a Donsker class or use cross-fitting
    Needed for asymptotic linearity of the efficient estimator; the paper asserts data-adaptive methods can satisfy this but does not prove it for the specific learners used.
  • standard math Existence and uniqueness of solutions to Fredholm equations (2) and (4)
    Assumed via square-integrability of the kernel and pathwise differentiability; needed to define h* and η*.
  • standard math Semiparametric efficiency theory (Bickel et al., Tsiatis, van der Laan & Robins)
    Background for tangent spaces, gradients, and one-step estimation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Efficient Inference for Time-to-Event Outcomes by Integrating Right-Censored and Current Status Data." pith.science (2026). https://pith.science/paper/VHKIOMC2

@misc{pith2026250810357,
  author       = {Pith},
  title        = {Pith review of: Efficient Inference for Time-to-Event Outcomes by Integrating Right-Censored and Current Status Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VHKIOMC2}},
  note         = {Machine review of arXiv:2508.10357}
}
read the original abstract

We propose a semiparametric data fusion framework for efficient inference on survival probabilities by integrating right-censored and current status data. Existing data fusion methods focus largely on fusing right-censored data only, while standard meta-analysis approaches are inadequate for combining right-censored and current status data, as estimators based on current status data alone typically converge at slower rates and have non-normal limiting distributions. In this work, we consider a semiparametric model under exchangeable event time distribution across data sources. We derive the canonical gradient of the survival probability at a given time, and develop one-step estimators along with the corresponding inference procedure. Specifically, we propose a doubly robust estimator and an efficient estimator that attains the semiparametric efficiency bound under mild conditions. Importantly, we show that incorporating current status data can lead to meaningful efficiency gains despite the slower convergence rate of current status-only estimators. We demonstrate the performance of our proposed method in simulations and discuss extensions to settings with covariate shift. We believe that this work has the potential to open new directions in data fusion methodology, particularly for settings involving mixed censoring types.

Figures

Figures reproduced from arXiv: 2508.10357 by the authors.

Figure 1
Figure 1. log(MSE) vs. log(N) when t ∗ = 0.7 for current status only estimator (CS), right censored only AIPCW estimator (RS), and the proposed data fusion estimator (Both). 16 [PITH_FULL_IMAGE:figures/full_fig_p016_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 9 canonical work pages

  1. [2]

    S Z (η∗(t, W)−η †(t, W))dMT (t) Z η(t, W)dMT (t) # +E

    Thus, we obtain the observed data tangent space. Moreover, the scores generated by perturbing different component distributions toPare orthogonal. See also Chapter 5 of Tsiatis [2006]. Proof of Lemma 3.2.Recall that with right-censored data and current status data, the augmented inverse probability of censoring weighted estimator using the right-censored ...

  2. [4]

    S Z Γ(u|W) bΓ(u|W) S(u|W) bET|W h bh(T, W)|T≥u, W i −E T|W h bh(T, W)|T≥u, W i · λR|W (u|W)− bλR|W (u|W) du # =E P

    Proof of Theorem 4.1.We prove Theorem 4.1 by showing ϕt∗ =E P bµ(W) + (1−S)(∆ C − bFT|W (C|W)) bFT|W (C|W)(1− bFT|W (C|W)) Z C 0 bh∗(u, W)dbFT|W (u|W) +S Z ∞ 0 bh∗(u, W)− bET|W [bh∗(T, W)|T≥u, W] bΓ(u|W) d I{Y≤u,∆ R = 1} − Z u 0 I{Y≥v}d bΛT|W (v|W) , (10) if either bFT|W =F T|W (and therefore bET|W [bh∗(T, W)|T≥u, W] =E T|W [bh∗(T, W)|T≥u, W]), or 41 {bg=...

  3. [12]

    S ∆R bΓ(T|W) − ∆R Γ(T|W) ! bh(T, W) # | {z } (term A-I) −E P

    with bγ(w) = (1−π) Z Z ∞ t bH(c, w)bg(c|w) bFT|W (c|w)(1− bFT|W (c|w)) bf(t|w)dcdt. 37 Putting these together, we get that term B = Z ( πbh(t, w) + (1−π) Z ∞ t bH(c, w)bg(c|w) bFT|W (c|w)(1− bFT|W (c|w)) dc ) f(t|w)p W (w)dtdw + (1−π) Z Z ∞ t bH(c, w){g(c|w)−bg(c|w)} bFT|W (c|w)(1− bFT|W (c|w)) dcf(t|w)p W (w)dtdw −(1−π) Z bFT|W (c|w) R c 0 bf(t|w) bh(t, ...

  4. [14]

    In addition, under bP,bg(u|w)−g(u|w) = 0 for anyu /∈ C. Hence the remainder term reduces to, Rem( bP , P) =π Z Z max(cu,t∗) 0 n Γ(t|w)ST|W (t|w)− bΓ(t|w) bST|W (t|w) o bη∗(t, w) n λT|W (t|w)− bλT|W (t|w) o pW (w)dtdw −(1−π) Z Z cu cl bST|W (c|w) bΘ∗(c, w) bFT|W (c|w) {bg(c|w)−g(c|w)} n ΛT|W (c|w)− bΛT|W (c|w) o pW (w)dcdw + (1−π) Z Z cu cl FT|W (c|w)− bFT...

  5. [15]

    (12) Given the derivations above, we obtain the canonical gradient of Φ i evaluated atP I relative to the modelM I 2 as follows: ξi,P I :x I 7→sh ∗(t, w)+ (1−s)(δ C −F T|W (c|w)) FT|W (c|w)(1−F T|W (c|w)) Z c 0 fT|W (u|w)h∗(u, w)du+1(s=i) Π(S=i) {µ(w)−µ i}, whereh ∗ solves the equation in (12). Finally, applying the results for coarsening-at-random, for e...

  6. [1955]

    B. R. Baer, A. Ertefaie, and R. L. Strawderman. On causal inference for the survivor function. arXiv preprint arXiv:2507.16691,

  7. [1979]

    Graham, M

    18 E. Graham, M. Carone, and A. Rotnitzky. Towards a unified theory for semiparametric data fusion with individual-level data.arXiv preprint arXiv:2409.09973,

  8. [2001]

    F. P. Havers, C. Reed, T. Lim, J. M. Montgomery, J. D. Klena, A. J. Hall, A. M. Fry, D. L. Cannon, C.-F. Chiang, A. Gibbons, et al. Seroprevalence of antibodies to sars-cov-2 in 10 sites in the united states, march 23-may 12, 2020.JAMA internal medicine, 180(12):1576–1586,

Show all 15 references
  1. [2003]

    A similar result is available in Tsiatis

    and results in van der Laan and Rubin [2007], a valid gradient relative to the observed data model is given by τ:x7→s Z ∞ 0 h∗(u, w)−ET|W [h∗(T, W)|T≥u, W=w] Γ(u|w) d I{y≤u, δ R = 1} − Z u 0 I{y≥v}dΛ T|W (v|w) + (1−s)(δ C −F T|W (c|w)) FT|W (c|w)(1−F T|W (c|w)) Z c 0 f(t|w)h ∗...

  2. [2007]

    W. Hu, R. Wang, W. Li, and W. Miao. Semiparametric efficient fusion of individual data and summary statistics.arXiv preprint arXiv:2210.00200,

  3. [2012]

    Y. Wang, A. Ying, and R. Xu. Doubly robust estimation under covariate-induced dependent left truncation.Biometrika, 111(3):789–808, 2024a. Y. Wang, A. Ying, and R. Xu. Learning treatment effects under covariate dependent left truncation and right censoring.arXiv preprint arXiv...

  4. [2013]

    C. Gao, S. Yang, M. Shan, W. W. Ye, I. Lipkovich, and D. Faries. Doubly protected estimation for survival outcomes utilizing external controls for randomized clinical trials.arXiv preprint arXiv:2410.18409,

  5. [2022]

    19 Y. Liu, A. W. Levis, K. Zhu, S. Yang, P. B. Gilbert, and L. Han. Targeted data fusion for causal survival analysis under distribution shift.arXiv preprint arXiv:2501.18798,

  6. [2023]

    D. Liu, R. Y. Liu, and M.-g. Xie. Nonparametric fusion learning for multiparameters: Synthesize inferences from diverse sources using data depth and confidence distribution.Journal of the American Statistical Association, 117(540):2086–2104,

  7. [2025]

    Xie and P

    M.-g. Xie and P. Wang. Repro samples method for finite-and large-sample inferences.arXiv preprint arXiv:2206.06421,

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.