REVIEW 2 major objections 4 minor 21 references
Survival analysis under label shift
T0 review · 2 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Under label shift, a target population's parametric survival model is identifiable and estimable from censored source data and target covariates alone, with $\sqrt{n_0}$-consistency and explicit variance.
desk verdict First real treatment of label shift under censoring, with serious asymptotics, but the identifiability lemma has a boundary gap that overclaims Theorem 3.1 as stated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is the approximate log-likelihood $\ell_n(\theta)$ in equation (3.3), in which the unknown source marginal survival function $p_T$ is replaced by its Kaplan-Meier estimator and the unknown target covariate distribution $q_Z$ by the empirical distribution of target covariates; maximizing it yields $\hat\theta$. Identifiability is guaranteed by Lemma 3.1's ratio condition, and the asymptotic proof works by decomposing the score into $\psi_1$, $\psi_{2n}$, and $\psi_{3n}$, where $\psi_{2n}$ and $\psi_{3n}$ carry the estimation error of $q_Z$ and $p_T$. Lemma C.1 supplies a uniform representation of Kaplan-Meier integrals, adapted from the Stute representation and from the product-limit integral theory, that lets those terms be linearized into U-statistic form with explicit influence components $\psi_{p_T}$ and $\psi_{q_Z}$.
What would settle it
Generate source data with $T$ drawn from a known Weibull model, let the censoring time $C$ have a hazard that depends on a covariate $Z$ (for example $C\sim\operatorname{Exp}(\lambda e^{\gamma Z})$ with $\gamma\neq 0$), and keep target data with only $Z$; run the proposed estimator against this data and compare $\hat\theta$ to the truth. If the estimator remains $\sqrt{n_0}$-consistent with nominal coverage, the independence assumption is not load-bearing; systematic bias or under-coverage that grows with sample size shows the method fails when censoring depends on covariates.
Extended reading notes
Core claim
Under label shift ($p_T(t)\neq q_T(t)$ but $p_{Z|T}(z,t)=q_{Z|T}(z,t)$) with right-censored event times in the source population and only covariates observed in the target, the paper claims the parametric conditional density $q_{T|Z}(t,z;\theta)$ of the target is identifiable whenever the ratio $q_{T|Z}(t,z;\theta)/q_{T|Z}(t,z;\tilde\theta)$ depends on $z$ for $\theta\neq\tilde\theta$ and $\operatorname{var}_q(Z)\neq 0$, and that the approximate maximum likelihood estimator obtained by replacing $p_T$ with the Kaplan-Meier estimator and $q_Z$ with the empirical distribution is consistent and $\sqrt{n_0}$-asymptotically normal with $n_0=\min(n_1,n_2)$, as stated in Theorems 3.1 and 3.2. The asymptotic variance is given explicitly as a sandwich formula whose middle factor combines the score variance with corrections for estimating $p_T$ and $q_Z$, so standard errors and confidence intervals can be formed without bootstrap. The claims cover parametric proportional hazards, accelerated failure time, proportional odds, and accelerated hazards models, and a corollary transfers the result to plug-in estimators of functionals such as conditional mean survival time.
Load-bearing premise
The load-bearing premise is that censoring in the source population is independent of both the event time and the covariates; if censoring depends on either one, the Kaplan-Meier estimate of the event-time distribution is inconsistent and the target inference collapses.
Editorial extensions
If this is right
- Target-population covariate effects and functionals like conditional mean survival can be estimated and tested using only source censored events plus target covariates, at rate $\sqrt{\min(n_1,n_2)}$.
- The convergence rate is set by the smaller of the two sample sizes: enlarging one population cannot overcome a small other population, which follows directly from Theorem 3.2.
- The same approximate-likelihood recipe covers parametric proportional hazards, accelerated failure time, proportional odds, and accelerated hazards models, so one implementation transfers across classical survival model families.
- The explicit asymptotic variance formula allows confidence intervals without bootstrap, with simulation coverage close to nominal level at moderate sample sizes.
- Applied to the liver transplant data, inference that ignores label shift can declare a covariate significant or non-significant opposite to the oracle conclusion, while the shift-aware estimator matches the oracle.
Reading between the lines
- An extension the paper leaves open: under conditional censoring $C\perp T\mid Z$, a modified estimator that weights by an estimated censoring model would likely restore consistency, at the cost of an extra variance term.
- A practical identifiability pre-check suggested by Lemma 3.1: for the fitted model, inspect whether the ratio of conditional densities at two covariate values varies with $t$; if it is constant in $z$, the parameters cannot be separated from target-only covariate data.
- The support-inclusion assumption implies the method is safest when the target's event times are stochastically no larger than the source's; for a target with longer survival, the estimated conditional density extrapolates beyond the source support.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies domain adaptation in survival analysis under label shift: the source population P provides right-censored survival times and covariates, while the target population Q provides only covariates, and the conditional distribution of covariates given survival time is assumed identical across populations. The authors model the target conditional density q_{T|Z}(t,z;θ) parametrically and estimate θ by maximizing an approximate likelihood in which the source marginal survival function is replaced by the Kaplan-Meier estimator and the target covariate distribution by the empirical distribution. They state a global identifiability lemma, prove consistency and asymptotic normality with an explicit variance formula (Theorems 3.1 and 3.2, Corollary 3.1), and report simulations over four parametric survival models plus an application to a liver transplant dataset, including a bootstrap test of the label-shift assumption.
Significance. If the identifiability issue described below is repaired, the paper would be a useful contribution: it is the first systematic treatment of label shift with censored survival responses, the approximate-likelihood construction is natural, the explicit asymptotic variance estimator avoids bootstrap computations, and the simulation study covers several classical survival models. The paper also supplies detailed regularity conditions and proof sketches in the supplement, and the real-data analysis includes a check of the label-shift assumption. The main weakness is that the identifiability argument underlying Theorem 3.1 is not valid over the full parameter space of the paper's own examples, so the consistency and normality claims are currently overstated; the restrictive independent-censoring and compact-support assumptions are stated and partially acknowledged by the authors.
major comments (2)
- [Section 3, Lemma 3.1 and Example 3.2; Supplement S1] The identifiability proof is incomplete for the paper's leading models, and the gap is load-bearing for Theorem 3.1. Lemma 3.1 requires that q_{T|Z}(t,z;θ)/q_{T|Z}(t,z;θ~) depend on z whenever θ≠θ~. For the Weibull proportional hazards model of Example 3.2, take θ=(β,γ,λ) and θ~=(β~,γ~,λ~) with β=β~=0. Then the ratio equals h(t;γ,λ) exp{-H(t;γ,λ)} / [h(t;γ~,λ~) exp{-H(t;γ~,λ~)}], which is independent of z for any γ,λ,γ~,λ~. The proof step 'setting t=0 leads to β=β~' only yields β=β~ and does not identify the baseline parameters. The same failure occurs in the accelerated failure time, proportional odds, and accelerated hazards examples whenever the covariate coefficients vanish. Because Conditions A1–A6 contain no identifiability assumption, the proof of Theorem 3.1 in Supplement S3.2.5, which requires a unique maximizer of Eℓ(θ), is not justified as stated. This is not merely a proof gap: if the target model has β=0 and Z⊥T in P, then any baseline parameters (γ,λ) give the same qZ(z) and the same approximate likelihood (3.3), so θ is genuinely unidentified. The authors should restate Lemma 3.1 and Theorems 3.1–3.2 with an explicit identifiability condition, correct the example verifications, and discuss the null covariate-effect case.
- [Supplement S3.2.2, Conditions A1–A6] Even if an identifiability condition is added, the current statement that 'the condition required in Lemma 3.1 is relatively mild' and that the listed models satisfy it over the whole parameter space is not accurate. The condition fails on a nonempty open set of parameter values when the covariate effect is zero, and the simulations and data application always use nonzero coefficients, so the numerical evidence does not exercise the problematic case. Please also state whether the proposed identifiability condition is intended to hold on the whole parameter space or only in a neighborhood of θ0, and adapt the consistency proof accordingly.
minor comments (4)
- [Section 5, Table 3] The row labeled 'P' is described as fitting a proportional hazards model with a Weibull baseline, whereas BIC in Table 2 selects an accelerated failure time model with an exponential baseline; please clarify whether this is deliberate or a typo, since the comparison with the oracle is not under the same selected model.
- [Supplement S2] The bootstrap label-shift test is used to justify the key assumption in the data application, but no regularity conditions or bandwidth choices are given for the kernel estimators of R_p and R_q; please add implementation details or a reference.
- [Equation (3.3)] In the third term of (3.3), the summation index k is used with r_k δ_k I(t_k > x_i), but t_k is not defined among the observed (X_i,Δ_i,Z_i); if t_k denotes the uncensored event times in the source sample, please state this explicitly.
- [Lemma 3.1] The condition var q(Z)≠0 is not used in the proof of Lemma 3.1; if it is needed, explain how, otherwise remove it.
Circularity Check
No circularity: the estimator plugs in nonparametrically estimated nuisance functions, and all cited technical tools are prior external results with stated assumptions.
full rationale
The paper's derivation is self-contained in the relevant sense. The target quantity θ is estimated by maximizing the approximate log-likelihood (3.3), which uses the Kaplan-Meier estimator bP_T for p_T and the empirical bQ_Z for q_Z; these are nuisance plug-ins, not fitted parameters that reappear as predictions. The score decomposition in Lemmas 3.2 and 3.3 corrects for estimation of these nuisance functions, and the asymptotic variance formula in Theorem 3.2 is an explicit function of the true densities, not of fitted target responses. Target responses are never used in estimation. The cited technical results (Stute 1995; Sánchez Sellero et al. 2005; Van der Vaart and Wellner 1996) are prior mathematical theorems with stated assumptions; even though two coauthors have prior work in the citation chain, those results are not equivalent to the present conclusion and do not smuggle in the target claim. The paper's own identifiability Lemma 3.1 is a sufficient condition rather than a restatement of the estimand. A separate correctness caveat, not a circularity, is that the proof of Example 3.2 in Supplement S1 sets t=0 to force β=β~, which fails when both β and β~ are zero; the same boundary gap affects the claims for the other example models. This would be relevant to theorem validity, but it is not a case of a prediction reducing to its input by construction.
Assumptions & free parameters
assumptions (6)
- domain assumption Label shift: p_T(t) != q_T(t) but p_{Z|T}(z,t) = q_{Z|T}(z,t)
- domain assumption Censoring time C is independent of (T,Z) in P
- domain assumption Parametric model q_{T|Z}(t,z;theta) is correctly specified in Q
- domain assumption Identifiability condition: q_{T|Z}(t,z;theta)/q_{T|Z}(t,z;theta_tilde) depends on z whenever theta != theta_tilde, and var_q(Z) != 0
- domain assumption Support inclusion and compact supports: support(q_T) subset support(p_T), supports compact, tau_T <= tau_C
- standard math Empirical process and Kaplan-Meier integral theory (Stute 1995; Sanchez Sellero et al. 2005; Van der Vaart and Wellner 1996)
Cite this review
Pith. "Pith review of Survival analysis under label shift." pith.science (2026). https://pith.science/paper/URZ4EHW4
@misc{pith2026250621190,
author = {Pith},
title = {Pith review of: Survival analysis under label shift},
year = {2026},
howpublished = {\url{https://pith.science/paper/URZ4EHW4}},
note = {Machine review of arXiv:2506.21190}
}
abstract
Let P represent the source population with complete data, containing covariate $\mathbf{Z}$ and response $T$, and Q the target population, where only the covariate $\mathbf{Z}$ is available. We consider a setting with both label shift and label censoring. Label shift assumes that the marginal distribution of $T$ differs between $P$ and $Q$, while the conditional distribution of $\mathbf{Z}$ given $T$ remains the same. Label censoring refers to the case where the response $T$ in $P$ is subject to random censoring. Our goal is to leverage information from the label-shifted and label-censored source population $P$ to conduct statistical inference in the target population $Q$. We propose a parametric model for $T$ given $\mathbf{Z}$ in $Q$ and estimate the model parameters by maximizing an approximate likelihood. This allows for statistical inference in $Q$ and accommodates a range of classical survival models. Under the label shift assumption, the likelihood depends not only on the unknown parameters but also on the unknown distribution of $T$ in $P$ and $\mathbf{Z}$ in $Q$, which we estimate nonparametrically. The asymptotic properties of the estimator are rigorously established and the effectiveness of the method is demonstrated through simulations and a real data application. This work is the first to combine survival analysis with label shift, offering a new research direction in this emerging topic.
Figures
Reference graph
Works this paper leans on
-
[1]
Alexandari, A., Kundaje, A., and Shrikumar, A. (2020). Maximum likelihood with bias-corrected calibration is hard-to-beat at label shift adaptation. InInternational Conference on Machine Learning, pages 222–232. PMLR. Azizzadenesheli, K., Liu, A., Yang, F., and Anandkumar, A. (2019). Regularized learning for domain adaptation under label shifts. arXiv pre...
arXiv 2020
-
[2]
, for some constant C. Note further that F12 and F14 have the second property of (1) and (3) of Lemma S3.4, hence by Van der Vaart and Wellner (1996)[Theorem 2.10.20], the class FS0 also has 28 this property, i.e., it satisfies the requirement on covering number in Condition (b2) of Lemma C.1. Combining the above results, FS0 satisfies Condition (b2) of L...
work page 1996
-
[3]
Cambridge university press. Van der Vaart, A. W. and Wellner, J. A. (1996). Weak Convergence and Empirical Processes: With Appli- cations to Statistics . Springer Series in Statistics. Springer. Van der Vaart, A. W. and Wellner, J. A. (2000). Preservation theorems for Glivenko-Cantelli and uniform Glivenko-Cantelli classes. In High dimensional probability...
work page 1996
-
[4]
Example S1.2 (Proportional odds model)
This implies that for every t and z, ∂ ∂z log qT |Z(t, z; θ) qT |Z(t, z; eθ) = (log(t) − µ − zTβ)β σ2 − (log(t) − eµ − zT eβ)eβ eσ2 = 0, which leads to θ = eθ contradicting the assumption θ ̸= eθ. Example S1.2 (Proportional odds model) . For the proportional odds model with LogLogistic baseline sur- vival function with θ = (βT, µ, σ)T, the conditional den...
work page 1993
-
[5]
B4 The functions qT |Z(t, z; θ0), ∂ ∂θ qT |Z(t, z; θ0), qT |Z(t, z; θ0)/qT (t) and ∂ ∂θ qT |Z(t, z; θ0)/qT (t) are bounded on ST × (SZp ∪ SZq). B5 The map θ 7→ E{ℓ(X, Z, ∆, R; θ)} is twice differentiable at θ0 with a non-singular second derivative matrix. Remark S3.3. Condition B3 guarantees that the classes defined in the proof of Lemma S3.4 have envelop...
work page 2005
-
[6]
Then, class F1 is Glivenko-Cantelli (Van der Vaart and Wellner, 1996, Theorem 2.4.1)
Since 0 < ∥c1∥L1(P ) < +∞ by Condition A1, N[] (ε, F1, L1(P )) < +∞ for every fixed ε >0. Then, class F1 is Glivenko-Cantelli (Van der Vaart and Wellner, 1996, Theorem 2.4.1). Result (4): We next prove the class F4 is Glivenko-Cantelli. Note that log( ·) is a continuous mapping. For every x ∈ ST , θ ∈ Θ, δ|logqT |Z(x, z; θ)| ≤δ|log{m1(x, z)}| + δ|log{M1(x...
work page 2000
-
[7]
Hence, F2 is Q-Glivenko-Cantelli
Since 0 < ∥c2∥L1(Q) < +∞ by Condition A2, N[] (ε, F1, L1(Q)) < +∞ for every fixed ε >0. Hence, F2 is Q-Glivenko-Cantelli. Result (3): As F3 is indexed by θ ∈ Θ, applying the same argument as in Result (1) we have N[](2ε∥c3∥L1(P ), F3, L1(P )) ≤ N(ε, Θ, ∥ · ∥1) ≤ K1ε−d. Since 0 < ∥c3∥L1(P ) < +∞ by Condition A3, N[] (ε, F3, L1(P )) < +∞ for every fixed ε >...
work page 2000
-
[8]
for i = 1, . . . , N. Now, sup θ,x,z Z t>x qT |Z(t, z; θ) bqT (t; θ) d bPT (t) − Z t>x qT |Z(t, z; θ) qT (t; θ) dPT (t) ≤ sup θ,x,z Z t>x qT |Z(t, z; θ) bqT (t; θ) − qT |Z(t, z; θ) qT (t; θ) d bPT (t) + Z t>x qT |Z(t, z; θ) qT (t; θ) d{ bPT (t) − PT (t)} . (S2) For the first term, we have sup θ,x,z Z t>x qT |Z(t, z; θ) bqT (t; θ) − qT |Z(t, z; θ) qT (t; θ...
work page 1993
Show all 21 references
-
[9]
Lemma S3.3
Therefore, we have sup θ∈Θ |ℓn(θ) − E{ℓ(X, Z, ∆; θ)}| p − →0 For any function g(·, h), we write its kth Gateaux derivative with respect to h at h1 in the direction h2 − h1 as ∂kg (·, h1) ∂hk [h2 − h1] ≡ ∂kg{·, h1 + ε(h2 − h1)} ∂εk ε=0 . Lemma S3.3. Under Conditions A2 and B2, ...
1996
-
[10]
Hence, F10 is Q-Donsker
Since 0 < ∥c2∥L2(Q) < +∞ by Condition A2, R ∞ 0 q logN[] (ε, F10, L2(Q))dε < +∞. Hence, F10 is Q-Donsker. By the definition of the Donsker class from Van der Vaart and Wellner (1996), the empirical process √n2(Qn − Q) weakly converges to a tight Borel measurable element in ℓ∞(...
1996
-
[11]
Then, F11 is also Q-Donsker
Since 0 < ∥c7∥L2(Q) < +∞ by Condition B2, R ∞ 0 q logN[] (ε, F11, L2(Q))dε <+∞. Then, F11 is also Q-Donsker. The empirical process √n2(Qn − Q) weakly converges to a tight Borel measurable element in ℓ∞(F11). Hence, we also have ∥bq∗ T (t) − q∗ T (t)∥∞ = Op(n−1/2 2 ) = op(n−1/4...
1996
-
[12]
Since 0 < ∥c6/qT ∥L1(P ) < +∞ by Condition B2, N[] (ε, F13, L1(P )) < ∞ for every ε >0. Note that for any norm ∥ · ∥, N (ε∥c6/qT ∥, F13, ∥ · ∥) ≤ N[](2ε∥c6/qT ∥, F13, ∥ · ∥) and for any probability measure eP , 0 < ∥M6∥L2( eP ) < +∞ is equivalent to 0 < ∥c6/qT ∥L2( eP ) < +∞. ...
1996
-
[13]
Let MF16(t) ≡ max{1, diam SZq}c5(t)/qT (t) + qT |Z(t, z0; θ0)/qT (t) for some fixed z0 ∈ SZq
Since 0 < ∥c5/qT ∥L1(P ) < +∞, N[] (ε, F16, L1(P )) < ∞ for every ε >0. Let MF16(t) ≡ max{1, diam SZq}c5(t)/qT (t) + qT |Z(t, z0; θ0)/qT (t) for some fixed z0 ∈ SZq. Then MF16(t) is an envelope function of F16. Since 0 < ∥c5/qT ∥L1(P ) < +∞ by Condition B1, 0 < ∥MF16∥L1(P ) < ...
1996
-
[14]
Since 0 < ∥MF16∥L1(P ) < +∞, we can obtain N[](ε, F14, ∥ · ∥L1(P )) < ∞ for every ε
Hence, N[]((K5 + 2)ε ∥MF16∥L1(P ) , F14, ∥ · ∥L1(P )) ≤ KKq 6 ε−(dz+1)Kq < ∞ for every ε. Since 0 < ∥MF16∥L1(P ) < +∞, we can obtain N[](ε, F14, ∥ · ∥L1(P )) < ∞ for every ε. Next we prove the second claim. RecallF14 = {t 7→ R qT |Z(t, z; θ0)d eQZ(z)/qT (t) : any empirical dis...
1996
-
[15]
Since F14 ⊂ convF16, we can obtain that F14 has a measurable envelope function M7(t) ≡ MF16(t) such that Z ∞ 0 sup eP r logN ε∥M7∥L2( eP ), F14, L2( eP ) dε <∞, where the sup is taken over all the probability measure eP on ST such that ∥M7∥L2( eP ) < ∞. Result (4): The proof i...
2007
-
[16]
Then, var(ηφS0 ) goes to 0 as n2 → ∞
Let φS0 ≡ I(t>x) q2 T (t) qT |Z(t, z; θ0){bqT (t) − qT (t)}. Then, var(ηφS0 ) goes to 0 as n2 → ∞. Hence, by result (3) of Lemma C.1, we have Z I(t > x) q2 T (t) qT |Z(t, z; θ0){bqT (t) − qT (t)}d{ bPT (t) − PT (t)} = op(n−1/2 1 ), uniformly in ( x, z). By the definition of (S...
2007
-
[17]
, for some constant C6. By Result (1) and Result (3) of Lemma S3.4, the classes F12 and F14 satisfies the uniform entropy condition defined in Van der Vaart and Wellner (1996)[Section 2.5.1] Then by Van der Vaart and Wellner (1996)[Theorem 2.10.20], the class FS21 also satisfi...
1996
-
[18]
Then, var( ηφS21 ) goes to 0 as n2 → ∞
Let φS21 ≡ I(t>x)q∗ T (t) q3 T (t) qT |Z(t, z; θ0){bqT (t) − qT (t)}. Then, var( ηφS21 ) goes to 0 as n2 → ∞. 34 Hence, by Result (3) of Lemma C.1, we have Z I(t > x)q∗ T (t) q3 T (t) qT |Z(t, z; θ0){bqT (t) − qT (t)}d{ bPT (t) − PT (t)} = op(n−1/2 1 ), uniformly in ( x, z). T...
2007
-
[19]
, for some constant C9. Note further both F12 and F15 satisfy the second property of (1) of Lemma S3.4, hence by Van der Vaart and Wellner (1996)[Theorem 2.10.20], the class FS22 also has this property, i.e., it satisfies the requirement on covering number in Condition (b2) of...
1996
-
[20]
R∆ π 1 − Ri 1 − π ∂ ∂θ qT |Z(X, Zi; θ0)qT (X) − q∗ T (X)qT |Z(X, Zi; θ0) {qT (X)}2 Zi, Ri # + op(n−1/2) + op(n−1/2 2 ) = n−1 nX i=1 1 − Ri 1 − π E
Let φS22 ≡ I(t>x) q2 T (t) qT |Z(t, z; θ0){bq∗ T (t)− q∗ T (t)}. Then, var( ηφS22 ) goes to 0 as n2 → ∞. Then we have Z I(t > x) q2 T (t) qT |Z(t, z; θ0){bq∗ T (t) − q∗ T (t)}d{ bPT (t) − PT (t)} = op(n−1/2 1 ), uniformly in ( x, z). Thus, ∂S2n(x, z, qT , q∗ T ) ∂q ∗ T [bq∗ T ...
2000
-
[21]
Hence, if π = 0, Σψ = varp{ψ(X, Z, ∆; θ0) + ψpT (X, ∆)}, and if π = 1, Σψ = varq{ψqZ(Z)}
We summarize the resulting expressions under specific values of π as follows: if π <1/2, Σψ = varp{ψ(X, Z, ∆; θ0) + ψpT (X, ∆)} + π 1 − π varq{ψqZ(Z)}, if π = 1/2, Σψ = varp{ψ(X, Z, ∆; θ0) + ψpT (X, ∆)} + varq{ψqZ(Z)}, and if π >1/2, Σψ = 1 − π π [varp{ψ(X, Z, ∆; θ0) + ψpT (X,...
1986
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.