REVIEW 2 major objections 6 minor 5 references
Results on standard estimators in the Cox model
T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read For the proportional hazards regression model, the rescaled maximum partial likelihood and Breslow estimators have uniformly bounded moments of every order.
desk verdict First uniform moment bounds for Cox estimators fill a real gap; Theorem 2's proof has a fixable gap on a random β* term. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument turns on three objects. The score of the partial likelihood is written as a sum of stochastic integrals of predictable processes with respect to counting-process martingales, so its moments are controlled through its predictable variation process. The empirical process term appearing in the Breslow estimator is treated as a supremum over a class of functions $f_t(u,z) = 1_{u \geq t} e^{\beta_0' z}$, whose bracketing numbers grow like $1/\varepsilon$ and whose envelope is $F(u,z) = e^{\beta_0' z}$; exponential moment assumptions bound that envelope in $L_{2 \vee p}(P)$. Finally, the Taylor expansion of the score around $\beta_0$ is controlled on an event $E_n$ on which the normalized matrix of second derivatives is uniformly close to a nonsingular matrix, so that the inverse stays bounded and the estimation error is dominated by the normalized score.
What would settle it
Set up the two-sample balanced Cox model with bounded covariates, unit exponential baseline, and independent censoring chosen so that $P(T = \tau_G) > 0$, then compute analytically or by high-precision simulation the quantities $\sup_n E[n^{p/2}|\hat\beta_n - \beta_0|^p]$ for $p=2$ and $p=4$; if any of these is infinite or grows without bound in $n$, Theorem 1 is false.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that the classical asymptotic normality results for the two standard estimators of the Cox model can be upgraded to fully quantitative moment bounds. Theorem 1 shows that under assumptions (A1), (A2), and (A4), with a continuous baseline hazard, there is a sequence of events $E_n$ with $P(E_n) \to 1$ and a constant $K$ such that $\limsup_n E[1_{E_n} n^{p/2} |\hat\beta_n - \beta_0|^p] \leq K$ for every $p \geq 1$. Theorem 2 shows the analogous uniform bound for $\sup_{t \in [0,\tau_G)} |\Lambda_n(t) - \Lambda_0(t)|$ under (A1)-(A4). Lemma 3 supplies an unconditional bound for the empirical process $n^{1/2} \sup_t |\Phi_n(t; \beta_0) - \Phi(t; \beta_0)|$ under the exponential-moment assumption (A3). The event restriction is a device to isolate the random fluctuation of the matrix of second derivatives; because the events have probability tending to one, the bounds are sufficient for later convergence-in-distribution arguments.
Load-bearing premise
The load-bearing extra condition is that the covariate vector has finite exponential moments in the direction of the true regression parameter, $E[e^{q \beta_0' Z}] < \infty$ for every $q \geq 1$, because without it the envelope $e^{\beta_0' z}$ of the empirical process class is not controlled enough for the moment bounds; the positivity assumption $P(T = \tau_G) > 0$ is also indispensable so that $\Phi$ does not vanish at the end of the observation interval.
Editorial extensions
If this is right
- The normalized coefficient error $n^{1/2}(\hat\beta_n - \beta_0)$ is uniformly integrable: its $p$-th moments are bounded for every $p$, not just for $p=1$ or $p=2$.
- The sup-norm error of the Breslow estimator over $[0,\tau_G)$ has bounded moments, so $L_p$ convergence of cumulative-hazard estimation follows at the parametric $\sqrt{n}$ rate.
- These bounds are exactly the ingredient needed to pass from sup-norm or pointwise asymptotics to global $L_p$ errors in shape-restricted baseline hazard estimation, such as isotonic-type estimators.
- Because the bounds hold for every $p$ simultaneously, they also cover higher-moment terms in expansions, enabling uniform confidence bands via moment-based arguments.
- Lemma 3 stands alone: the rescaled empirical process for the at-risk function $\Phi_n$ has bounded moments without any event restriction, under only the exponential moment condition.
Reading between the lines
- A natural testable extension is to ask whether the exponential moment conditions (A3) and (A4) can be weakened to finite moments of sufficiently high order; the empirical-process peeling argument suggests that some higher-moment analogue should hold, but the present proof genuinely needs the exponential envelope.
- The high-probability event restriction could likely be removed if one had direct moment control on the inverse of the observed information matrix; the paper's event device is one way to carry the argument, not an intrinsic obstruction.
- The same martingale-plus-empirical-process template should transfer to other partial-likelihood settings, such as stratified or time-dependent covariate versions of the proportional hazards model, where the same missing moment bounds would be needed for global errors.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proves uniform L^p moment bounds for the maximum partial likelihood estimator β̂_n and the Breslow estimator Λ_n in the Cox proportional hazards model, on events with probability tending to one, under assumptions (A1)–(A4), with (A3) required for the Breslow result. It also states Lemma 3, a uniform moment bound for the rescaled empirical process n^{1/2} sup_t |Φ_n(t;β0)−Φ(t;β0)| under the exponential moment condition (A3). The motivation is to provide tools for L^p global error analysis of shape-restricted estimators of the baseline hazard, where the usual CLTs for β̂_n and Λ_n are not enough. Theorem 1 controls the regression coefficient; Theorem 2 controls the Breslow estimator; Lemma 3 is auxiliary but of independent interest.
Significance. If the moment bounds are correct, they close a real gap in the Cox model literature: standard asymptotic results give weak convergence but not uniform boundedness of moments of arbitrary order, which is needed for L^p errors of shape-restricted baseline hazard estimators. The paper's high-probability-event formulation is honest and useful, and the proofs combine martingale calculations with empirical-process inequalities in a mostly careful way. The results are plausible and the statements are clean. However, two technical problems in the proofs—the random-β* factor in Theorem 2 and the covering-number claim in Lemma 3—must be repaired before the results can be considered established.
major comments (2)
- [§3, proof of Theorem 2, term (13)] The control of term (13) is not established. After the Cauchy-Schwarz step, the proof requires a uniform bound on E[|n^{-1} Σ Z_i^2 e^{(β*)'Z_i}|^p], where β* is the random point from the mean value theorem. The text says this factor is bounded by the proof of Theorem 1 using assumption (A4), but (A4) only controls sup_{|β−β0|≤ε} E[Z_k^{2q} e^{qβ'Z}] for fixed β in a fixed neighborhood; it does not control an empirical average with a data-dependent β* that can leave that neighborhood on a set of positive probability for finite n. For example, β0=1 and Z=−Y with Y having density proportional to y^{-3} on [1,∞) satisfy (A1)–(A4), but E[Y^2 e^{aY}] = ∞ for every a>0, so the cited argument cannot yield a finite bound. The proof needs either an additional event forcing |β*−β0|≤ε (with the complement handled separately) or a different bound exploiting the event A2_n. As written, the proof of Theorem 2 is incomplete.
- [§3, proof of Lemma 3 and term (11)] The covering-number bound N(ε||F||_{L2(Q)}, F, L2(Q)) ≲ 1/ε is not justified and is in general false for the class of monotone functions. Theorem 2.7.5 in van der Vaart and Wellner gives an upper bound of the form log N_{[]}(ε, F, L2(Q)) ≲ 1/ε, which does not imply N ≲ 1/ε; the covering number can grow exponentially in 1/ε. The proof of Lemma 3 (and the analogous argument for term (11)) should use the logarithmic bound instead. The entropy integral J(1,F) is still finite under this corrected bound, so the conclusions of Lemma 3 and of term (11) remain plausible, but the displayed inequality as written is wrong.
minor comments (6)
- [Abstract] The phrase 'of the the Breslow estimator' contains a duplicated article; it should read 'of the Breslow estimator'.
- [§3, proof of Theorem 1, equations (5)–(6)] The definitions of D1_n(t;β) and D2_n(t;β) use e^{β'_0 Z_i}; the derivative with respect to β should contain e^{β' Z_i}.
- [§3, proof of Theorem 2, term (13)] In the Cauchy-Schwarz display, the first factor contains e^{β'_0 Z_i} while the second contains e^{(β*)' Z_i}; both should contain e^{(β*)' Z_i} to match the preceding bound.
- [§3, proof of Lemma 3] The sentence 'From Theorem 2.7.5 and in van der Vaart and Wellner (1996)' has an extra 'and in', and several integrals are written over R×Rp rather than R×R^d.
- [§3, proof of Theorem 2, definition of A2_n] The convergence sup_t |Φ_n(t;β*)−Φ(t;β0)| → 0 is asserted by analogy with a fixed-β lemma, but β* is random; a uniform-in-β argument over a neighborhood of β0 should be spelled out.
- [Assumptions] Assumption (A3) requires finite exponential moments of every order and is substantially stronger than the moment conditions in Andersen and Gill (1982); a remark on how restrictive this is in applications would be helpful.
Circularity Check
No circularity: the moment bounds are derived from external asymptotic results and stated moment assumptions; self-citations are only motivational.
full rationale
The central claims do not reduce to their inputs. Lemma 3 proves an unconditional uniform moment bound for the empirical process by bounding the entropy of a class of functions using van der Vaart and Wellner's Theorem 2.14.1 and assumption (A3); no part of the target bound is assumed. Theorem 1 derives the moment bound for the maximum partial likelihood estimator from the Tsiatis score representation, the Andersen-Gill convergence of the observed information matrix, and the Kalbfleisch-Prentice martingale representation, with assumption (A4) supplying the needed moment control of Z^2 e^{beta'Z}. Theorem 2 combines Lemma 3, Theorem 1, and the external result (4) from Lopuhaa and Nane. The self-references to Durot and Musta (2019) and Lopuhaa and Musta (2018) appear only as motivation for why such moment bounds are needed and as a precedent for working on an event of probability tending to one; they are not load-bearing for any theorem. The skeptical concern about the random beta* factor in the proof of Theorem 2 is a possible correctness gap in the Cauchy-Schwarz step, not a circularity: the desired bound is not built into the assumptions or into the cited results. No fitted parameter is renamed as a prediction, and no uniqueness theorem or ansatz is imported from the authors' prior work. The derivation is self-contained conditional on the stated assumptions and external asymptotics.
Assumptions & free parameters
assumptions (8)
- domain assumption Cox proportional hazards model with conditionally independent censoring given Z and non-informative censoring.
- domain assumption (A1) τG < τF ≤ ∞ and P(T = τG) > 0.
- domain assumption (A2) sup_{|β−β0|≤ε} E[|Z|^2 e^{2β'Z}] < ∞.
- domain assumption (A3) E[e^{q β0'Z}] < ∞ for all q ≥ 1.
- domain assumption (A4) sup_{|β−β0|≤ε} E[Z_k^{2q} e^{qβ'Z}] < ∞ for all q ≥ 1 and k = 1,...,d.
- standard math Consistency and asymptotic normality of β̂_n and convergence of n^{-1}S'' to a nonsingular Σ (Tsiatis 1981; Andersen and Gill 1982).
- standard math Maximal inequality for empirical processes, Theorem 2.14.1 in van der Vaart and Wellner (1996).
- standard math Counting process martingale representation of the score (Kalbfleisch and Prentice 2002).
Cite this review
Pith. "Pith review of Results on standard estimators in the Cox model." pith.science (2026). https://pith.science/paper/XN653TQN
@misc{pith2026190807456,
author = {Pith},
title = {Pith review of: Results on standard estimators in the Cox model},
year = {2026},
howpublished = {\url{https://pith.science/paper/XN653TQN}},
note = {Machine review of arXiv:1908.07456}
}
abstract
We consider the Cox regression model and prove some properties of the maximum partial likelihood estimator $\hat\beta_n$ and of the the Breslow estimator $\Lambda_n$. The asymptotic properties of these estimators have been widely studied in the literature but we are not aware of a reference where it is shown that they have uniformly bounded moments. These results are needed, for example, when studying global errors of shape restricted estimators of the baseline hazard function.
Reference graph
Works this paper leans on
-
[1]
Andersen, P. K. and Gill, R. D. (1982), Cox’s regression model for c ounting processes: a large sample study. Ann. Statist. , 10.4, 1100–1120
work page 1982
-
[2]
Cox, D. R. (1972), Regression models and life-tables. J. Roy. Statist. Soc. Ser. B , 34, 187–220
work page 1972
-
[3]
Cox, D. R.(1975), Partial likelihood. Biometrika, 62.2, 269–276
work page 1975
-
[4]
On the $L_p$-error of the Grenander-type estimator in the Cox model
Durot, C. and Musta, E. (2019), On the Lp error of the Grenander-type estimator in the Cox model. https://arxiv.org/abs/1907.06933 . Kalbfleisch, J. D. and Prentice, R. L. (2002), The statistical analy sis of failure time data. Wiley Series in Probability and Statistics , Second edition, Wiley-Interscience [John Wiley & Sons], Hoboken, NJ, xiv+439. Lopuha¨...
work page Pith review arXiv 2019
-
[5]
Tsiatis, A. A. (1981), A large sample study of Cox’s regression mode l. Ann. Statist. , 9.1, 93–108. van der Vaart, A. W. and Wellner, J. A. (1996), Weak convergence and empirical processes. Springer Series in Statistics , Springer-Verlag, New York, with applications to statistics, xvi+508. 16
work page 1981
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.