REVIEW 2 major objections 5 minor 50 references
Proportional asymptotics of piecewise exponential proportional hazards models
T0 review · 2 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read A scalar saddle point captures the high-dimensional survival model's prediction error.
desk verdict A real first CGMT treatment of a survival model with a fixable proof gap and a likely-typo'd fixed-point equation; deserves refereeing, not desk rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the Convex Gaussian Min-Max theorem (CGMT), a comparison principle that transfers Gaussian min-max optimization problems into simpler surrogate problems with the same optimal value and optimizer. Applied after rewriting the objective as a saddle point with a Lagrange multiplier $\varphi$, the theorem leads to a scalar surrogate built from the Moreau envelope of $g(x,\omega,\Delta,T)=\Lambda(T|\omega)e^x-\Delta x$; its proximal operator, expressed through the Lambert W function, closes the self-consistent equations. The number of baseline-hazard parameters $\ell$ is fixed as $n\to\infty$, which is what keeps the limiting problem finite-dimensional.
What would settle it
Set $n=400$, $p=ζn$ for several $ζ$, generate log-logistic survival data with 40% censoring as in Section 5, and trace the empirical minimizer along a ridge path $η\in(0.1,6)$. If the squared norm of $β̂$ or the orthogonal projection $ṥ̂_n$ grows without bound as $ζ$ approaches 1 or $η$ decreases, while equations (17)-(21) still have a finite solution, Theorem 2's convergence claim would fail.
Extended reading notes
Core claim
The central claim is that for data generated as $Y|X \sim f_0(\cdot|X'\beta_0)$ with $X \sim N(0,I_p)$ and right censoring independent of covariates, the minimum of the penalized piecewise exponential log-likelihood $L_n(\omega,\beta)$ converges in probability to the saddle point of the scalar function $L(\omega,w,v,\varphi,\tau)$ built from a Moreau envelope, and the estimator projections $\beta_0'\hat\beta_n/\|\beta_0\|$, $\|P_{\beta_0^\perp}\hat\beta_n\|$, and $\hat\omega_n$ converge to the unique solution $(w^\star,v^\star,\omega^\star)$ of the self-consistent equations. Because the out-of-sample linear predictor converges in distribution to $w^\star Z_0 + v^\star Q$, quantities such as the c-index and the ideal integrated Brier score can be evaluated exactly from the scalar limit. The proof uses the Convex Gaussian Min-Max theorem to replace the original high-dimensional min-max problem with an asymptotically equivalent scalar process, then identifies that process's saddle point.
Load-bearing premise
The argument depends on the assumption that the true solution of the fitting problem stays inside a fixed bounded region that does not grow with the number of observations; the paper asserts this restriction rather than proving that unconstrained minimizers remain bounded in probability.
Editorial extensions
If this is right
- For any $ζ=p/n>0$, the expected test c-index is nearly flat in ridge strength $η$, while the value of $η$ that minimizes the integrated Brier score ratio grows with $ζ$, so denser problems need more regularization just to beat the null model.
- The estimator's behavior splits cleanly: one scalar measures alignment with the true coefficient direction and another measures shrinkage in the orthogonal space, and each has a deterministic limit.
- Prediction metrics can be computed from the saddle point without fitting the high-dimensional model, making regularization-path comparisons cheap.
- The scalar limit provides a rigorous benchmark that heuristic replica calculations for the Cox model must reproduce.
Reading between the lines
- If the compactness assumption were proved rather than asserted, the same proof structure would likely extend to correlated Gaussian designs, with the population covariance spectrum entering the self-consistent equations.
- Because the replica heuristic only requires asymptotic Gaussianity of the linear predictor, the saddle-point characterization may survive for sub-Gaussian covariates; the author leaves this as a conjecture.
- If the number of intervals $ℓ$ grows with $n$ so that the baseline hazard becomes saturated, the scalar limit might approach a Cox-model limit, but the proof as written requires $ℓ$ fixed and would need an independent argument.
- The convergence-in-probability result suggests that fixed-point iteration on the self-consistent equations is a reliable computational shortcut; one could test how its accuracy degrades for small $n$ or $ζ$ near zero.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes a ridge-penalized piecewise exponential proportional hazards model in a high-dimensional regime where p/n tends to a positive constant, with Gaussian uncorrelated covariates and independent censoring. The main results are Theorem 1, which asserts that the normalized minimum of the penalized log-likelihood converges in probability to the saddle point of the scalar surrogate objective L in Eq. (13), and Theorem 2, which asserts that the estimator's projections w_hat_n, v_hat_n and the baseline log-hazard parameters omega_hat_n converge to the solution of the self-consistent equations (17)-(21). From these, the paper derives asymptotic expressions for prediction metrics such as Harrell's c-index and an oracle integrated Brier score. The proof follows the standard CGMT roadmap: the problem is rewritten as a min-max problem, Gordon's comparison is applied, directions are optimized, Moreau envelopes are introduced, and convexity arguments pass to the deterministic limit. Numerical experiments in Section 5 compare the theory with simulations for several values of the aspect ratio and ridge strength. Appendices A-D contain technical lemmas on integrability, derivatives of the Moreau envelope, pointwise convergence of the auxiliary objective, and localization of the minimizer.
Significance. If the proof is completed and the displayed equations are corrected, this would be a useful rigorous contribution: it extends Convex Gaussian Min-Max 'exact asymptotics' from generalized linear models to a parametric survival model with censoring, and it yields explicit, falsifiable predictions for discrimination and calibration metrics. The surrogate objective and the self-consistent equations are derived from CGMT rather than fitted to simulations, and the numerical experiments are an independent check; the paper also provides reproducible code. The main caveats are the unproved compactness/coercivity step in the CGMT application and the mismatch between Eq. (21) and the derivative computation in Proposition 8. Both are localizable and appear fixable, but they currently prevent the theorems, as stated, from being considered established.
major comments (2)
- [Section 4 and Theorem 3; Eqs. (12), (26)-(29)] The proof restricts the minimization over beta, xi and the maximization over phi to compact convex sets, with the statement that 'Intuition suggests that if a saddle point exists and the set is sufficiently large, then there is not going to be any difference between the bounded and unbounded problem.' No argument is supplied that the unconstrained minimizers of L_n in Eq. (8), or the optimal dual variables, stay in a fixed compact set with probability tending to one. Since the CGMT in Theorem 3 applies only to compact, convex sets, the comparison in Eqs. (26)-(29) and the limit in Eq. (12) are established only for a compact-constrained problem, not for the actual objective in Theorem 1. The localization part of Theorem 2 inherits this issue: the use of implication (28) and of Eqs. (109)-(111) in Appendix D requires exactly this boundedness-in-probability. A coercivity argument based on lower bounds for min_x g(x, omega, T, Delta) together with the ridge penalties would close the gap, but such an argument is absent from the manuscript.
- [Section 3, Theorem 2, Eq. (21)] Equation (21) does not follow from the stationarity condition computed in Proposition 8. Differentiating L in Eq. (13) with respect to omega_k and setting the derivative to zero gives alpha * omega_k = E[Delta * psi_k(T)] - exp(omega_k) * E[Psi_k(T) * exp(xi_hat)], where xi_hat is the proximal point. Solving this equation yields omega_k = (1/alpha) * E[Delta * psi_k(T)] - W0( (1/alpha) * E[Psi_k(T) * exp(xi_hat)] * exp{ (1/alpha) * E[Delta * psi_k(T)] } ). The displayed equation instead uses E[exp(xi_hat) * psi_k(T)] as the coefficient, E[Delta * Psi_k(T)] inside the exponential, and an unexplained factor eta in the denominators. Because Eq. (21) is used in Section 5 to produce the theoretical curves, this discrepancy must be corrected and the derivation shown; if it is a typographical error, the numerical predictions should be re-checked against the corrected equation.
minor comments (5)
- [Section 3 and 4] The symbol tau is used both for the user-specified knots of the piecewise exponential model and for the variational parameter tau in the surrogate objective L. This is confusing; consider renaming one of them.
- [Appendix C, proof of Proposition 11] In the sentence after Eq. (91), the text reads 'L^{omega,w,v}_n(.) converges in probability to L^{omega,w,v}_n(.)'; the limit should be L^{omega,w,v}(.), not L^{omega,w,v}_n(.).
- [Section 3, Theorems 1 and 2] The theorems are stated without the regularity conditions that the appendices actually use, such as finite moments of T, positive mass in each interval, and bounded knots. These conditions should be stated in the main text so that the statements are self-contained.
- [Section 5] In the comparison of theory and simulations, the paper says 'because of the convergence in probability established in (2)'; this should refer to Eq. (12) or to Theorem 2, not to Eq. (2).
- [Appendix A, Proposition 1 and Proposition 4] The minimum min_x g(x, omega, Delta, T) is used in several places, but for Delta = 0 the infimum is not attained; replace 'min' by 'inf' or add a brief note on the Delta = 0 case.
Circularity Check
No significant circularity: the surrogate objective and replica-symmetric equations are derived from the CGMT and convex analysis, not fitted to simulations; author self-citations are motivational only.
full rationale
The paper's central claims derive the asymptotic optimal value of the penalized log-likelihood from the Convex Gaussian Min-Max Theorem, reducing the high-dimensional problem to the scalar saddle-point problem L in equation (13). The replica-symmetric equations (17)-(21) are then obtained by differentiating the expected Moreau-envelope surrogate, as in Propositions 7 and 8, and solving the stationarity conditions. No parameter appearing in the surrogate or in the RS equations is fitted to the simulations; the numerical experiments in Section 5 are an independent check against the theoretical curves. The author's prior replica-based papers [44] and [9] are cited only as motivation or as heuristic results to be rigorously justified, not as input to the proof of Theorems 1 and 2. The compactness restriction in Section 4 is a genuine rigor gap, since Theorem 3 requires compact convex sets and the paper asserts without proof that the unconstrained optimizers remain in a fixed compact set; however, this is a correctness or proof-completeness concern, not a circularity. Similarly, the apparent mismatch in equation (21) relative to Proposition 8 is a typographical or correctness issue, not a circular reduction. The derivation chain is therefore self-contained conditional on the cited external CGMT theorem, and the predictions are not equivalent to the inputs by construction.
Assumptions & free parameters
free parameters (4)
- zeta
- eta
- alpha =
0.01 in simulations
- ell =
11 intervals in simulations
assumptions (5)
- domain assumption Covariates are Gaussian and uncorrelated, X ~ N(0, I_p).
- domain assumption Censoring is uninformative and independent of covariates.
- domain assumption The true event time depends on covariates only through the linear predictor X' beta0.
- domain assumption The baseline hazard is piecewise constant on a finite set of intervals, with ell fixed as n tends to infinity.
- ad hoc to paper The global minimizers and saddle point variables stay in fixed compact sets independent of n.
Cite this review
Pith. "Pith review of Proportional asymptotics of piecewise exponential proportional hazards models." pith.science (2026). https://pith.science/paper/HD4O3OQY
@misc{pith2026250118995,
author = {Pith},
title = {Pith review of: Proportional asymptotics of piecewise exponential proportional hazards models},
year = {2026},
howpublished = {\url{https://pith.science/paper/HD4O3OQY}},
note = {Machine review of arXiv:2501.18995}
}
abstract
We study the flexible piecewise exponential model in a high dimensional setting where the number of covariates $p$ grows proportionally to the number of observations $n$ and under the hypothesis of random uncorrelated Gaussian designs. We prove rigorously that the optimal ridge penalized log-likelihood of the model converges in probability to the saddle point of a surrogate objective function. The technique of proof is the Convex Gaussian Min-Max theorem of Thrampoulidis, Oymak and Hassibi. An important consequence of this result, is that we can study the impact of the ridge regularization on the estimates of the parameter of the model and the prediction error as a function of the ratio $p/n > 0$. Furthermore, these results represent a first step toward rigorously proving the (conjectured) correctness of several results obtained with the heuristic replica method for the Cox semi-parametric model.
Figures
Reference graph
Works this paper leans on
-
[1]
Survival Analysis: Techniques for Censored and Truncated Data
JP Klein and ML Moeschberger. Survival Analysis: Techniques for Censored and Truncated Data. Statistics for Biology and Health. Springer New York, 2005
work page 2005
-
[2]
The Statistical Analysis of Failure Time Data
JD Kalbfleisch and RL Prentice. The Statistical Analysis of Failure Time Data. Wiley Series in Probability and Statistics. Wiley, 2011
work page 2011
-
[3]
J Ramjith, C Andolina, T Bousema, and MA Jonker. Flexible time-to-event models for double-interval-censored infectious disease data with clearance of the infection as a competing risk.Frontiers in Applied Mathematics and Statistics, 8, 2022
work page 2022
-
[4]
Survival analysis in infectious disease research: describing events in time
Cole SR and Michael GH. Survival analysis in infectious disease research: describing events in time. AIDS (London, England), 24, 2010
work page 2010
-
[5]
Regression models and life-tables
DR Cox. Regression models and life-tables. Journal of the Royal Statistical Society. Series B (Methodological), 34(2):187–220, 1972
work page 1972
-
[6]
Cox’s regression model for counting processes: A large sample study
PK Andersen and RD Gill. Cox’s regression model for counting processes: A large sample study. The Annals of Statistics, 10(4):1100–1120, 1982
work page 1982
-
[7]
A large sample study of cox’s regression model
AA Tsiatis. A large sample study of cox’s regression model. The Annals of Statistics, 9(1):93 – 108, 1981
work page 1981
-
[8]
Replica analysis of overfitting in regression models for time-to-event data
ACC Coolen, JE Barrett, P Paga, and CJ Perez-Vicente. Replica analysis of overfitting in regression models for time-to-event data. Journal of Physics A: Mathematical and Theoretical, 50(37):375001, 2017
work page 2017
Show all 50 references
-
[9]
Replica analysis of overfitting in regression models for time to event data: the impact of censoring
E Massa, A Mozeika, and ACC Coolen. Replica analysis of overfitting in regression models for time to event data: the impact of censoring. Journal of Physics A: Mathematical and Theoretical, 57(12):125003, 2024
2024
-
[10]
High-Dimensional Statistics: A Non-Asymptotic Viewpoint
MJ Wainwright. High-Dimensional Statistics: A Non-Asymptotic Viewpoint. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 2019
2019
-
[11]
On robust regression with high-dimensional predictors
N El Karoui, D Bean, PJ Bickel, C Lim, and B Yu. On robust regression with high-dimensional predictors. Proceedings of the National Academy of Sciences, 110(36):14557–14562, 2013
2013
-
[12]
High-dimensional generalized linear models and the lasso
SA van de Geer. High-dimensional generalized linear models and the lasso. The Annals of Statistics, 36(2):614 – 645, 2008
2008
-
[13]
Regularization for cox’s proportional hazards model with np-dimensionality
J Bradic, J Fan, and J Jiang. Regularization for cox’s proportional hazards model with np-dimensionality. The Annals of Statistics, 39(6):3092–3120, 2011
2011
-
[14]
Non-asymptotic oracle inequalities for the high-dimensional cox regression via lasso
K Shengchun and B Nan. Non-asymptotic oracle inequalities for the high-dimensional cox regression via lasso. Statistica Sinica, 2014
2014
-
[15]
Regression shrinkage and selection via the lasso
R Tibshirani. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society. Series B (Methodological), 58(1):267–288, 1996
1996
-
[16]
Statistics for High-Dimensional Data: Methods, Theory and Applications
P Bühlmann and S van de Geer. Statistics for High-Dimensional Data: Methods, Theory and Applications . Springer Series in Statistics. Springer Berlin Heidelberg, 2011
2011
-
[17]
Weak Convergence and Empirical Processes: With Applications to Statistics
A Van der Vaart and J Wellner. Weak Convergence and Empirical Processes: With Applications to Statistics . Springer Series in Statistics. Springer New York, 2013
2013
-
[18]
Asymptotic Statistics
AW Van Der Vaart. Asymptotic Statistics. Asymptotic Statistics. Cambridge University Press, 2000
2000
-
[19]
Empirical Processes in M-Estimation
S van de Geer. Empirical Processes in M-Estimation. Cambridge Series in Statistical and Probabilistic Mathe- matics. Cambridge University Press, 2000. 9 Proportional asymptotics of piecewise exponential proportional hazards models A PREPRINT
2000
-
[20]
Fundamentals of High-Dimensional Statistics: With Exercises and R Labs
J Lederer. Fundamentals of High-Dimensional Statistics: With Exercises and R Labs. Springer Texts in Statistics. Springer International Publishing, 2021
2021
-
[21]
The Lasso with general Gaussian designs with applica- tions to hypothesis testing
Michael Celentano, Andrea Montanari, and Yuting Wei. The Lasso with general Gaussian designs with applica- tions to hypothesis testing. The Annals of Statistics, 51(5):2194 – 2220, 2023
2023
-
[22]
Surprises in high-dimensional ridgeless least squares interpolation
TJ Hastie, A Montanari, S Rosset, and RJ Tibshirani. Surprises in high-dimensional ridgeless least squares interpolation. Annals of statistics, 50 2:949–986, 2019
2019
-
[23]
High-dimensional asymptotics of prediction: Ridge regression and classifi- cation
Edgar Dobriban and Stefan Wager. High-dimensional asymptotics of prediction: Ridge regression and classifi- cation. arXiv: Statistics Theory, 2015
2015
-
[24]
The distribution of the Lasso: Uniform control over sparse balls and adaptive parameter tuning
L Miolane and A Montanari. The distribution of the Lasso: Uniform control over sparse balls and adaptive parameter tuning. The Annals of Statistics, 49(4):2313 – 2335, 2021
2021
-
[25]
Optimal storage properties of neural network models
E Gardner and B Derrida. Optimal storage properties of neural network models. Journal of Physics A: Mathe- matical and General, 21(1):271, 1988
1988
-
[26]
On the impact of predictor geometry on the performance on high-dimensional ridge-regularized generalized robust regression estimators
N El Karoui. On the impact of predictor geometry on the performance on high-dimensional ridge-regularized generalized robust regression estimators. Probability Theory and Related Fields, 170:95–175, 2018
2018
-
[27]
A modern maximum-likelihood theory for high-dimensional logistic regression
P Sur and EJ Candès. A modern maximum-likelihood theory for high-dimensional logistic regression. Proceed- ings of the National Academy of Sciences, 116(29):14516–14525, 2019
2019
-
[28]
High dimensional robust m-estimation: asymptotic variance via approximate message passing
DL Donoho and A Montanari. High dimensional robust m-estimation: asymptotic variance via approximate message passing. Probability Theory and Related Fields, 166:935–969, 2013
2013
-
[29]
The gaussian min-max theorem in the presence of convexity, 2015
C Thrampoulidis, S Oymak, and B Hassibi. The gaussian min-max theorem in the presence of convexity, 2015
2015
-
[30]
Precise error analysis of regularized m -estimators in high dimen- sions
C Thrampoulidis, E Abbasi, and B Hassibi. Precise error analysis of regularized m -estimators in high dimen- sions. IEEE Transactions on Information Theory, 64(8):5592–5628, 2018
2018
-
[31]
Learning curves of generic features maps for realistic datasets with a teacher-student model*
B Loureiro, C Gerbelot, H Cui, S Goldt, F Krzakala, M Mézard, and L Zdeborová. Learning curves of generic features maps for realistic datasets with a teacher-student model*. Journal of Statistical Mechanics: Theory and Experiment, 2022(11):114001, 2022
2022
-
[32]
Piecewise Exponential Models for Survival Data with Covariates
M Friedman. Piecewise Exponential Models for Survival Data with Covariates. The Annals of Statistics , 10(1):101 – 113, 1982
1982
-
[33]
Marginal likelihoods based on Cox’s regression and life model
JD Kalbfleisch and RL Prentice. Marginal likelihoods based on Cox’s regression and life model. Biometrika, 60(2):267–278, 1973
1973
-
[34]
Discussion on professor cox’s paper
NE Breslow. Discussion on professor cox’s paper. Journal of the Royal Statistical Society: Series B (Method- ological), 34(2):202–220, 1972
1972
-
[35]
Analysis of survival data under the proportional hazards model
NE Breslow. Analysis of survival data under the proportional hazards model. International Statistical Review / Revue Internationale de Statistique, 43(1):45–57, 1975
1975
-
[36]
Regression Modeling Strategies: With Applications to Linear Models, Logistic Regression, and Survival Analysis
FE Harrell. Regression Modeling Strategies: With Applications to Linear Models, Logistic Regression, and Survival Analysis. Graduate Texts in Mathematics. Springer, 2001
2001
-
[37]
Time-to-event prediction with neural networks and cox regression.J
H Kvamme, Ø Borgan, and I Scheel. Time-to-event prediction with neural networks and cox regression.J. Mach. Learn. Res., 20:129:1–129:30, 2019
2019
-
[38]
Continuous and discrete-time survival prediction with neural networks.Lifetime Data Analysis, 2022
H Kvamme and Ø Borgan. Continuous and discrete-time survival prediction with neural networks.Lifetime Data Analysis, 2022
2022
-
[39]
Convex Analysis
RT Rockafellar. Convex Analysis. Princeton Landmarks in Mathematics and Physics. Princeton University Press, 1997
1997
-
[40]
On the lambert w function
RM Corless, GH Gonnet, DEG Hare, DJ Jeffrey, and DE Knuth. On the lambert w function. Advances in Computational mathematics, 5:329–359, 1996
1996
-
[41]
Some inequalities for gaussian processes and applications
Y Gordon. Some inequalities for gaussian processes and applications. Israel Journal of Mathematics, 50, 1985
1985
-
[42]
Assessing the performance of prediction models: a framework for traditional and novel measures
EW Steyerberg, AJ Vickers, NR Cook, T Gerds, M Gonen, N Obuchowski, MJ Pencina, and MW Kattan. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology (Cambridge, Mass.), 2010
2010
-
[43]
Evaluating the Yield of Medical Tests
FE Harrell, RM Califf, DB Pryor, KL Lee, and RA Rosati. Evaluating the Yield of Medical Tests. JAMA, 247(18):2543–2546, 1982
1982
-
[44]
Penalization-induced shrinking without rotation in high dimensional glm regression: a cavity analysis
E Massa, MA Jonker, and ACC Coolen. Penalization-induced shrinking without rotation in high dimensional glm regression: a cavity analysis. Journal of Physics A: Mathematical and Theoretical, 55(48):485002, 2022. 10 Proportional asymptotics of piecewise exponential proportional...
2022
-
[45]
Universality of regularized regression estimators in high dimensions.The Annals of Statistics, 51(4):1799 – 1823, 2023
Q Han and Y Shen. Universality of regularized regression estimators in high dimensions.The Annals of Statistics, 51(4):1799 – 1823, 2023
2023
-
[46]
Universality laws for high-dimensional learning with random features
H Hu and YM Lu. Universality laws for high-dimensional learning with random features. IEEE Transactions on Information Theory, 69:1932–1964, 2020
1932
-
[47]
Universality of empirical risk minimization
A Montanari and B Saeed. Universality of empirical risk minimization. ArXiv, abs/2202.08832, 2022. A PROPERTIES OF THE MOREAU ENVELOPE FOR THE PIECEWISE EXPONENTIAL MODEL In the following we assume that ∥ω∥ ≤Cbomega, w, v≤ Cβ and τ >0, furthermore τ1 < τ2 < · · ·< τℓ+1 < ∞. Pr...
2022 arXiv
-
[48]
E h wZ0 + vQ 4i = w4E h Z 4 0 i + v4E h Q4 i is finite,
-
[49]
E h g(0, ω, ∆, T) 2i ≤ Pℓ k=1 exp(ωk)(τk+1 − τk) 2 ≤ Pℓ k=1 exp(2ωk) Pℓ k=1(τk+1 − τk)2 which is finite,
-
[50]
Ψk(T ) exp(ωk) W0 τ Λ(T |ω)e∆τ +wZ0+vQ τ Λ(T |ω) # ≤ (τk+1 − τk)EZ0,Q
the term min x g(x, ω, ∆, T) is bounded by a deterministic constant (49), and hence so it is its square. The cross terms can be similarly bounded via the Cauchy-Schwartz inequality. Proposition 3 (FINITE V ARIANCE OF THE DERIV ATIVE) . The random function (ω, w, v, τ) → ˙g(., ...
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.