{"id":"47a2e3cc-12ec-45a8-8cf0-719c37c1742f","arxiv_id":"2501.18995","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"For Gaussian designs, the ridge-penalized piecewise exponential survival model has an asymptotically exact scalar surrogate whose saddle point predicts the estimator and test error.","lead":"This paper develops an exact asymptotic theory for a common survival-analysis model, the piecewise exponential proportional hazards model, when the number of predictors grows proportionally to the number of patients. The result lets researchers compute the effect of ridge regularization on prediction quality and parameter estimates in high-dimensional settings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The CGMT proof applies only to compactly constrained optima, yet no argument shows the unconstrained minimizers of Ln stay in a fixed compact set; this is load-bearing for Theorems 1 and 2.","rationale":"The paper's central claim is the CGMT-based asymptotic equivalence for the unconstrained penalized maximum likelihood problem. The single most load-bearing condition for that claim is that the compact sets introduced in Section 4 do not change the optimum: Theorem 3 is explicitly a compact-set comparison, and the 'no difference' assertion is not proved. This is precisely the reader's weakest_assumption, and the concern is legitimate. The issue is not an external disagreement with consensus; it is an internal gap in the proof of the main theorem. However, the gap is plausibly repairable by a standard coercivity argument, and the matching simulations provide empirical support that the asymptotic characterization is correct. The reader's verdict of CONDITIONAL is therefore appropriate: the proof currently has a missing step, but the central claim is credible and likely fixable. The omega equation inconsistency is real but secondary: it affects the explicit self-consistent characterization in Theorem 2 and the reproducibility of the numerical solution, yet it does not undermine the proof of convergence of the optimal value, and it is probably a typo in the displayed equation rather than a structural failure. For these reasons the stress-test does not change the reader's verdict.","tokens_in":19608,"tokens_out":17843,"duration_ms":193315,"concrete_test":"Insert a lemma proving boundedness in probability of the unconstrained minimizers: for every epsilon > 0 there exists R such that P(||(beta_hat, omega_hat)|| <= R) >= 1 - epsilon for all sufficiently large n, using the lower bound (1/n) sum_i min_x[g(x,omega,T_i,Delta_i) - Delta_i log lambda(T_i|omega)] >= -C_n with C_n = O_P(1), plus ridge coercivity in beta and omega. If this lemma holds, the compact-set restriction in Section 4 is justified and Theorems 1-2 follow as stated; if it fails, the theorem statements need modification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 restricts the min over beta, xi and the max over phi to convex compact sets, asserting that 'intuition suggests' there is no difference from the unbounded problem. This is load-bearing because Theorem 3 is stated only for compact convex sets, and the convergence of the optimal value in Theorem 1 and the minimizer localization in Theorem 2 both require that, with probability tending to 1, the unconstrained minimizers of Ln lie in a fixed compact set and that the maximizing phi is attained at bounded values. No such argument is supplied; without it, the CGMT comparison in equations (26)-(29) applies to a different, compact-constrained problem, and the limit in (12) is not established for the actual objective. The gap appears closable: for each observation, min_x[g(x,omega,T,Delta)-Delta log lambda(T|omega)] is bounded below by a quantity of order -|log Psi_k(T)| up to constants, and the 1/n average of these lower bounds converges by the law of large numbers to a finite constant under mild conditions on the event-time density; together with the ridge penalties this yields high-probability norm bounds on the minimizers. But this coercivity argument is absent from the paper. A secondary issue flagged by the reader is that the displayed RS equation (21) for omega does not match the derivative calculation in Proposition 8; that appears to be a separate typo/correctness issue in the statement of Theorem 2, less central than the compactness gap but still requiring correction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes a ridge-penalized piecewise exponential proportional hazards model in a high-dimensional regime where p/n tends to a positive constant, with Gaussian uncorrelated covariates and independent censoring. The main results are Theorem 1, which asserts that the normalized minimum of the penalized log-likelihood converges in probability to the saddle point of the scalar surrogate objective L in Eq. (13), and Theorem 2, which asserts that the estimator's projections w_hat_n, v_hat_n and the baseline log-hazard parameters omega_hat_n converge to the solution of the self-consistent equations (17)-(21). From these, the paper derives asymptotic expressions for prediction metrics such as Harrell's c-index and an oracle integrated Brier score. The proof follows the standard CGMT roadmap: the problem is rewritten as a min-max problem, Gordon's comparison is applied, directions are optimized, Moreau envelopes are introduced, and convexity arguments pass to the deterministic limit. Numerical experiments in Section 5 compare the theory with simulations for several values of the aspect ratio and ridge strength. Appendices A-D contain technical lemmas on integrability, derivatives of the Moreau envelope, pointwise convergence of the auxiliary objective, and localization of the minimizer.","tokens_in":19911,"tokens_out":14700,"duration_ms":141347,"significance":"If the proof is completed and the displayed equations are corrected, this would be a useful rigorous contribution: it extends Convex Gaussian Min-Max 'exact asymptotics' from generalized linear models to a parametric survival model with censoring, and it yields explicit, falsifiable predictions for discrimination and calibration metrics. The surrogate objective and the self-consistent equations are derived from CGMT rather than fitted to simulations, and the numerical experiments are an independent check; the paper also provides reproducible code. The main caveats are the unproved compactness/coercivity step in the CGMT application and the mismatch between Eq. (21) and the derivative computation in Proposition 8. Both are localizable and appear fixable, but they currently prevent the theorems, as stated, from being considered established.","major_comments":[{"comment":"The proof restricts the minimization over beta, xi and the maximization over phi to compact convex sets, with the statement that 'Intuition suggests that if a saddle point exists and the set is sufficiently large, then there is not going to be any difference between the bounded and unbounded problem.' No argument is supplied that the unconstrained minimizers of L_n in Eq. (8), or the optimal dual variables, stay in a fixed compact set with probability tending to one. Since the CGMT in Theorem 3 applies only to compact, convex sets, the comparison in Eqs. (26)-(29) and the limit in Eq. (12) are established only for a compact-constrained problem, not for the actual objective in Theorem 1. The localization part of Theorem 2 inherits this issue: the use of implication (28) and of Eqs. (109)-(111) in Appendix D requires exactly this boundedness-in-probability. A coercivity argument based on lower bounds for min_x g(x, omega, T, Delta) together with the ridge penalties would close the gap, but such an argument is absent from the manuscript.","section":"Section 4 and Theorem 3; Eqs. (12), (26)-(29)"},{"comment":"Equation (21) does not follow from the stationarity condition computed in Proposition 8. Differentiating L in Eq. (13) with respect to omega_k and setting the derivative to zero gives alpha * omega_k = E[Delta * psi_k(T)] - exp(omega_k) * E[Psi_k(T) * exp(xi_hat)], where xi_hat is the proximal point. Solving this equation yields omega_k = (1/alpha) * E[Delta * psi_k(T)] - W0( (1/alpha) * E[Psi_k(T) * exp(xi_hat)] * exp{ (1/alpha) * E[Delta * psi_k(T)] } ). The displayed equation instead uses E[exp(xi_hat) * psi_k(T)] as the coefficient, E[Delta * Psi_k(T)] inside the exponential, and an unexplained factor eta in the denominators. Because Eq. (21) is used in Section 5 to produce the theoretical curves, this discrepancy must be corrected and the derivation shown; if it is a typographical error, the numerical predictions should be re-checked against the corrected equation.","section":"Section 3, Theorem 2, Eq. (21)"}],"minor_comments":[{"comment":"The symbol tau is used both for the user-specified knots of the piecewise exponential model and for the variational parameter tau in the surrogate objective L. This is confusing; consider renaming one of them.","section":"Section 3 and 4"},{"comment":"In the sentence after Eq. (91), the text reads 'L^{omega,w,v}_n(.) converges in probability to L^{omega,w,v}_n(.)'; the limit should be L^{omega,w,v}(.), not L^{omega,w,v}_n(.).","section":"Appendix C, proof of Proposition 11"},{"comment":"The theorems are stated without the regularity conditions that the appendices actually use, such as finite moments of T, positive mass in each interval, and bounded knots. These conditions should be stated in the main text so that the statements are self-contained.","section":"Section 3, Theorems 1 and 2"},{"comment":"In the comparison of theory and simulations, the paper says 'because of the convergence in probability established in (2)'; this should refer to Eq. (12) or to Theorem 2, not to Eq. (2).","section":"Section 5"},{"comment":"The minimum min_x g(x, omega, Delta, T) is used in several places, but for Delta = 0 the infimum is not attained; replace 'min' by 'inf' or add a brief note on the Delta = 0 case.","section":"Appendix A, Proposition 1 and Proposition 4"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision rather than rejection because the compactness gap appears closable with a coercivity argument, and the Eq. (21) issue looks like a combination of typos. However, the extra eta factor and the swapped psi/Psi and Delta-psi/Delta-Psi terms could change the numerical curves, so the authors need to re-derive Eq. (21) and re-run the experiments after the correction. If both issues are resolved, the paper would be a solid contribution to high-dimensional survival analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read it. First: this is, as far as I can tell, the first rigorous proportional-limit characterization for a parametric survival model — Massa runs the convex Gaussian min-max machinery on the piecewise exponential proportional hazards likelihood with right censoring and gets a scalar surrogate plus replica-symmetric equations for the estimator projections. That is a real extension of the GLM results in [30,31]. Second: the proof as written has a load-bearing gap, and the omega equation in Theorem 2 looks wrong as printed. The verdict is conditional, not a rejection.\n\nWhat it does well. The RS equations are derived, not fitted; the simulations (c-index, integrated Brier score, w and v projections) agree well and the code and data are on GitHub. The paper is honest about scope: uncorrelated Gaussian designs, a finite number of baseline pieces, no Cox limit. Self-citations to the author's replica heuristics are used as motivation, not as input to the proof — that's fine.\n\nSoft spots, in proportion. The bigger one is compactness. Section 4 restricts the min over beta, xi and the max over phi to compact convex sets on 'intuition suggests', but Theorems 1 and 2 are stated for the unconstrained problem, and the CGMT comparison in (26)-(29) only applies to the restricted one. Without an argument that the unconstrained minimizers are tight, neither the value convergence nor the minimizer localization follows. The stress-test sketch of a coercivity proof (per-observation lower bound of order -|log Psi_k(T)|, plus LLN and the ridge penalties) looks plausible, but it is not in the paper. Closable, but real.\n\nThe smaller one is eq (21). Differentiating the surrogate L and using Proposition 8, the omega fixed point should read omega_k = (1/alpha) E[Delta psi_k] - W0((1/alpha) E[Psi_k e^{xihat}] exp{(1/alpha) E[Delta psi_k]}). The printed version has psi and Psi swapped and a spurious eta in the prefactors. I read that as a typo in a theorem statement — the numerics presumably used the intended equations, and the GitHub repo makes that checkable — but it has to be fixed before the theorem can be trusted.\n\nMinor: the appendix caps phi at lambda_max and then extends to all phi >= 0; the extension is sketched, not proved. Secondary.\n\nWho this is for: the CGMT/replica crowd and anyone comparing estimators in high-dimensional survival. I'd send it to a serious referee. The strategy is credible, the target is useful, and the two fixes are the price of admission. I'd cite it once the fixed-point equations check out.","headline":"A real first CGMT treatment of a survival model with a fixable proof gap and a likely-typo'd fixed-point equation; deserves refereeing, not desk rejection.","tokens_in":20416,"tokens_out":16781,"would_cite":true,"duration_ms":137823,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62N01","62J07","62E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"A scalar saddle point captures the high-dimensional survival model's prediction error.","keywords":["high-dimensional statistics","survival analysis","piecewise exponential model","proportional hazards","ridge regularization","Convex Gaussian Min-Max theorem","replica method","censored data"],"falsifier":"Set $n=400$, $p=ζn$ for several $ζ$, generate log-logistic survival data with 40% censoring as in Section 5, and trace the empirical minimizer along a ridge path $η\\in(0.1,6)$. If the squared norm of $β̂$ or the orthogonal projection $ṥ̂_n$ grows without bound as $ζ$ approaches 1 or $η$ decreases, while equations (17)-(21) still have a finite solution, Theorem 2's convergence claim would fail.","tokens_in":19377,"feed_emoji":"⏳","tokens_out":6905,"duration_ms":67641,"temperature":0.7,"pith_summary":"This paper claims that in the proportional asymptotics regime, where $p/n \\to \\zeta > 0$, the ridge-regularized piecewise exponential proportional hazards model has a complete low-dimensional description. For uncorrelated Gaussian covariates, the paper proves that the optimal penalized log-likelihood converges in probability to the saddle point of a scalar surrogate function, and that the estimator's projections along and orthogonal to the true coefficient vector converge to solutions of self-consistent equations. A reader should care because this turns a high-dimensional survival-fitting problem into a small system that yields exact expected prediction metrics—Harrell's c-index and an integrated Brier score ratio—as functions of $p/n$ and of the ridge penalty. The paper presents this as a rigorous step toward previously heuristic replica-method results for the Cox model.","feed_headline":"Ridge survival model gets exact high-dimensional asymptotics","feed_subtitle":"A scalar saddle point predicts the penalized estimator's c-index and Brier score as p/n grows.","key_machinery":"The carrying object is the Convex Gaussian Min-Max theorem (CGMT), a comparison principle that transfers Gaussian min-max optimization problems into simpler surrogate problems with the same optimal value and optimizer. Applied after rewriting the objective as a saddle point with a Lagrange multiplier $\\varphi$, the theorem leads to a scalar surrogate built from the Moreau envelope of $g(x,\\omega,\\Delta,T)=\\Lambda(T|\\omega)e^x-\\Delta x$; its proximal operator, expressed through the Lambert W function, closes the self-consistent equations. The number of baseline-hazard parameters $\\ell$ is fixed as $n\\to\\infty$, which is what keeps the limiting problem finite-dimensional.","core_discovery":"The central claim is that for data generated as $Y|X \\sim f_0(\\cdot|X'\\beta_0)$ with $X \\sim N(0,I_p)$ and right censoring independent of covariates, the minimum of the penalized piecewise exponential log-likelihood $L_n(\\omega,\\beta)$ converges in probability to the saddle point of the scalar function $L(\\omega,w,v,\\varphi,\\tau)$ built from a Moreau envelope, and the estimator projections $\\beta_0'\\hat\\beta_n/\\|\\beta_0\\|$, $\\|P_{\\beta_0^\\perp}\\hat\\beta_n\\|$, and $\\hat\\omega_n$ converge to the unique solution $(w^\\star,v^\\star,\\omega^\\star)$ of the self-consistent equations. Because the out-of-sample linear predictor converges in distribution to $w^\\star Z_0 + v^\\star Q$, quantities such as the c-index and the ideal integrated Brier score can be evaluated exactly from the scalar limit. The proof uses the Convex Gaussian Min-Max theorem to replace the original high-dimensional min-max problem with an asymptotically equivalent scalar process, then identifies that process's saddle point.","pith_inferences":["If the compactness assumption were proved rather than asserted, the same proof structure would likely extend to correlated Gaussian designs, with the population covariance spectrum entering the self-consistent equations.","Because the replica heuristic only requires asymptotic Gaussianity of the linear predictor, the saddle-point characterization may survive for sub-Gaussian covariates; the author leaves this as a conjecture.","If the number of intervals $ℓ$ grows with $n$ so that the baseline hazard becomes saturated, the scalar limit might approach a Cox-model limit, but the proof as written requires $ℓ$ fixed and would need an independent argument.","The convergence-in-probability result suggests that fixed-point iteration on the self-consistent equations is a reliable computational shortcut; one could test how its accuracy degrades for small $n$ or $ζ$ near zero."],"forward_implications":["For any $ζ=p/n>0$, the expected test c-index is nearly flat in ridge strength $η$, while the value of $η$ that minimizes the integrated Brier score ratio grows with $ζ$, so denser problems need more regularization just to beat the null model.","The estimator's behavior splits cleanly: one scalar measures alignment with the true coefficient direction and another measures shrinkage in the orthogonal space, and each has a deterministic limit.","Prediction metrics can be computed from the saddle point without fitting the high-dimensional model, making regularization-path comparisons cheap.","The scalar limit provides a rigorous benchmark that heuristic replica calculations for the Cox model must reproduce."],"supporting_citations":[{"why":"Supplies the Convex Gaussian Min-Max theorem that converts the high-dimensional saddle point into a scalar surrogate.","marker":"[29]"},{"why":"Provides the precise-error-analysis framework, including the norm variational representation and the lemmas used to pass from pointwise convergence to saddle-point convergence.","marker":"[30]"},{"why":"Gives the base result that Theorem 1 generalizes and identifies the spectral complications for correlated designs.","marker":"[31]"},{"why":"Defines the piecewise exponential proportional hazards model whose likelihood is the object of study.","marker":"[32]"},{"why":"Provides the earlier cavity-method heuristics that this paper puts on rigorous footing for the piecewise exponential setting.","marker":"[44]"},{"why":"Uses replica analysis for survival regression and frames the Cox-model conjectures this work takes a first step toward proving.","marker":"[8]"},{"why":"Supplies Gordon's Gaussian comparison inequalities that underlie the CGMT.","marker":"[41]"}],"fun_headline_variants":["Ridge survival model gets exact high-dimensional limits","Proportional asymptotics: saddle point solves piecewise exponential","High-dim survival: ridge penalty yields precise p/n predictions","Survival regression: CGMT proves scalar limit for penalized fit","Exact asymptotics for ridge penalized piecewise exponential hazards"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on the assumption that the true solution of the fitting problem stays inside a fixed bounded region that does not grow with the number of observations; the paper asserts this restriction rather than proving that unconstrained minimizers remain bounded in probability.","fun_headline_variants_meta":{"raw":{"variants":["Ridge survival model gets exact high-dimensional limits","Proportional asymptotics: saddle point solves piecewise exponential","High-dim survival: ridge penalty yields precise p/n predictions","Survival regression: CGMT proves scalar limit for penalized fit","Exact asymptotics for ridge penalized piecewise exponential hazards"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1260,"prompt_tokens":927,"completion_tokens":333,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":251}},"tokens_in":543,"tokens_out":333,"duration_ms":4212,"temperature":1.0,"reasoning_tokens":251,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T21:42:07.167850+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Set $n=400$, $p=ζn$ for several $ζ$, generate log-logistic survival data with 40% censoring as in Section 5, and trace the empirical minimizer along a ridge path $η\\in(0.1,6)$. If the squared norm of $β̂$ or the orthogonal projection $ṥ̂_n$ grows without bound as $ζ$ approaches 1 or $η$ decreases, while equations (17)-(21) still have a finite solution, Theorem 2's convergence claim would fail.","supporting_citations":[{"cited_title":"The gaussian min-max theorem in the presence of convexity, 2015","cited_arxiv_id":null,"evidence_quote":"Supplies the Convex Gaussian Min-Max theorem that converts the high-dimensional saddle point into a scalar surrogate."},{"cited_title":"Precise error analysis of regularized m -estimators in high dimen- sions","cited_arxiv_id":null,"evidence_quote":"Provides the precise-error-analysis framework, including the norm variational representation and the lemmas used to pass from pointwise convergence to saddle-point convergence."},{"cited_title":"Learning curves of generic features maps for realistic datasets with a teacher-student model*","cited_arxiv_id":null,"evidence_quote":"Gives the base result that Theorem 1 generalizes and identifies the spectral complications for correlated designs."},{"cited_title":"Piecewise Exponential Models for Survival Data with Covariates","cited_arxiv_id":null,"evidence_quote":"Defines the piecewise exponential proportional hazards model whose likelihood is the object of study."},{"cited_title":"Penalization-induced shrinking without rotation in high dimensional glm regression: a cavity analysis","cited_arxiv_id":null,"evidence_quote":"Provides the earlier cavity-method heuristics that this paper puts on rigorous footing for the piecewise exponential setting."},{"cited_title":"Replica analysis of overfitting in regression models for time-to-event data","cited_arxiv_id":null,"evidence_quote":"Uses replica analysis for survival regression and frames the Cox-model conjectures this work takes a first step toward proving."},{"cited_title":"Some inequalities for gaussian processes and applications","cited_arxiv_id":null,"evidence_quote":"Supplies Gordon's Gaussian comparison inequalities that underlie the CGMT."}],"review_version":1}