REVIEW 4 major objections 7 minor 41 references
The existence of maximum likelihood estimate in high-dimensional binary response generalized linear models
T0 review · 4 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read In high-dimensional binary regression with elliptical covariates, the maximum likelihood estimate exists with probability tending to one below an explicit threshold $h_{\mathrm{MLE}}$ and with probability tending to zero above it.
desk verdict A plausible extension of the Candès–Sur phase transition to elliptical covariates, but the proof of the key convergence theorem has gaps that need fixing before the result is fully rigorous. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key object is the threshold $h_{\mathrm{MLE}}$, the minimal expected squared positive part of $\lambda_0 Y+\lambda_1 X-Z$ over $(\lambda_0,\lambda_1)$. It does the heavy lifting through conic geometry: for log-concave links the MLE fails exactly when the data are separable, and for spherical covariates separability is equivalent to a random $(p-1)$-dimensional subspace $L$ (spanned by the nuisance coordinates) intersecting a cone $C(W)$ built from the labels and the signal coordinate. The approximate kinematic formula of convex geometry says such an intersection becomes overwhelmingly likely when $p-1+\delta(C(W))$ exceeds $n$, where the statistical dimension $\delta(C(W))$ is $n$ minus the expected squared distance from a standard normal vector to $C(W)$; the limit of that distance is $n\,h_{\mathrm{MLE}}(\alpha_0,\beta_0,\gamma_0)$. A stochastic approximation lemma, proved by epi-convergence, shows that the empirical minimization defining the threshold converges to its expectation, and condition (2.7) bounds the probability of separation using only the signal coordinate. Assumption 2.6, that the projections $U^{(p)}$ converge in distribution, makes the limit $(Y,X)$ well defined; the moment criterion (2.3) gives a checkable sufficient condition.
What would settle it
Run the linear-programming separation test (3.1) for n=1000, p=$\kappa n$, with spherical covariates whose radial component is Gamma(1,$\theta_p$) scaled so $\mathbb{E}R^2=p+1$, logit link, $\beta_0=0$, across a grid of $\gamma_0$ and $\kappa$, and compare the empirical 50% existence boundary with $h_{\mathrm{MLE}}$ from (2.6); the theorem predicts agreement within the uncertainty band, so a systematic gap would falsify it. The log-normal design, where the projection-limit assumption fails, is the built-in negative control: there the gap is real.
Extended reading notes
Core claim
The central claim is Theorem 2.7. After rotating the covariates so all signal lies in one coordinate, write $(Y^{(p)},X^{(p)})=(V^{(p)},V^{(p)}U^{(p)})$, where $U^{(p)}$ is a single coordinate of the standardized elliptical covariate vector and $P(V^{(p)}=1\,|\,U^{(p)})=\sigma(\beta_0+(\gamma_0/\alpha_0)U^{(p)})$. Assume the link satisfies log-concavity of $\sigma$ and $1-\sigma$, the covariates are full-rank elliptical with $\mathbb{E}R^2/p\to\alpha_0^2$, the coefficient scaling obeys $|\Sigma^{1/2}\beta|\to\gamma_0/\alpha_0$, and the projections converge in distribution, $U^{(p)}\Rightarrow U$. Let $Z\sim N(0,1)$ be independent and define $h_{\mathrm{MLE}}(\alpha_0,\beta_0,\gamma_0)=\min_{\lambda_0,\lambda_1}\mathbb{E}(\lambda_0 Y+\lambda_1 X-Z)_+^2$. Then $\kappa=p/n>h_{\mathrm{MLE}}$ implies the maximum likelihood estimate exists with probability tending to $0$, while $\kappa<h_{\mathrm{MLE}}$ implies it exists with probability tending to $1$, subject to the finite-eighth-moment condition and condition (2.7). This is the same phase transition known for Gaussian logistic regression, now extended to any elliptical family whose one-dimensional projections stabilize, with the boundary computed from the limiting projection distribution.
Load-bearing premise
The load-bearing premise is that the one-dimensional projections of the standardized covariates converge in distribution to a fixed limit, because the threshold $h_{\mathrm{MLE}}$ is computed from that limiting distribution; when it fails, the paper's own simulations show the formula does not hold.
Editorial extensions
If this is right
- Below the threshold $\kappa<h_{\mathrm{MLE}}$, a practitioner can expect the MLE to exist with probability tending to one; above it, data separation occurs almost surely and standard fitting algorithms have no finite solution to find.
- The threshold can be evaluated in advance by solving a two-dimensional convex program for the limiting projection distribution, so the result supplies a practical pre-fit diagnostic.
- The phase transition holds across logit, probit, and cloglog links, and across Gaussian, Gamma, Pareto (with enough moments), and half-normal elliptical covariates, because these satisfy the projection-limit and univariate-separation conditions.
- When the projection-limit assumption fails, as for log-normal covariates, the theoretical curve no longer matches simulations, so the assumption marks the real boundary of the phenomenon rather than a technical convenience.
Reading between the lines
- If, as the authors conjecture, the same conic-geometry proof carries over to multinomial, Poisson, or log-linear models, the threshold method would become a general identifiability diagnostic for large categorical regressions; that extension is not proved here.
- The log-normal failure suggests that for radial distributions with indeterminate moment sequences, MLE existence may depend on the full radial law and not just its low moments; a natural experiment is to hold $\mathbb{E}R^2$ fixed and vary the log-normal scale to see whether the empirical phase transition moves with the moment-indeterminate part of the distribution.
- The stochastic approximation lemma is stated for a general convex loss bounded below, so the same phase-transition machinery may apply to other convex classification losses such as hinge or squared hinge; in that case the threshold would mark where the corresponding risk-minimizing classifier stops being well defined.
- Practically, one could use the gap between empirical MLE-existence frequencies and the formula (2.6) as a diagnostic for whether covariate projections stabilize, checking the moment condition (2.3) on data before trusting Gaussian-based asymptotics.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the existence of the maximum likelihood estimate in high-dimensional binary-response generalized linear models with elliptical covariates. Its main result, Theorem 2.7, states that under Assumptions 2.3–2.6 and the technical condition (2.7), the probability that the MLE exists tends to 0 when κ > h_MLE and to 1 when κ < h_MLE, where h_MLE is the minimum over λ0,λ1 of E(λ0 Y + λ1 X − Z)_+^2. This extends the Candès–Sur Gaussian logistic-regression phase transition to a general class of elliptical covariate distributions. The proof proceeds by translating non-existence into data separation, applying the approximate kinematic formula of Amelunxen et al., and proving a stochastic-approximation result (Theorem 4.3) for the normalized projection cost.
Significance. If the theorem is correct, the result is a meaningful advance: it shows that the phase-transition boundary depends on the covariate distribution only through the limiting law of a one-dimensional projection, and it exposes the log-normal case as a genuine failure of the projection-limit assumption. The paper is also useful for its checkable Carleman-type condition (2.3) and for introducing a stochastic-approximation framework that may apply to other problems. The simulations for gamma, Pareto, and half-normal covariates support the formula, and the authors are transparent about the log-normal counterexample. The proof, however, is not fully rigorous as written: the strong-convexity lemma and the bounded-minimizer lemma contain gaps, and the final derivation of Theorem 2.7 from the kinematic formula is sketched.
major comments (4)
- [Section 4.4, Lemma 4.5] The proof of Lemma 4.5 does not establish the claimed uniform strong convexity (4.11). The argument that certain expectations 'can be approximated by' simpler quantities for large λ0,λ1 is made without explicit error bounds, and it only considers λ0,λ1>0, leaving the remaining quadrants untreated. The uniform positive lower bound on ∇²G over all λ is therefore not proven. Since the bounded-minimizer proof in Lemma 4.6 invokes (4.11) with a fixed α0 for all λ, Theorem 4.3 is not rigorously established as written. Please supply explicit approximation bounds and cover all sign cases, or replace the strong-convexity assumption with a condition that can be verified.
- [Section 4.4, Lemma 4.6] In the union bound (4.15), the quantity σ² is written as sup_n Var[g(λ,ξ_n)] < ∞, but this cannot be a single constant independent of λ: on the circle |λ−λ0| = x, Var[g(λ,ξ_n)] grows like (1+|λ|)^4. The displayed Chebyshev bound is therefore not valid as written. The proof can likely be repaired by tracking the x-dependence of the variance and choosing d and y accordingly (e.g., using y = α0 x²/6 as already introduced), but the repair needs to be stated explicitly.
- [Section 4.1, proof of Theorem 2.7] The proof of Theorem 2.7 is only a sketch. The step from the approximate kinematic formula (4.3) to the asserted probability statements is not written out; in particular, the conditioning on (X,Y), the claim that E(Q_{p,n}|X,Y) converges in probability to h_MLE, and the treatment of the uncertainty band all require detailed arguments. The paper says the proof proceeds as in [12] with two modifications, but the two modifications are not fully spelled out. Please provide a complete derivation or a precise reference to the corresponding steps in [12].
- [Section 2.3, Theorem 2.7 and condition (2.7)] The hypothesis (2.7) is not derived from Assumptions 2.3–2.6, and the paper only gives a sufficient condition (2.8) that is checked for examples. As a result, the abstract's claim of a phase transition for 'a wide range' of elliptical covariate distributions goes beyond what is proven: Theorem 2.7 is conditional on a technical assumption that may fail for some elliptical families. Please either prove (2.7) under the stated assumptions, or state the condition explicitly in the abstract and title and discuss its scope.
minor comments (7)
- [Section 3] The phrase 'link functino' should be 'link function'.
- [Equation (4.9)] The last term on the right-hand side has unbalanced parentheses: it should be n E[p_-(X^(p))(G_{p,+}(X^(p)) + G_{p,-}(X^(p)))^{n-1}], with the exponent applied to the full sum.
- [Section 2.3, Theorem 2.7] 'supp E[(X^(p))^8]' should presumably be 'sup_p E[(X^(p))^8]'; please correct the notation.
- [Section 4.4, proof of Lemma 4.4] The strong law of large numbers for triangular arrays is invoked, but the stated fourth-moment condition typically gives convergence in probability, not almost-sure convergence. Either state a suitable theorem with all hypotheses or phrase the conclusion as convergence in probability.
- [Section 2.3, equation (2.5)] The same symbols G_{p,±} are defined for both the cumulative and tail integrals; this is confusing. Clarify by using different notation for the two sets of functions.
- [Section 4.4, Lemma 4.5 proof] The phrase 'consider λ1,λ2>0' appears to be a typo for 'λ0,λ1>0'.
- [Section 4.1] The approximate kinematic formula [2, Theorem I] is cited without checking that its hypotheses (e.g., closed convex cone, Gaussian subspace) are satisfied for C(W) and L; please verify or state the necessary conditions.
Circularity Check
No circularity: the phase-transition threshold is computed from the model distribution and checked against independent simulations, not fitted to the existence outcome.
full rationale
The paper's central claim is that the MLE-existence phase transition for binary-response GLMs with elliptical covariates is governed by h_MLE(α0,β0,γ0) = min_{λ0,λ1} E(λ0Y + λ1X − Z)_+^2, where (Y,X) is the projection limit of the covariate model under Assumption 2.6. This quantity is defined purely from the assumed covariate distribution, link function, and scaling parameters; it is not defined in terms of the probability that the MLE exists, and it is not fitted to MLE-existence data. The theorem then proves that the existence probability tends to 0 or 1 according as κ exceeds or falls below this computed value. The simulation section independently evaluates h_MLE by convex optimization and separately estimates existence frequencies by linear programming, so the comparison is an external check rather than a construction. The only citations with load-bearing content are to prior work by Candès and Sur, Montanari et al., and Amelunxen et al.; these are external results, not self-citations by the present authors, and the present paper extends rather than renames them. The paper also honestly identifies a genuine limitation: Assumption 2.6 fails for log-normal covariates, and the reported simulations show the theoretical curve does not match the empirical phase transition in that case; this is falsifiable behavior, not circularity. The proof of Theorem 4.3 contains gaps that are correctness risks rather than circularity: Lemma 4.5 asserts strong convexity using phrases like 'can be approximated by' without uniform error bounds, and Lemma 4.6 uses σ^2 := sup_n Var[g(λ,ξ_n)] as if it were uniform in λ on the circle. These are unproved technical steps, but they do not make the derivation equivalent to its inputs. No step reduces by definition or by fitted-parameter renaming to the claimed conclusion. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (8)
- domain assumption Assumption 2.6: the projection U^(p) converges in distribution to U.
- domain assumption Assumption 2.5: ER^2/p → α0^2 and |Σ^{1/2}β| → γ0/α0.
- domain assumption Assumption 2.3: σ and 1−σ are log-concave.
- domain assumption Assumption 2.4: full-rank elliptical covariate distribution, with r = p = rank(Σ) and FR absolutely continuous.
- ad hoc to paper Condition (2.7): E[p±(X^(p))(G_{p,∓}(X^(p)) + G_{p,±}(X^(p)))^{n−1}] = o(1/n).
- ad hoc to paper Moment bound sup_p E[(X^(p))^8] < ∞.
- standard math Approximate kinematic formula [2, Theorem I] and the statistical dimension identity [12, Lemma 3].
- standard math Epi-convergence and stochastic approximation results from [3, Theorem 2.3], [17, Proposition 3.3], [24, Theorem 2.1], and [8, Lemma 8.3].
Cite this review
Pith. "Pith review of The existence of maximum likelihood estimate in high-dimensional binary response generalized linear models." pith.science (2026). https://pith.science/paper/OYOMOKLB
@misc{pith2026190806208,
author = {Pith},
title = {Pith review of: The existence of maximum likelihood estimate in high-dimensional binary response generalized linear models},
year = {2026},
howpublished = {\url{https://pith.science/paper/OYOMOKLB}},
note = {Machine review of arXiv:1908.06208}
}
read the original abstract
Motivated by recent works on the high-dimensional logistic regression, we establish that the existence of the maximum likelihood estimate exhibits a phase transition for a wide range of generalized linear models with binary outcome and elliptical covariates. This extends a previous result of Cand\`es and Sur who proved the phase transition for the logistic regression with Gaussian covariates. Our result reveals a rich structure in the phase transition phenomenon, which is simply overlooked by Gaussianity. The main tools for deriving the result are data separation, convex geometry and stochastic approximation. We also conduct simulation studies to corroborate our theoretical findings, and explore other features of the problem.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[12]
Emmanuel J. Cand` es and Pragya Sur. The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression. Ann. Statist., 48(1):27–42, 2020
work page 2020
-
[1]
On the existence of maximum likelihood estimates in logistic regression models
Adelin Albert and John Anderson. On the existence of maximum likelihood estimates in logistic regression models. Biometrika, 71(1):1–10, 1984
work page 1984
-
[2]
Living on the edge: Phase transitions in convex programs with random data
Dennis Amelunxen, Martin Lotz, Michael McCoy, and Joel Tropp. Living on the edge: Phase transitions in convex programs with random data. Information and Inference: A Journal of the IMA , 3(3):224–294, 2014
work page 2014
-
[3]
Consistency of minimizers and the slln for stochastic programs
Zvi Artstein and Roger Wets. Consistency of minimizers and the slln for stochastic programs. Journal of Convex Analysis, 2(1/2):1–17, 1995
work page 1995
-
[4]
H´ edy Attouch.Variational convergence for functions and operators . Pitman, Boston, 1984
work page 1984
-
[5]
Approximation and convergence in nonlinear optimization
H´ edy Attouch and Roger Wets. Approximation and convergence in nonlinear optimization. InNonlinear Programming, pages 367–394. 1981
work page 1981
-
[6]
The dynamics of message passing on dense graphs, with applica- tions to compressed sensing
Mohsen Bayati and Andrea Montanari. The dynamics of message passing on dense graphs, with applica- tions to compressed sensing. IEEE Transactions on Information Theory , 57(2):764–785, 2011
work page 2011
-
[7]
Application of the logistic function to bio-assay
Joseph Berkson. Application of the logistic function to bio-assay. Journal of the American Statistical Association, 39(227):357–365, 1944
work page 1944
Show all 41 references
-
[8]
Some asymptotic theory for the bootstrap
Peter Bickel and David Freedman. Some asymptotic theory for the bootstrap. The annals of statistics , 9(6):1196–1217, 1981
1981
-
[9]
The calculation of the dosage-mortality curve
Chester Bliss. The calculation of the dosage-mortality curve. Annals of Applied Biology , 22(1):134–167, 1935
1935
-
[10]
The normal distribution: characterizations with applications , volume 100 of Lecture Notes in Statistics
Wlodzimierz Bryc. The normal distribution: characterizations with applications , volume 100 of Lecture Notes in Statistics . Springer-Verlag, New York, 1995
1995
-
[11]
On the theory of elliptically contoured distribu- tions
Stamatis Cambanis, Steel Huang, and Gordon Simons. On the theory of elliptically contoured distribu- tions. Journal of Multivariate Analysis , 11(3):368–385, 1981
1981
-
[13]
Measuring overlap in binary regression
Andreas Christmann and Peter Rousseeuw. Measuring overlap in binary regression. Computational Sta- tistics & Data Analysis , 37(1):65–75, 2001
2001
-
[14]
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
Thomas M Cover. Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition. IEEE transactions on electronic computers , (3):326–334, 1965
1965
-
[15]
Generalized maximum likelihood estimates for exponential families
Imre Csisz´ ar and Frantiˇ sek Mat´ uˇ s. Generalized maximum likelihood estimates for exponential families. Probability Theory and Related Fields , 141(1-2):213–246, 2008
2008
-
[16]
Computational aspects of probit model
Eugene Demidenko. Computational aspects of probit model. Mathematical Communications, 6(2):233– 247, 2001
2001
-
[17]
Asymptotic behavior of statistical estimators and of optimal solutions of stochastic optimization problems
Jitka Dupacov´ a and Roger Wets. Asymptotic behavior of statistical estimators and of optimal solutions of stochastic optimization problems. The Annals of Statistics , pages 1517–1549, 1988
1988
-
[18]
On the impact of predictor geometry on the performance on high-dimensional ridge-regularized generalized robust regression estimators
Noureddine El Karoui. On the impact of predictor geometry on the performance on high-dimensional ridge-regularized generalized robust regression estimators. Probability Theory and Related Fields , 170(1- 2):95–175, 2018
2018
-
[19]
Symmetric multivariate and related distributions
Kai-Tai Fang, Samuel Kotz, and Kai Wang Ng. Symmetric multivariate and related distributions . Chap- man & Hall, 1990
1990
-
[20]
Maximum likelihood estimation in log-linear models
Stephen Fienberg and Alessandro Rinaldo. Maximum likelihood estimation in log-linear models. The Annals of Statistics , 40(2):996–1023, 2012
2012
-
[21]
R. A. Fisher. On the mathematical foundations of theoretical statistics. Phil. Trans. R. Soc. , 222(594- 604):309–368, 1922
1922
-
[22]
The analysis of frequency data , volume 4
Shelby Haberman. The analysis of frequency data , volume 4. University of Chicago Press, 1977
1977
-
[23]
Epi-convergence of sequences of normal integrands and strong consistency of the maximum likelihood estimator
Christian Hess. Epi-convergence of sequences of normal integrands and strong consistency of the maximum likelihood estimator. The Annals of Statistics , 24(3):1298–1315, 1996
1996
-
[24]
Kanniappan and S
P. Kanniappan and S. Sastry. Uniform convergence of convex optimization problems. Journal of mathe- matical analysis and applications , 96(1):1–12, 1983
1983
-
[25]
Distribution theory of spherical distributions and a location-scale parameter generaliza- tion
Douglas Kelker. Distribution theory of spherical distributions and a location-scale parameter generaliza- tion. Sankhy¯ a, Series A, pages 419–430, 1970
1970
-
[26]
Stochastic estimation of the maximum of a regression function
Jack Kiefer and Jacob Wolfowitz. Stochastic estimation of the maximum of a regression function. The Annals of Mathematical Statistics , 23(3):462–466, 1952. EXISTENCE OF GLM MLE 23
1952
-
[27]
Epi-consistency of convex stochastic programs
Alan King and Roger Wets. Epi-consistency of convex stochastic programs. Stochastics and Stochastic Reports, 34(1-2):83–92, 1991
1991
-
[28]
Linear programming algorithms for detecting separated data in binary logistic regression models
Kjell Konis. Linear programming algorithms for detecting separated data in binary logistic regression models. PhD thesis, University of Oxford, 2007
2007
-
[29]
Existence and uniqueness of the maximum likelihood estimator for a multivariate probit model
Emmanuel Lesaffre and Heinz Kaufmann. Existence and uniqueness of the maximum likelihood estimator for a multivariate probit model. Journal of the American statistical Association , 87(419):805–811, 1992
1992
-
[30]
Recent developments on the moment problem
Gwo Dong Lin. Recent developments on the moment problem. Journal of Statistical Distributions and Applications, 4(1):5, 2017
2017
-
[31]
De Loera and T.A
J.A. De Loera and T.A. Hogan. Stochastic Tverberg theorems and their applications in multi-class logistic regression, data separability, and centerpoints of data. arXiv:11907.09698, 2019
2019 arXiv
-
[32]
Generalized linear models, volume 37
Peter McCullagh and John Nelder. Generalized linear models, volume 37. CRC Press, 1989
1989
-
[33]
The generalization error of max-margin linear classifiers: high-dimensional asymptotics in the overparametrized regime
Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan. The generalization error of max-margin linear classifiers: high-dimensional asymptotics in the overparametrized regime. arXiv:1911.01544, 2019
1911 arXiv
-
[34]
Generalized linear models
John Nelder and Robert Wedderburn. Generalized linear models. Journal of the Royal Statistical Society: Series A, 135(3):370–384, 1972
1972
-
[35]
A stochastic approximation method
Herbert Robbins and Sutton Monro. A stochastic approximation method. The Annals of Mathematical Statistics, 22(3):400–407, 1951
1951
-
[36]
A note on A
Thomas Santner and Diane Duffy. A note on A. Albert and J. A. Anderson’s conditions for the existence of maximum likelihood estimates in logistic regression models. Biometrika, 73(3):755–758, 1986
1986
-
[37]
Asymptotic analysis of stochastic programs
Alexander Shapiro. Asymptotic analysis of stochastic programs. Annals of Operations Research , 30(1):169–186, 1991
1991
-
[38]
On the existence of maximum likelihood estimators for the binomial response models
Mervyn Silvapulle. On the existence of maximum likelihood estimators for the binomial response models. Journal of the Royal Statistical Society. Series B (Methodological) , pages 310–313, 1981
1981
-
[39]
A modern maximum-likelihood theory for high-dimensional logistic regression
Pragya Sur and Emmanuel Cand` es. A modern maximum-likelihood theory for high-dimensional logistic regression. Proceedings of the National Academy of Sciences , 116(29):14516–14525, 2019
2019
-
[40]
The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled chi-square
Pragya Sur, Yuxin Chen, and Emmanuel Cand` es. The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled chi-square. arXiv:1706.01191, 2017. To appear in Probability Theory and Related Fields
2017 arXiv
-
[41]
Wedderburn
R. Wedderburn. On the existence and uniqueness of the maximum likelihood estimates for certain gen- eralized linear models. Biometrika, 63(1):27–32, 1976. Department of Industrial Engineering and Operations Research, Columbia University. Email address: wt2319@columbia.edu Divi...
1976
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.