Pith. sign in

REVIEW 4 major objections 7 minor 41 references

The existence of maximum likelihood estimate in high-dimensional binary response generalized linear models

T0 review · 4 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read In high-dimensional binary regression with elliptical covariates, the maximum likelihood estimate exists with probability tending to one below an explicit threshold $h_{\mathrm{MLE}}$ and with probability tending to zero above it.

desk verdict A plausible extension of the Candès–Sur phase transition to elliptical covariates, but the proof of the key convergence theorem has gaps that need fixing before the result is fully rigorous. read the letter →

arxiv 1908.06208 v2 pith:OYOMOKLB submitted 2019-08-17 math.ST stat.TH

classification math.STstat.TH MSC 62F1262J1262H05
keywords maximumlikelihoodestimationphasetransitionbinaryresponsegeneralizedlinearmodelsellipticaldistributionshigh-dimensionalasymptoticsdataseparationconvexgeometrystochasticapproximation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes that, for binary-response generalized linear models with elliptical covariates in high dimensions, whether the maximum likelihood estimate exists is governed by a sharp phase transition in the ratio $\kappa = p/n$. When $\kappa$ lies below an explicit threshold $h_{\mathrm{MLE}}$, the MLE exists with probability tending to one; above it, the data are separable and the MLE fails with probability tending to one. The threshold depends on the link function, the intercept, and the limiting scaling of the regression signal, and it can be computed by solving a small convex optimization problem. This matters because it tells practitioners, before fitting, when a model can be fitted at all, and it shows that the Gaussian assumption in earlier work can be relaxed to a broad class of elliptical distributions. The paper also identifies the boundary of its own result: the phase transition breaks down when the one-dimensional projections of the covariates do not stabilize in distribution, as with log-normal covariates.

What carries the argument

The key object is the threshold $h_{\mathrm{MLE}}$, the minimal expected squared positive part of $\lambda_0 Y+\lambda_1 X-Z$ over $(\lambda_0,\lambda_1)$. It does the heavy lifting through conic geometry: for log-concave links the MLE fails exactly when the data are separable, and for spherical covariates separability is equivalent to a random $(p-1)$-dimensional subspace $L$ (spanned by the nuisance coordinates) intersecting a cone $C(W)$ built from the labels and the signal coordinate. The approximate kinematic formula of convex geometry says such an intersection becomes overwhelmingly likely when $p-1+\delta(C(W))$ exceeds $n$, where the statistical dimension $\delta(C(W))$ is $n$ minus the expected squared distance from a standard normal vector to $C(W)$; the limit of that distance is $n\,h_{\mathrm{MLE}}(\alpha_0,\beta_0,\gamma_0)$. A stochastic approximation lemma, proved by epi-convergence, shows that the empirical minimization defining the threshold converges to its expectation, and condition (2.7) bounds the probability of separation using only the signal coordinate. Assumption 2.6, that the projections $U^{(p)}$ converge in distribution, makes the limit $(Y,X)$ well defined; the moment criterion (2.3) gives a checkable sufficient condition.

What would settle it

Run the linear-programming separation test (3.1) for n=1000, p=$\kappa n$, with spherical covariates whose radial component is Gamma(1,$\theta_p$) scaled so $\mathbb{E}R^2=p+1$, logit link, $\beta_0=0$, across a grid of $\gamma_0$ and $\kappa$, and compare the empirical 50% existence boundary with $h_{\mathrm{MLE}}$ from (2.6); the theorem predicts agreement within the uncertainty band, so a systematic gap would falsify it. The log-normal design, where the projection-limit assumption fails, is the built-in negative control: there the gap is real.

Watch

Extended reading notes

Core claim

The central claim is Theorem 2.7. After rotating the covariates so all signal lies in one coordinate, write $(Y^{(p)},X^{(p)})=(V^{(p)},V^{(p)}U^{(p)})$, where $U^{(p)}$ is a single coordinate of the standardized elliptical covariate vector and $P(V^{(p)}=1\,|\,U^{(p)})=\sigma(\beta_0+(\gamma_0/\alpha_0)U^{(p)})$. Assume the link satisfies log-concavity of $\sigma$ and $1-\sigma$, the covariates are full-rank elliptical with $\mathbb{E}R^2/p\to\alpha_0^2$, the coefficient scaling obeys $|\Sigma^{1/2}\beta|\to\gamma_0/\alpha_0$, and the projections converge in distribution, $U^{(p)}\Rightarrow U$. Let $Z\sim N(0,1)$ be independent and define $h_{\mathrm{MLE}}(\alpha_0,\beta_0,\gamma_0)=\min_{\lambda_0,\lambda_1}\mathbb{E}(\lambda_0 Y+\lambda_1 X-Z)_+^2$. Then $\kappa=p/n>h_{\mathrm{MLE}}$ implies the maximum likelihood estimate exists with probability tending to $0$, while $\kappa<h_{\mathrm{MLE}}$ implies it exists with probability tending to $1$, subject to the finite-eighth-moment condition and condition (2.7). This is the same phase transition known for Gaussian logistic regression, now extended to any elliptical family whose one-dimensional projections stabilize, with the boundary computed from the limiting projection distribution.

Load-bearing premise

The load-bearing premise is that the one-dimensional projections of the standardized covariates converge in distribution to a fixed limit, because the threshold $h_{\mathrm{MLE}}$ is computed from that limiting distribution; when it fails, the paper's own simulations show the formula does not hold.

Editorial extensions

If this is right

  • Below the threshold $\kappa<h_{\mathrm{MLE}}$, a practitioner can expect the MLE to exist with probability tending to one; above it, data separation occurs almost surely and standard fitting algorithms have no finite solution to find.
  • The threshold can be evaluated in advance by solving a two-dimensional convex program for the limiting projection distribution, so the result supplies a practical pre-fit diagnostic.
  • The phase transition holds across logit, probit, and cloglog links, and across Gaussian, Gamma, Pareto (with enough moments), and half-normal elliptical covariates, because these satisfy the projection-limit and univariate-separation conditions.
  • When the projection-limit assumption fails, as for log-normal covariates, the theoretical curve no longer matches simulations, so the assumption marks the real boundary of the phenomenon rather than a technical convenience.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If, as the authors conjecture, the same conic-geometry proof carries over to multinomial, Poisson, or log-linear models, the threshold method would become a general identifiability diagnostic for large categorical regressions; that extension is not proved here.
  • The log-normal failure suggests that for radial distributions with indeterminate moment sequences, MLE existence may depend on the full radial law and not just its low moments; a natural experiment is to hold $\mathbb{E}R^2$ fixed and vary the log-normal scale to see whether the empirical phase transition moves with the moment-indeterminate part of the distribution.
  • The stochastic approximation lemma is stated for a general convex loss bounded below, so the same phase-transition machinery may apply to other convex classification losses such as hinge or squared hinge; in that case the threshold would mark where the corresponding risk-minimizing classifier stops being well defined.
  • Practically, one could use the gap between empirical MLE-existence frequencies and the formula (2.6) as a diagnostic for whether covariate projections stabilize, checking the moment condition (2.3) on data before trusting Gaussian-based asymptotics.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper studies the existence of the maximum likelihood estimate in high-dimensional binary-response generalized linear models with elliptical covariates. Its main result, Theorem 2.7, states that under Assumptions 2.3–2.6 and the technical condition (2.7), the probability that the MLE exists tends to 0 when κ > h_MLE and to 1 when κ < h_MLE, where h_MLE is the minimum over λ0,λ1 of E(λ0 Y + λ1 X − Z)_+^2. This extends the Candès–Sur Gaussian logistic-regression phase transition to a general class of elliptical covariate distributions. The proof proceeds by translating non-existence into data separation, applying the approximate kinematic formula of Amelunxen et al., and proving a stochastic-approximation result (Theorem 4.3) for the normalized projection cost.

Significance. If the theorem is correct, the result is a meaningful advance: it shows that the phase-transition boundary depends on the covariate distribution only through the limiting law of a one-dimensional projection, and it exposes the log-normal case as a genuine failure of the projection-limit assumption. The paper is also useful for its checkable Carleman-type condition (2.3) and for introducing a stochastic-approximation framework that may apply to other problems. The simulations for gamma, Pareto, and half-normal covariates support the formula, and the authors are transparent about the log-normal counterexample. The proof, however, is not fully rigorous as written: the strong-convexity lemma and the bounded-minimizer lemma contain gaps, and the final derivation of Theorem 2.7 from the kinematic formula is sketched.

major comments (4)
  1. [Section 4.4, Lemma 4.5] The proof of Lemma 4.5 does not establish the claimed uniform strong convexity (4.11). The argument that certain expectations 'can be approximated by' simpler quantities for large λ0,λ1 is made without explicit error bounds, and it only considers λ0,λ1>0, leaving the remaining quadrants untreated. The uniform positive lower bound on ∇²G over all λ is therefore not proven. Since the bounded-minimizer proof in Lemma 4.6 invokes (4.11) with a fixed α0 for all λ, Theorem 4.3 is not rigorously established as written. Please supply explicit approximation bounds and cover all sign cases, or replace the strong-convexity assumption with a condition that can be verified.
  2. [Section 4.4, Lemma 4.6] In the union bound (4.15), the quantity σ² is written as sup_n Var[g(λ,ξ_n)] < ∞, but this cannot be a single constant independent of λ: on the circle |λ−λ0| = x, Var[g(λ,ξ_n)] grows like (1+|λ|)^4. The displayed Chebyshev bound is therefore not valid as written. The proof can likely be repaired by tracking the x-dependence of the variance and choosing d and y accordingly (e.g., using y = α0 x²/6 as already introduced), but the repair needs to be stated explicitly.
  3. [Section 4.1, proof of Theorem 2.7] The proof of Theorem 2.7 is only a sketch. The step from the approximate kinematic formula (4.3) to the asserted probability statements is not written out; in particular, the conditioning on (X,Y), the claim that E(Q_{p,n}|X,Y) converges in probability to h_MLE, and the treatment of the uncertainty band all require detailed arguments. The paper says the proof proceeds as in [12] with two modifications, but the two modifications are not fully spelled out. Please provide a complete derivation or a precise reference to the corresponding steps in [12].
  4. [Section 2.3, Theorem 2.7 and condition (2.7)] The hypothesis (2.7) is not derived from Assumptions 2.3–2.6, and the paper only gives a sufficient condition (2.8) that is checked for examples. As a result, the abstract's claim of a phase transition for 'a wide range' of elliptical covariate distributions goes beyond what is proven: Theorem 2.7 is conditional on a technical assumption that may fail for some elliptical families. Please either prove (2.7) under the stated assumptions, or state the condition explicitly in the abstract and title and discuss its scope.
minor comments (7)
  1. [Section 3] The phrase 'link functino' should be 'link function'.
  2. [Equation (4.9)] The last term on the right-hand side has unbalanced parentheses: it should be n E[p_-(X^(p))(G_{p,+}(X^(p)) + G_{p,-}(X^(p)))^{n-1}], with the exponent applied to the full sum.
  3. [Section 2.3, Theorem 2.7] 'supp E[(X^(p))^8]' should presumably be 'sup_p E[(X^(p))^8]'; please correct the notation.
  4. [Section 4.4, proof of Lemma 4.4] The strong law of large numbers for triangular arrays is invoked, but the stated fourth-moment condition typically gives convergence in probability, not almost-sure convergence. Either state a suitable theorem with all hypotheses or phrase the conclusion as convergence in probability.
  5. [Section 2.3, equation (2.5)] The same symbols G_{p,±} are defined for both the cumulative and tail integrals; this is confusing. Clarify by using different notation for the two sets of functions.
  6. [Section 4.4, Lemma 4.5 proof] The phrase 'consider λ1,λ2>0' appears to be a typo for 'λ0,λ1>0'.
  7. [Section 4.1] The approximate kinematic formula [2, Theorem I] is cited without checking that its hypotheses (e.g., closed convex cone, Gaussian subspace) are satisfied for C(W) and L; please verify or state the necessary conditions.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the phase-transition threshold is computed from the model distribution and checked against independent simulations, not fitted to the existence outcome.

full rationale

The paper's central claim is that the MLE-existence phase transition for binary-response GLMs with elliptical covariates is governed by h_MLE(α0,β0,γ0) = min_{λ0,λ1} E(λ0Y + λ1X − Z)_+^2, where (Y,X) is the projection limit of the covariate model under Assumption 2.6. This quantity is defined purely from the assumed covariate distribution, link function, and scaling parameters; it is not defined in terms of the probability that the MLE exists, and it is not fitted to MLE-existence data. The theorem then proves that the existence probability tends to 0 or 1 according as κ exceeds or falls below this computed value. The simulation section independently evaluates h_MLE by convex optimization and separately estimates existence frequencies by linear programming, so the comparison is an external check rather than a construction. The only citations with load-bearing content are to prior work by Candès and Sur, Montanari et al., and Amelunxen et al.; these are external results, not self-citations by the present authors, and the present paper extends rather than renames them. The paper also honestly identifies a genuine limitation: Assumption 2.6 fails for log-normal covariates, and the reported simulations show the theoretical curve does not match the empirical phase transition in that case; this is falsifiable behavior, not circularity. The proof of Theorem 4.3 contains gaps that are correctness risks rather than circularity: Lemma 4.5 asserts strong convexity using phrases like 'can be approximated by' without uniform error bounds, and Lemma 4.6 uses σ^2 := sup_n Var[g(λ,ξ_n)] as if it were uniform in λ on the circle. These are unproved technical steps, but they do not make the derivation equivalent to its inputs. No step reduces by definition or by fitted-parameter renaming to the claimed conclusion. Therefore the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 8 assumptions · 0 invented entities

The theorem rests on model assumptions about the link function, the covariate distribution, parameter scaling, the projection-limit condition, a technical tail condition (2.7), and a moment bound. No new physical or mathematical entities are introduced, and no parameter is fitted to the existence data. The largest burden is Assumption 2.6, which is not automatically satisfied by all elliptical distributions.

assumptions (8)
  • domain assumption Assumption 2.6: the projection U^(p) converges in distribution to U.
    Required to define the limit distribution F_{α0,β0,γ0} and the threshold h_MLE; without it the phase transition can fail, as shown by the log-normal example in Section 3.3.
  • domain assumption Assumption 2.5: ER^2/p → α0^2 and |Σ^{1/2}β| → γ0/α0.
    Parameter scaling needed for a meaningful high-dimensional limit; it makes Var(x^T β) converge to γ0^2.
  • domain assumption Assumption 2.3: σ and 1−σ are log-concave.
    Used so that the MLE exists if and only if the data overlap, via Theorem 2.1. Logit, probit, and cloglog links satisfy it.
  • domain assumption Assumption 2.4: full-rank elliptical covariate distribution, with r = p = rank(Σ) and FR absolutely continuous.
    Excludes singular or discrete cases and ensures the geometric separation characterization applies.
  • ad hoc to paper Condition (2.7): E[p±(X^(p))(G_{p,∓}(X^(p)) + G_{p,±}(X^(p)))^{n−1}] = o(1/n).
    Technical condition introduced to prove that univariate separation is rare, P(No MLE Single) = o(1). A sufficient condition is (2.8), which is assumed for the examples.
  • ad hoc to paper Moment bound sup_p E[(X^(p))^8] < ∞.
    Used for the law of large numbers for triangular arrays in Theorem 4.3. The authors state it is purely technical and that simulations suggest it can be relaxed.
  • standard math Approximate kinematic formula [2, Theorem I] and the statistical dimension identity [12, Lemma 3].
    Used in Section 4.1 to convert the intersection probability P(L ∩ C(W) ≠ {0}) into a comparison between p/n and the statistical dimension δ(C(W)).
  • standard math Epi-convergence and stochastic approximation results from [3, Theorem 2.3], [17, Proposition 3.3], [24, Theorem 2.1], and [8, Lemma 8.3].
    Used in Lemma 4.4 and Theorem 4.3 to establish convergence of the stochastic optimization values for triangular arrays.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The existence of maximum likelihood estimate in high-dimensional binary response generalized linear models." pith.science (2026). https://pith.science/paper/OYOMOKLB

@misc{pith2026190806208,
  author       = {Pith},
  title        = {Pith review of: The existence of maximum likelihood estimate in high-dimensional binary response generalized linear models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OYOMOKLB}},
  note         = {Machine review of arXiv:1908.06208}
}
read the original abstract

Motivated by recent works on the high-dimensional logistic regression, we establish that the existence of the maximum likelihood estimate exhibits a phase transition for a wide range of generalized linear models with binary outcome and elliptical covariates. This extends a previous result of Cand\`es and Sur who proved the phase transition for the logistic regression with Gaussian covariates. Our result reveals a rich structure in the phase transition phenomenon, which is simply overlooked by Gaussianity. The main tools for deriving the result are data separation, convex geometry and stochastic approximation. We also conduct simulation studies to corroborate our theoretical findings, and explore other features of the problem.

Figures

Figures reproduced from arXiv: 1908.06208 by the authors.

Figure 1
Figure 1. Phase transition for the existence of the maximum likelihood es￾timate for R ∼ chi distribution with degree of freedom p. The red curve is the theoretical hMLE boundary given by (2.6). distributions, and the resulting theoretical phase transition curves agree with the simulations. More precisely, assume R ∼ Gamma(k, θ) where k is the shape parameter and θ is the scale parameter. The second moment condition (3.2) giv… view at source ↗
Figure 2
Figure 2. Convergence of hMLE by (2.6) for Gamma distributions. hˆ (0) MLE is computed by taking the average of hMLE for κ ≥ 0.3. 3.3. The moment condition and the tail behavior of R. First we explore a case where the eighth moment of X(p) does not exist as required by Theorem 2.7. To this end, we sample R from the Pareto distribution of type I. We specify the shape parameter α, and set the scale parameter xm = p (α − 2)/α · … view at source ↗
Figure 3
Figure 3. Phase transition of the MLE existence for R ∼ chi distribution (df = p) with different link functions. Upper: The value of each grid in the heatmap is the proportion of times that the MLE exists across the 100 replicates. The red curve is the theoretical hMLE boundary given by (2.6). Middle: The same setup as the upper figures, but A is generated so that each entry Aij is sampled from the standard Gaussian distribut… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Phase transition of the MLE existence for R ∼ Gamma distribu￾tions with different parameters and the logit link (see Section 3.2). Upper: The value of each grid in the heatmap is the proportion of times that the MLE exists across the 100 replicates. The red curve is th…
Figure 5
Figure 5. Figure 5: Phase transition of the MLE existence for the Pareto distribu￾tions with different parameters and the logit link (see Section 3.3). Upper: The value of each grid in the heatmap is the proportion of times that the MLE exists across the 100 replicates. The red curve is t…
Figure 6
Figure 6. Figure 6: Phase transition of the MLE existence for the half-normal dis￾tribution and the logit link function. Left: The value of each grid in the heatmap is the proportion of times that the MLE exists across the 100 repli￾cates. The red curve is the theoretical hMLE boundary gi…
Figure 7
Figure 7. Figure 7: Phase transition of the MLE existence for the log-Normal distribu￾tion and the logit link function. Left: The value of each grid in the heatmap is the proportion of times that the MLE exists across the 100 replicates. The red curve is the theoretical hMLE boundary give…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

41 extracted references · 40 canonical work pages

  1. [12]

    Cand` es and Pragya Sur

    Emmanuel J. Cand` es and Pragya Sur. The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression. Ann. Statist., 48(1):27–42, 2020

  2. [1]

    On the existence of maximum likelihood estimates in logistic regression models

    Adelin Albert and John Anderson. On the existence of maximum likelihood estimates in logistic regression models. Biometrika, 71(1):1–10, 1984

  3. [2]

    Living on the edge: Phase transitions in convex programs with random data

    Dennis Amelunxen, Martin Lotz, Michael McCoy, and Joel Tropp. Living on the edge: Phase transitions in convex programs with random data. Information and Inference: A Journal of the IMA , 3(3):224–294, 2014

  4. [3]

    Consistency of minimizers and the slln for stochastic programs

    Zvi Artstein and Roger Wets. Consistency of minimizers and the slln for stochastic programs. Journal of Convex Analysis, 2(1/2):1–17, 1995

  5. [4]

    Pitman, Boston, 1984

    H´ edy Attouch.Variational convergence for functions and operators . Pitman, Boston, 1984

  6. [5]

    Approximation and convergence in nonlinear optimization

    H´ edy Attouch and Roger Wets. Approximation and convergence in nonlinear optimization. InNonlinear Programming, pages 367–394. 1981

  7. [6]

    The dynamics of message passing on dense graphs, with applica- tions to compressed sensing

    Mohsen Bayati and Andrea Montanari. The dynamics of message passing on dense graphs, with applica- tions to compressed sensing. IEEE Transactions on Information Theory , 57(2):764–785, 2011

  8. [7]

    Application of the logistic function to bio-assay

    Joseph Berkson. Application of the logistic function to bio-assay. Journal of the American Statistical Association, 39(227):357–365, 1944

Show all 41 references
  1. [8]

    Some asymptotic theory for the bootstrap

    Peter Bickel and David Freedman. Some asymptotic theory for the bootstrap. The annals of statistics , 9(6):1196–1217, 1981

  2. [9]

    The calculation of the dosage-mortality curve

    Chester Bliss. The calculation of the dosage-mortality curve. Annals of Applied Biology , 22(1):134–167, 1935

  3. [10]

    The normal distribution: characterizations with applications , volume 100 of Lecture Notes in Statistics

    Wlodzimierz Bryc. The normal distribution: characterizations with applications , volume 100 of Lecture Notes in Statistics . Springer-Verlag, New York, 1995

  4. [11]

    On the theory of elliptically contoured distribu- tions

    Stamatis Cambanis, Steel Huang, and Gordon Simons. On the theory of elliptically contoured distribu- tions. Journal of Multivariate Analysis , 11(3):368–385, 1981

  5. [13]

    Measuring overlap in binary regression

    Andreas Christmann and Peter Rousseeuw. Measuring overlap in binary regression. Computational Sta- tistics & Data Analysis , 37(1):65–75, 2001

  6. [14]

    Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition

    Thomas M Cover. Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition. IEEE transactions on electronic computers , (3):326–334, 1965

  7. [15]

    Generalized maximum likelihood estimates for exponential families

    Imre Csisz´ ar and Frantiˇ sek Mat´ uˇ s. Generalized maximum likelihood estimates for exponential families. Probability Theory and Related Fields , 141(1-2):213–246, 2008

  8. [16]

    Computational aspects of probit model

    Eugene Demidenko. Computational aspects of probit model. Mathematical Communications, 6(2):233– 247, 2001

  9. [17]

    Asymptotic behavior of statistical estimators and of optimal solutions of stochastic optimization problems

    Jitka Dupacov´ a and Roger Wets. Asymptotic behavior of statistical estimators and of optimal solutions of stochastic optimization problems. The Annals of Statistics , pages 1517–1549, 1988

  10. [18]

    On the impact of predictor geometry on the performance on high-dimensional ridge-regularized generalized robust regression estimators

    Noureddine El Karoui. On the impact of predictor geometry on the performance on high-dimensional ridge-regularized generalized robust regression estimators. Probability Theory and Related Fields , 170(1- 2):95–175, 2018

  11. [19]

    Symmetric multivariate and related distributions

    Kai-Tai Fang, Samuel Kotz, and Kai Wang Ng. Symmetric multivariate and related distributions . Chap- man & Hall, 1990

  12. [20]

    Maximum likelihood estimation in log-linear models

    Stephen Fienberg and Alessandro Rinaldo. Maximum likelihood estimation in log-linear models. The Annals of Statistics , 40(2):996–1023, 2012

  13. [21]

    R. A. Fisher. On the mathematical foundations of theoretical statistics. Phil. Trans. R. Soc. , 222(594- 604):309–368, 1922

  14. [22]

    The analysis of frequency data , volume 4

    Shelby Haberman. The analysis of frequency data , volume 4. University of Chicago Press, 1977

  15. [23]

    Epi-convergence of sequences of normal integrands and strong consistency of the maximum likelihood estimator

    Christian Hess. Epi-convergence of sequences of normal integrands and strong consistency of the maximum likelihood estimator. The Annals of Statistics , 24(3):1298–1315, 1996

  16. [24]

    Kanniappan and S

    P. Kanniappan and S. Sastry. Uniform convergence of convex optimization problems. Journal of mathe- matical analysis and applications , 96(1):1–12, 1983

  17. [25]

    Distribution theory of spherical distributions and a location-scale parameter generaliza- tion

    Douglas Kelker. Distribution theory of spherical distributions and a location-scale parameter generaliza- tion. Sankhy¯ a, Series A, pages 419–430, 1970

  18. [26]

    Stochastic estimation of the maximum of a regression function

    Jack Kiefer and Jacob Wolfowitz. Stochastic estimation of the maximum of a regression function. The Annals of Mathematical Statistics , 23(3):462–466, 1952. EXISTENCE OF GLM MLE 23

  19. [27]

    Epi-consistency of convex stochastic programs

    Alan King and Roger Wets. Epi-consistency of convex stochastic programs. Stochastics and Stochastic Reports, 34(1-2):83–92, 1991

  20. [28]

    Linear programming algorithms for detecting separated data in binary logistic regression models

    Kjell Konis. Linear programming algorithms for detecting separated data in binary logistic regression models. PhD thesis, University of Oxford, 2007

  21. [29]

    Existence and uniqueness of the maximum likelihood estimator for a multivariate probit model

    Emmanuel Lesaffre and Heinz Kaufmann. Existence and uniqueness of the maximum likelihood estimator for a multivariate probit model. Journal of the American statistical Association , 87(419):805–811, 1992

  22. [30]

    Recent developments on the moment problem

    Gwo Dong Lin. Recent developments on the moment problem. Journal of Statistical Distributions and Applications, 4(1):5, 2017

  23. [31]

    De Loera and T.A

    J.A. De Loera and T.A. Hogan. Stochastic Tverberg theorems and their applications in multi-class logistic regression, data separability, and centerpoints of data. arXiv:11907.09698, 2019

  24. [32]

    Generalized linear models, volume 37

    Peter McCullagh and John Nelder. Generalized linear models, volume 37. CRC Press, 1989

  25. [33]

    The generalization error of max-margin linear classifiers: high-dimensional asymptotics in the overparametrized regime

    Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan. The generalization error of max-margin linear classifiers: high-dimensional asymptotics in the overparametrized regime. arXiv:1911.01544, 2019

  26. [34]

    Generalized linear models

    John Nelder and Robert Wedderburn. Generalized linear models. Journal of the Royal Statistical Society: Series A, 135(3):370–384, 1972

  27. [35]

    A stochastic approximation method

    Herbert Robbins and Sutton Monro. A stochastic approximation method. The Annals of Mathematical Statistics, 22(3):400–407, 1951

  28. [36]

    A note on A

    Thomas Santner and Diane Duffy. A note on A. Albert and J. A. Anderson’s conditions for the existence of maximum likelihood estimates in logistic regression models. Biometrika, 73(3):755–758, 1986

  29. [37]

    Asymptotic analysis of stochastic programs

    Alexander Shapiro. Asymptotic analysis of stochastic programs. Annals of Operations Research , 30(1):169–186, 1991

  30. [38]

    On the existence of maximum likelihood estimators for the binomial response models

    Mervyn Silvapulle. On the existence of maximum likelihood estimators for the binomial response models. Journal of the Royal Statistical Society. Series B (Methodological) , pages 310–313, 1981

  31. [39]

    A modern maximum-likelihood theory for high-dimensional logistic regression

    Pragya Sur and Emmanuel Cand` es. A modern maximum-likelihood theory for high-dimensional logistic regression. Proceedings of the National Academy of Sciences , 116(29):14516–14525, 2019

  32. [40]

    The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled chi-square

    Pragya Sur, Yuxin Chen, and Emmanuel Cand` es. The likelihood ratio test in high-dimensional logistic regression is asymptotically a rescaled chi-square. arXiv:1706.01191, 2017. To appear in Probability Theory and Related Fields

  33. [41]

    Wedderburn

    R. Wedderburn. On the existence and uniqueness of the maximum likelihood estimates for certain gen- eralized linear models. Biometrika, 63(1):27–32, 1976. Department of Industrial Engineering and Operations Research, Columbia University. Email address: wt2319@columbia.edu Divi...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.