{"id":"03504496-e779-4e97-9f0a-756e4f59ef56","arxiv_id":"1908.06208","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Under a projection-limit assumption, the existence of the MLE in binary-response generalized linear models with elliptical covariates changes sharply at a threshold p/n that generalizes the Gaussian formula of Candès and Sur.","lead":"This paper proves a phase transition for whether the maximum likelihood estimate exists in high-dimensional binary classification models with elliptical, non-Gaussian covariates. It extends the Gaussian case to a broad class, with a threshold that depends on model parameters and on a projection-limit distribution.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The phase transition rests on an unproved strong-convexity step: Lemma 4.5 is only sketched and Lemma 4.6's variance bound is not uniform, so Theorem 4.3 is not rigorously established as written.","rationale":"","tokens_in":17701,"tokens_out":34011,"duration_ms":342961,"concrete_test":"Re-derive Lemma 4.5 rigorously: for a fixed direction u=(u0,u1) and t=|lambda| -> infinity, use dominated convergence on the two Hessian terms in (4.12) to obtain explicit positive-definite limits, then verify liminf_t lambda_min(grad^2 G(tu)) > 0 uniformly in u. A numerical cross-check: evaluate lambda_min(grad^2 G(lambda)) for the logit link and a non-degenerate U satisfying Assumption 2.6 on a grid of lambda up to |lambda|=1e4 and along rays lambda0/lambda1 in {-infinity,...,infinity}; if any eigenvalue tends to 0 while the limiting matrices are positive definite, the sketched proof has a missing step. Also redo Lemma 4.6 replacing the constant sigma^2 with Var[g(lambda, xi_n)] <= C(1+|lambda|^4), and check that n y^2 = n alpha0^2 x^4/36 absorbs the x^4 growth so the Chebyshev probabilities are o(1).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proof of Theorem 2.7 goes through Theorem 4.3, which asserts that the empirical min Q_{p,n} converges in probability to h_MLE. To apply the stochastic-approximation Lemma 4.4, the paper must show (i) the limit objective is strongly convex and (ii) the empirical minimizers are bounded in probability. Lemma 4.5 attempts (i) but only establishes pointwise strict convexity (det>0) and then asserts asymptotic approximations of the Hessian for large lambda0,lambda1 with phrases like 'can be approximated by' and 'by strict convexity and the approximation'. No explicit error bounds are given, and the argument does not cover all directions lambda/lambda1 or negative lambda. Lemma 4.6 attempts (ii) but its union-bound step (4.15) uses sigma^2 := sup_n Var[g(lambda, xi_n)] < infinity as if it were uniform in lambda; on the circle |lambda - lambda0| = x the variance can grow like x^4. The bound can likely be repaired because y = alpha0 x^2/6 makes Chebyshev still give o(1), but the written proof does not state this. If either the uniform positive-definiteness in Lemma 4.5 or the bounded-minimizer conclusion in Lemma 4.6 fails, the convergence of Q_{p,n} to h_MLE is not justified and the phase transition is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the existence of the maximum likelihood estimate in high-dimensional binary-response generalized linear models with elliptical covariates. Its main result, Theorem 2.7, states that under Assumptions 2.3–2.6 and the technical condition (2.7), the probability that the MLE exists tends to 0 when κ > h_MLE and to 1 when κ < h_MLE, where h_MLE is the minimum over λ0,λ1 of E(λ0 Y + λ1 X − Z)_+^2. This extends the Candès–Sur Gaussian logistic-regression phase transition to a general class of elliptical covariate distributions. The proof proceeds by translating non-existence into data separation, applying the approximate kinematic formula of Amelunxen et al., and proving a stochastic-approximation result (Theorem 4.3) for the normalized projection cost.","tokens_in":18015,"tokens_out":10236,"duration_ms":89317,"significance":"If the theorem is correct, the result is a meaningful advance: it shows that the phase-transition boundary depends on the covariate distribution only through the limiting law of a one-dimensional projection, and it exposes the log-normal case as a genuine failure of the projection-limit assumption. The paper is also useful for its checkable Carleman-type condition (2.3) and for introducing a stochastic-approximation framework that may apply to other problems. The simulations for gamma, Pareto, and half-normal covariates support the formula, and the authors are transparent about the log-normal counterexample. The proof, however, is not fully rigorous as written: the strong-convexity lemma and the bounded-minimizer lemma contain gaps, and the final derivation of Theorem 2.7 from the kinematic formula is sketched.","major_comments":[{"comment":"The proof of Lemma 4.5 does not establish the claimed uniform strong convexity (4.11). The argument that certain expectations 'can be approximated by' simpler quantities for large λ0,λ1 is made without explicit error bounds, and it only considers λ0,λ1>0, leaving the remaining quadrants untreated. The uniform positive lower bound on ∇²G over all λ is therefore not proven. Since the bounded-minimizer proof in Lemma 4.6 invokes (4.11) with a fixed α0 for all λ, Theorem 4.3 is not rigorously established as written. Please supply explicit approximation bounds and cover all sign cases, or replace the strong-convexity assumption with a condition that can be verified.","section":"Section 4.4, Lemma 4.5"},{"comment":"In the union bound (4.15), the quantity σ² is written as sup_n Var[g(λ,ξ_n)] < ∞, but this cannot be a single constant independent of λ: on the circle |λ−λ0| = x, Var[g(λ,ξ_n)] grows like (1+|λ|)^4. The displayed Chebyshev bound is therefore not valid as written. The proof can likely be repaired by tracking the x-dependence of the variance and choosing d and y accordingly (e.g., using y = α0 x²/6 as already introduced), but the repair needs to be stated explicitly.","section":"Section 4.4, Lemma 4.6"},{"comment":"The proof of Theorem 2.7 is only a sketch. The step from the approximate kinematic formula (4.3) to the asserted probability statements is not written out; in particular, the conditioning on (X,Y), the claim that E(Q_{p,n}|X,Y) converges in probability to h_MLE, and the treatment of the uncertainty band all require detailed arguments. The paper says the proof proceeds as in [12] with two modifications, but the two modifications are not fully spelled out. Please provide a complete derivation or a precise reference to the corresponding steps in [12].","section":"Section 4.1, proof of Theorem 2.7"},{"comment":"The hypothesis (2.7) is not derived from Assumptions 2.3–2.6, and the paper only gives a sufficient condition (2.8) that is checked for examples. As a result, the abstract's claim of a phase transition for 'a wide range' of elliptical covariate distributions goes beyond what is proven: Theorem 2.7 is conditional on a technical assumption that may fail for some elliptical families. Please either prove (2.7) under the stated assumptions, or state the condition explicitly in the abstract and title and discuss its scope.","section":"Section 2.3, Theorem 2.7 and condition (2.7)"}],"minor_comments":[{"comment":"The phrase 'link functino' should be 'link function'.","section":"Section 3"},{"comment":"The last term on the right-hand side has unbalanced parentheses: it should be n E[p_-(X^(p))(G_{p,+}(X^(p)) + G_{p,-}(X^(p)))^{n-1}], with the exponent applied to the full sum.","section":"Equation (4.9)"},{"comment":"'supp E[(X^(p))^8]' should presumably be 'sup_p E[(X^(p))^8]'; please correct the notation.","section":"Section 2.3, Theorem 2.7"},{"comment":"The strong law of large numbers for triangular arrays is invoked, but the stated fourth-moment condition typically gives convergence in probability, not almost-sure convergence. Either state a suitable theorem with all hypotheses or phrase the conclusion as convergence in probability.","section":"Section 4.4, proof of Lemma 4.4"},{"comment":"The same symbols G_{p,±} are defined for both the cumulative and tail integrals; this is confusing. Clarify by using different notation for the two sets of functions.","section":"Section 2.3, equation (2.5)"},{"comment":"The phrase 'consider λ1,λ2>0' appears to be a typo for 'λ0,λ1>0'.","section":"Section 4.4, Lemma 4.5 proof"},{"comment":"The approximate kinematic formula [2, Theorem I] is cited without checking that its hypotheses (e.g., closed convex cone, Gaussian subspace) are satisfied for C(W) and L; please verify or state the necessary conditions.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a plausible and useful extension of Candès–Sur, and the simulations are supportive, but the proof has gaps in key lemmas (Lemma 4.5 and Lemma 4.6) that are load-bearing for Theorem 4.3. The authors should be asked to either repair these arguments or clearly state the additional assumptions under which the theorem is proven. The reliance on [12] and [2] without full details is acceptable in this field, but the proof of Theorem 2.7 should be substantially expanded. The abstract's 'wide range' claim should be tempered unless (2.7) is proven under more primitive conditions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: the paper probably has the right answer, but the proof as written has a couple of holes in the stochastic-approximation argument. It is worth engaging with; it deserves a referee, but a referee should ask for revisions.\n\nWhat is new: the phase transition threshold for MLE existence in high-dimensional binary GLMs is extended from Gaussian covariates (Candès–Sur) to elliptical covariates, under a 'projection limit' assumption on the one-dimensional marginal. The threshold takes the same convex-program form, but the paper identifies a condition (Assumption 2.6) that is exactly what allows the Gaussian case to go through, and shows that when it fails (log-normal covariates) the phase transition formula breaks. The sufficient moment condition and the stochastic-approximation route are also new, at least in this area. The simulations match theory for Gamma, Pareto, half-normal, and chi, and the paper honestly reports the log-normal mismatch.\n\nWhat is not so solid: Theorem 4.3 is the load-bearing step, and its proof relies on Lemma 4.5 and Lemma 4.6. Lemma 4.5 claims strong convexity of the limit objective G(λ), but the argument only shows strict convexity pointwise and then waves at approximations for large λ0,λ1. No error bounds are given, and the case of arbitrary signs/directions is not covered. This is a real gap, though probably fillable. Lemma 4.6 has a more minor issue: the variance bound in (4.15) is written as if it were uniform in λ, but on a circle of radius x the variance of g grows like x^4. The Chebyshev argument still works for fixed x, so the conclusion of boundedness in probability can likely be repaired with a small clarification. The condition (2.7) is also technical and only checked via the sufficient condition (2.8); the paper doesn't fully verify (2.8) for all claimed examples, but it looks plausible. No code or data is shipped, and simulation plots have no error bars — minor.\n\nOn circularity: there is none. The threshold h_MLE is computed from the model distribution, not fit to existence data, and the phase transition is tested independently.\n\nWho is this for: anyone working on high-dimensional GLM asymptotics, likelihood ratio tests, or convex geometry of random cones. The paper is a useful extension of [12], and the projection-limit assumption is a genuine conceptual contribution.\n\nMy recommendation: accept for peer review with the expectation of revision. A serious referee should push on Lemma 4.5 and ask for a complete proof of Theorem 4.3. If those get fixed, this is a solid paper.","headline":"A plausible extension of the Candès–Sur phase transition to elliptical covariates, but the proof of the key convergence theorem has gaps that need fixing before the result is fully rigorous.","tokens_in":18476,"tokens_out":5209,"would_cite":true,"duration_ms":40733,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62J12","62H05"],"pacs":[],"model":"deepseek-v4-flash","headline":"In high-dimensional binary regression with elliptical covariates, the maximum likelihood estimate exists with probability tending to one below an explicit threshold $h_{\\mathrm{MLE}}$ and with probability tending to zero above it.","keywords":["maximum likelihood estimation","phase transition","binary response generalized linear models","elliptical distributions","high-dimensional asymptotics","data separation","convex geometry","stochastic approximation"],"falsifier":"Run the linear-programming separation test (3.1) for n=1000, p=$\\kappa n$, with spherical covariates whose radial component is Gamma(1,$\\theta_p$) scaled so $\\mathbb{E}R^2=p+1$, logit link, $\\beta_0=0$, across a grid of $\\gamma_0$ and $\\kappa$, and compare the empirical 50% existence boundary with $h_{\\mathrm{MLE}}$ from (2.6); the theorem predicts agreement within the uncertainty band, so a systematic gap would falsify it. The log-normal design, where the projection-limit assumption fails, is the built-in negative control: there the gap is real.","tokens_in":17468,"feed_emoji":"📊","tokens_out":15060,"duration_ms":132694,"temperature":0.7,"pith_summary":"The paper establishes that, for binary-response generalized linear models with elliptical covariates in high dimensions, whether the maximum likelihood estimate exists is governed by a sharp phase transition in the ratio $\\kappa = p/n$. When $\\kappa$ lies below an explicit threshold $h_{\\mathrm{MLE}}$, the MLE exists with probability tending to one; above it, the data are separable and the MLE fails with probability tending to one. The threshold depends on the link function, the intercept, and the limiting scaling of the regression signal, and it can be computed by solving a small convex optimization problem. This matters because it tells practitioners, before fitting, when a model can be fitted at all, and it shows that the Gaussian assumption in earlier work can be relaxed to a broad class of elliptical distributions. The paper also identifies the boundary of its own result: the phase transition breaks down when the one-dimensional projections of the covariates do not stabilize in distribution, as with log-normal covariates.","feed_headline":"A computable threshold decides when the MLE exists","feed_subtitle":"For elliptical covariates, the maximum likelihood estimate appears below the threshold and is absent above it, with probability tending to…","key_machinery":"The key object is the threshold $h_{\\mathrm{MLE}}$, the minimal expected squared positive part of $\\lambda_0 Y+\\lambda_1 X-Z$ over $(\\lambda_0,\\lambda_1)$. It does the heavy lifting through conic geometry: for log-concave links the MLE fails exactly when the data are separable, and for spherical covariates separability is equivalent to a random $(p-1)$-dimensional subspace $L$ (spanned by the nuisance coordinates) intersecting a cone $C(W)$ built from the labels and the signal coordinate. The approximate kinematic formula of convex geometry says such an intersection becomes overwhelmingly likely when $p-1+\\delta(C(W))$ exceeds $n$, where the statistical dimension $\\delta(C(W))$ is $n$ minus the expected squared distance from a standard normal vector to $C(W)$; the limit of that distance is $n\\,h_{\\mathrm{MLE}}(\\alpha_0,\\beta_0,\\gamma_0)$. A stochastic approximation lemma, proved by epi-convergence, shows that the empirical minimization defining the threshold converges to its expectation, and condition (2.7) bounds the probability of separation using only the signal coordinate. Assumption 2.6, that the projections $U^{(p)}$ converge in distribution, makes the limit $(Y,X)$ well defined; the moment criterion (2.3) gives a checkable sufficient condition.","core_discovery":"The central claim is Theorem 2.7. After rotating the covariates so all signal lies in one coordinate, write $(Y^{(p)},X^{(p)})=(V^{(p)},V^{(p)}U^{(p)})$, where $U^{(p)}$ is a single coordinate of the standardized elliptical covariate vector and $P(V^{(p)}=1\\,|\\,U^{(p)})=\\sigma(\\beta_0+(\\gamma_0/\\alpha_0)U^{(p)})$. Assume the link satisfies log-concavity of $\\sigma$ and $1-\\sigma$, the covariates are full-rank elliptical with $\\mathbb{E}R^2/p\\to\\alpha_0^2$, the coefficient scaling obeys $|\\Sigma^{1/2}\\beta|\\to\\gamma_0/\\alpha_0$, and the projections converge in distribution, $U^{(p)}\\Rightarrow U$. Let $Z\\sim N(0,1)$ be independent and define $h_{\\mathrm{MLE}}(\\alpha_0,\\beta_0,\\gamma_0)=\\min_{\\lambda_0,\\lambda_1}\\mathbb{E}(\\lambda_0 Y+\\lambda_1 X-Z)_+^2$. Then $\\kappa=p/n>h_{\\mathrm{MLE}}$ implies the maximum likelihood estimate exists with probability tending to $0$, while $\\kappa<h_{\\mathrm{MLE}}$ implies it exists with probability tending to $1$, subject to the finite-eighth-moment condition and condition (2.7). This is the same phase transition known for Gaussian logistic regression, now extended to any elliptical family whose one-dimensional projections stabilize, with the boundary computed from the limiting projection distribution.","pith_inferences":["If, as the authors conjecture, the same conic-geometry proof carries over to multinomial, Poisson, or log-linear models, the threshold method would become a general identifiability diagnostic for large categorical regressions; that extension is not proved here.","The log-normal failure suggests that for radial distributions with indeterminate moment sequences, MLE existence may depend on the full radial law and not just its low moments; a natural experiment is to hold $\\mathbb{E}R^2$ fixed and vary the log-normal scale to see whether the empirical phase transition moves with the moment-indeterminate part of the distribution.","The stochastic approximation lemma is stated for a general convex loss bounded below, so the same phase-transition machinery may apply to other convex classification losses such as hinge or squared hinge; in that case the threshold would mark where the corresponding risk-minimizing classifier stops being well defined.","Practically, one could use the gap between empirical MLE-existence frequencies and the formula (2.6) as a diagnostic for whether covariate projections stabilize, checking the moment condition (2.3) on data before trusting Gaussian-based asymptotics."],"forward_implications":["Below the threshold $\\kappa<h_{\\mathrm{MLE}}$, a practitioner can expect the MLE to exist with probability tending to one; above it, data separation occurs almost surely and standard fitting algorithms have no finite solution to find.","The threshold can be evaluated in advance by solving a two-dimensional convex program for the limiting projection distribution, so the result supplies a practical pre-fit diagnostic.","The phase transition holds across logit, probit, and cloglog links, and across Gaussian, Gamma, Pareto (with enough moments), and half-normal elliptical covariates, because these satisfy the projection-limit and univariate-separation conditions.","When the projection-limit assumption fails, as for log-normal covariates, the theoretical curve no longer matches simulations, so the assumption marks the real boundary of the phenomenon rather than a technical convenience."],"supporting_citations":[{"why":"Supplies the Gaussian logistic phase transition and the proof template this paper extends.","marker":"[12]"},{"why":"Establishes that for log-concave links the MLE exists iff the data overlap, turning the question into geometry.","marker":"[1]"},{"why":"Provides the approximate kinematic formula for random cones used to convert overlap into a comparison with n and the statistical dimension.","marker":"[2]"},{"why":"Defines elliptical distributions and the marginal and conditional structure used to reduce to a univariate projection.","marker":"[11]"},{"why":"Gives the consistency of minimizers for stochastic programs on which Lemma 4.4 relies.","marker":"[3]"},{"why":"Supplies the epi-convergence principle used to pass from convex objective convergence to optimal-value convergence.","marker":"[17]"},{"why":"Provides uniform convergence of convex optimization problems, used in the proof that the threshold program converges.","marker":"[24]"},{"why":"Offers the lemma transferring convergence in distribution to convergence of expectations in the threshold formula.","marker":"[8]"},{"why":"Gives the moment criterion used to check the projection-limit assumption.","marker":"[30]"}],"fun_headline_variants":["Elliptical covariates shift the MLE threshold","Phase transition proven for elliptical binary GLMs","Beyond Gaussian: MLE existence phase transition","MLE existence: a sharp threshold for elliptical data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the one-dimensional projections of the standardized covariates converge in distribution to a fixed limit, because the threshold $h_{\\mathrm{MLE}}$ is computed from that limiting distribution; when it fails, the paper's own simulations show the formula does not hold.","fun_headline_variants_meta":{"raw":{"variants":["Elliptical covariates shift the MLE threshold","Phase transition proven for elliptical binary GLMs","Beyond Gaussian: MLE existence phase transition","MLE existence: a sharp threshold for elliptical data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00068,"raw_usage":{"total_tokens":3103,"prompt_tokens":971,"completion_tokens":2132,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":2074}},"tokens_in":587,"tokens_out":2132,"duration_ms":17071,"temperature":1.0,"reasoning_tokens":2074,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:52:40.530374+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the linear-programming separation test (3.1) for n=1000, p=$\\kappa n$, with spherical covariates whose radial component is Gamma(1,$\\theta_p$) scaled so $\\mathbb{E}R^2=p+1$, logit link, $\\beta_0=0$, across a grid of $\\gamma_0$ and $\\kappa$, and compare the empirical 50% existence boundary with $h_{\\mathrm{MLE}}$ from (2.6); the theorem predicts agreement within the uncertainty band, so a systematic gap would falsify it. The log-normal design, where the projection-limit assumption fails, is the built-in negative control: there the gap is real.","supporting_citations":[{"cited_title":"Cand` es and Pragya Sur","cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian logistic phase transition and the proof template this paper extends."},{"cited_title":"On the existence of maximum likelihood estimates in logistic regression models","cited_arxiv_id":null,"evidence_quote":"Establishes that for log-concave links the MLE exists iff the data overlap, turning the question into geometry."},{"cited_title":"Living on the edge: Phase transitions in convex programs with random data","cited_arxiv_id":null,"evidence_quote":"Provides the approximate kinematic formula for random cones used to convert overlap into a comparison with n and the statistical dimension."},{"cited_title":"On the theory of elliptically contoured distribu- tions","cited_arxiv_id":null,"evidence_quote":"Defines elliptical distributions and the marginal and conditional structure used to reduce to a univariate projection."},{"cited_title":"Consistency of minimizers and the slln for stochastic programs","cited_arxiv_id":null,"evidence_quote":"Gives the consistency of minimizers for stochastic programs on which Lemma 4.4 relies."},{"cited_title":"Asymptotic behavior of statistical estimators and of optimal solutions of stochastic optimization problems","cited_arxiv_id":null,"evidence_quote":"Supplies the epi-convergence principle used to pass from convex objective convergence to optimal-value convergence."},{"cited_title":"Kanniappan and S","cited_arxiv_id":null,"evidence_quote":"Provides uniform convergence of convex optimization problems, used in the proof that the threshold program converges."},{"cited_title":"Some asymptotic theory for the bootstrap","cited_arxiv_id":null,"evidence_quote":"Offers the lemma transferring convergence in distribution to convergence of expectations in the threshold formula."},{"cited_title":"Recent developments on the moment problem","cited_arxiv_id":null,"evidence_quote":"Gives the moment criterion used to check the projection-limit assumption."}],"review_version":1}