REVIEW 3 major objections 5 minor 15 references
Get me out of this hole: a profile likelihood approach to identifying and avoiding inferior local optima in choice models
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that a systematic profile-likelihood search, iterated from any improved constrained solution, escapes inferior local optima in latent class and mixed logit models and yields solutions whose asymptotic-normality…
desk verdict A practical heuristic that works in its case studies, but the 'guarantee' language outstrips the unverified constrained optimization underneath it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the profile log-likelihood $\mathrm{LL}_{PL}(\tilde{\beta}_k)$ — the highest log-likelihood achievable when parameter $k$ is held fixed at $\tilde{\beta}_k$ and all other parameters are freely re-estimated. The algorithm evaluates this profile on a symmetric grid of 51 trial values per parameter, spaced at $\pm 4$ estimated robust standard errors around the base estimate, and treats any constrained solution with $\mathrm{LL} > \mathrm{LL}_{\mathrm{base}}$ as evidence of an inferior starting optimum. An outer loop (steps B1–B4) turns each such improved constrained solution into a starting point for unconstrained estimation, removes duplicate solutions, and repeats the profiling from the survivors until no further improvements and no non-quadratic profile shapes are found.
What would settle it
A decisive test is a simulated likelihood with a planted superior optimum reachable only by changing two parameters at once; if the iterative profile search, started from the inferior basin, never reaches the planted optimum, the one-parameter-grid premise is false.
Extended reading notes
Core claim
The central discovery, stated as the authors intend it, is that the maximum-likelihood solution a routine estimation run returns from reasonable starting values is often only one of many local optima, and that the 'best' of these can be reached by a structured one-parameter search rather than by hoping random starting values land in the right basin. For each parameter $k$, the paper defines the profile log-likelihood $\mathrm{LL}_{PL}(\tilde{\beta}_k) = \max_{\beta_k=\tilde{\beta}_k} \mathrm{LL}_N(\beta)$, evaluates it at 51 evenly spaced shifts $\tilde{\beta}_{k,m} = \hat{\beta}_{k,\mathrm{base}} + \gamma_m \hat{\sigma}_{k,\mathrm{base}}$ with $\gamma_m$ running from $-4$ to $4$, and re-estimates the remaining parameters at each shift using the base estimates as starting values. Any constrained model whose log-likelihood exceeds the base log-likelihood is a candidate better optimum; unconstrained re-estimation from those candidates and filtering of duplicates yields new solutions, and the whole loop repeats from each new solution. On the case-study data this recovers better optima for both latent class and mixed logit, and the improved solutions' profiles show the log-likelihood drop at $\pm 1.96\hat{\sigma}_k$ close to the asymptotic value of $1.92$, indicating the quadratic approximation is more reliable at the new solutions than at the original one.
Load-bearing premise
The method assumes that any better optimum can be reached by changing one parameter at a time along a grid of 51 points out to plus or minus four standard errors from the initial estimate, and that each constrained re-estimation returns the best fit over the remaining parameters.
Editorial extensions
If this is right
- Estimators that report a converged local optimum can now be audited with a structured search that does not depend on the luck of random starting values.
- For latent class models, the case study shows that solutions with nearly equal in-sample fit can imply materially different willingness-to-pay, so policy analysis should compare profile-derived solutions before relying on one run.
- For mixed logit, the search found 76 distinct better solutions, and the best one improves fit only slightly but yields narrower confidence intervals and more stable profiles.
- The final solutions' adherence to the expected 1.92 log-likelihood drop at ±1.96 standard errors gives analysts a practical check on whether asymptotic inference is trustworthy at the reported optimum.
Reading between the lines
- Because the search moves one parameter at a time, it could miss optima that require simultaneous shifts in several coordinates; a two-parameter or blockwise version of the same profile idea would be a natural extension to test.
- The same profiling procedure transfers directly to other extremum estimators such as nonlinear least squares or GMM, though the constrained re-estimations would need to be checked for their own multimodality.
- The observation that similar log-likelihoods hide different valuations suggests reporting profile traces as standard output of choice-model estimation, letting readers assess how much of the policy result is an artifact of which optimum the optimizer found.
- The improved asymptotic-normality behaviour of the final solutions hints that profile shape could be used as a model-selection criterion among competing local optima, not just as a search device.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a profile-likelihood-based procedure for diagnosing and escaping inferior local optima in maximum likelihood estimation of choice models. Starting from a base MLE, Step A1 defines a one-dimensional grid around each parameter using multiples of its standard error; Step A2 re-estimates the model with each parameter fixed at each grid point; Step A3 compares the constrained log-likelihoods with the base log-likelihood to detect better adjacent solutions. The paper then iterates the procedure from the best improved solutions (Steps B1–B4). Two empirical case studies on a Swiss stated-preference route-choice dataset are presented: a two-class latent class model and a multivariate mixed logit model. In both cases the procedure finds solutions with higher log-likelihood than the base solution, and a second round of profiling from the best new solution finds no further improvements. The paper also reports that the improved solutions exhibit profile shapes closer to what would be expected under asymptotic normality, and that differences in willingness-to-pay across local optima can be substantial even when log-likelihood differences are small.
Significance. If the central claim holds, the paper offers a practical, relatively easy-to-implement addition to the current toolbox for handling multimodality in choice model estimation. The topic is important: local optima are routinely acknowledged as a risk in latent class and mixed logit models, but published evidence on their prevalence and on systematic ways to escape them is scarce. The empirical demonstrations are genuine: the improved candidate solutions are obtained from constrained fits whose log-likelihoods exceed the base value, and the subsequent unconstrained re-estimation and filtering in Steps B2–B3 provide a real check that the improvements are not artifacts of a single optimization run. The paper also usefully connects its diagnostic to McCullough and Vinod's recommendation to profile the likelihood, and it reports Hessian eigenvalue and conditioning checks for the candidate solutions.
major comments (3)
- [§2.3, Steps A2–A3] The profile values fLL_{mk} are obtained from a single constrained optimization of the remaining K−1 parameters starting from the base MLE vector. Nothing in Step A2 ensures that this optimization reaches the global maximum of the constrained problem. If the constrained log-likelihood is itself multimodal, the reported fLL_{mk} can be too low, and the Step A3 statement that 'LLbase > fLL_{mk} for all k,m' offers some 'guarantee' of a good local optimum is therefore not supported. The same issue undermines the round-2 conclusions in Sections 3.4 and 3.5 that no further improvements exist after profiling from s*_1 and s*_22, because a missed constrained maximum could hide a better adjacent optimum. The authors should either verify each constrained maximum (for example by multi-start runs or by checking first-order and second-order conditions at every profile point) or explicitly reframe the method as a heuristic and remove the guarantee and no-further-improvement language.
- [§2.3, Eq. (9) and §2.4] The search space is restricted to univariate shifts of the form β_k = βhat_k,base + γ_m σhat_k,base with γ_a = −4, γ_b = 4 and M = 51. A superior optimum that requires simultaneous movement of several parameters, or that lies outside this grid range, will not be detected by Step A3. The abstract's and conclusions' phrase 'systematically analyses the parameter space' therefore overstates the coverage of the method. The authors should state this limitation prominently in the methodology section and, ideally, include a sensitivity check with a wider grid or with bivariate profiles for at least one parameter pair, so that readers can judge how likely the current grid is to miss relevant optima.
- [§3.4–§3.5] The counts of 99 and 210 'better local optima' reported in Step A3 refer to constrained profile points with fLL_{mk} > LLbase, not to distinct unconstrained local optima. The relevant number of distinct new solutions after the unconstrained Step B2 and filtering in Step B3 is 4 for the latent class model and 76 for the mixed logit model. The manuscript would be clearer if this distinction were made explicit, since otherwise readers may overestimate the multiplicity of genuine local maxima identified by the procedure.
minor comments (5)
- [§2.2, Eq. (6)] The statement that the likelihood ratio statistic is distributed chi-squared with one degree of freedom should mention that this is an asymptotic result holding under the usual regularity conditions, rather than a finite-sample property.
- [§3.4, Table 3] The column headers 'Profile likelihoods*1' and similar are confusingly formatted; using 's*_1' consistently, as in the text, would improve readability.
- [Figures 2–9] The legend label 'MLE/MLE LL' is unclear; it should be defined in the captions, for example as the base-model log-likelihood at the MLE.
- [§3.4, Figure 7 and §3.5, Figures 10–11] The claim that the improved solution is 'more in line with asymptotic normality' is based on visual inspection of profiles and on the single necessary condition that a ±1.96σ shift produces roughly a 1.92 drop in log-likelihood. This is suggestive but not a formal test; the paper should either add a quantitative diagnostic or soften the wording.
- [§4, first paragraph of the final discussion] There is a typo in 'While analyst may be concerned about the computational cost'; it should read 'While an analyst may be concerned'.
Circularity Check
No significant circularity: the profile likelihood values and iterative improvements are computed from the data and compared directly, not derived from the paper's conclusions.
full rationale
The paper's derivation chain is self-contained. The central quantity fLL_mk is defined in Eq. (7) as the constrained maximum of LL_N(beta) with beta_k fixed, and Step A2 computes it by re-estimating the remaining parameters; Step A3 then compares these values with LL_base. A candidate point is declared better only when fLL_mk > LL_base, so the empirical finding of superior local optima is a data-dependent outcome, not an input built into the grid or the profiling procedure. The iterative algorithm (B1-B4) uses the better constrained solutions as starting values for unconstrained estimation and re-profiles each new optimum, so the final solution's status is again determined by fresh log-likelihood comparisons. The claimed asymptotic-normality benefit is also not part of the objective: the search maximizes LL, and the more quadratic profile shape of s*_1/s*_22 is an observed property reported after the fact. The use of ±4 standard-error grid (Eq. 9) is a heuristic scaling choice, and the one-parameter-at-a-time movement is an acknowledged limitation that could miss optima, but neither reduces the result to an input. Self-citations (Apollo software, Hess and Palma 2019; Bunch 2024; the Axhausen et al. 2008 dataset) supply estimation tools, optimization background, and data, not the load-bearing mathematical argument, and the profile-likelihood concept is credited to McCullough and Vinod (2003). No equation or fitted parameter was found that is equivalent by construction to the paper's conclusions.
Assumptions & free parameters
free parameters (4)
- profile grid width (gamma_a, gamma_b) =
-4 to 4
- number of grid points M =
51
- standard error type for grid scaling =
robust
- number of Sobol draws =
500
assumptions (3)
- standard math Standard asymptotic MLE theory: under correct specification and a unique true parameter vector, the MLE is consistent and asymptotically normal with sandwich variance.
- domain assumption The log-likelihood of a linear-in-parameters multinomial logit model is globally concave, so its MLE is unique.
- domain assumption Latent class and mixed logit models can have multiple local optima in their log-likelihood functions.
Cite this review
Pith. "Pith review of Get me out of this hole: a profile likelihood approach to identifying and avoiding inferior local optima in choice models." pith.science (2026). https://pith.science/paper/A5O7ZJCI
@misc{pith2026250602722,
author = {Pith},
title = {Pith review of: Get me out of this hole: a profile likelihood approach to identifying and avoiding inferior local optima in choice models},
year = {2026},
howpublished = {\url{https://pith.science/paper/A5O7ZJCI}},
note = {Machine review of arXiv:2506.02722}
}
read the original abstract
Choice modellers routinely acknowledge the risk of convergence to inferior local optima when using structures other than a simple linear-in-parameters logit model. At the same time, there is no consensus on appropriate mechanisms for addressing this issue. Most analysts seem to ignore the problem, while others try a set of different starting values, or put their faith in what they believe to be more robust estimation approaches. This paper puts forward the use of a profile likelihood approach that systematically analyses the parameter space around an initial maximum likelihood estimate and tests for the existence of better local optima in that space. We extend this to an iterative algorithm which then progressively searches for the best local optimum under given settings for the algorithm. Using a well known stated choice dataset, we show how the approach identifies better local optima for both latent class and mixed logit, with the potential for substantially different policy implications. In the case studies we conduct, an added benefit of the approach is that the new solutions exhibit properties that more closely adhere to the property of asymptotic normality, also highlighting the benefits of the approach in analysing the statistical properties of a solution.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
author Amemiya, T. , year 1985 . title Advanced Econometrics . publisher Harvard University Press , address Cambridge, MA
work page 1985
-
[2]
author Armstrong, P. , author Garrido, R.A. , author Ort \' u zar, J. de D . , year 2001 . title Confidence interval to bound the value of time . journal Transportation Research Part E volume 37 , pages 143--161
work page 2001
-
[3]
author Axhausen, K.W. , author Hess, S. , author K \"o nig, A. , author Abay, G. , author Bates, J.J. , author Bierlaire, M. , year 2008 . title State of the art estimates of the swiss value of travel time savings . journal Transport Policy volume 15 , pages 173--185
work page 2008
-
[4]
author Ben-Akiva, M. , author Swait, J. , year 1986 . title The Akaike Likelihood Ratio Index . journal Transportation Science volume 20 , pages 133--136
work page 1986
-
[5]
author Bierlaire, M. , author Th \'e mans, M. , author Zufferey, N. , year 2010 . title A heuristic for nonlinear global optimization . journal INFORMS Journal on Computing volume 22 , pages 59--70
work page 2010
-
[6]
author Bunch, D.S. , year 2024 . title Numerical methods for optimization-based model estimation and inference , in: editor Hess, S. , editor Daly, A. (Eds.), booktitle Handbook of Choice Modelling, second edition . publisher Edward Elgar , p. pages 594–629
work page 2024
-
[7]
author Bunch, D.S. , author Gay, D.M. , author Welsch, R.E. , year 1993 . title Algorithm 717: Subroutines for maximum likelihood and quasi-likelihood estimation of parameters in nonlinear regression models . journal ACM Trans. Math. Softw. volume 19 , pages 109–130 . https://doi.org/10.1145/151271.151279, :10.1145/151271.151279
-
[8]
author Cameron, A.C. , author Trivedi, P.K. , year 2005 . title Microeconometrics : methods and applications . publisher Cambridge University Press , address Cambridge
work page 2005
Show all 15 references
-
[9]
, author Hess, S
author Daly, A. , author Hess, S. , author de Jong, G. , year 2012 . title Calculating errors for measures derived from choice modelling estimates . journal Transportation Research Part B volume 46 , pages 333--341
2012
-
[10]
, author MacKinnon, J.G
author Davidson, R. , author MacKinnon, J.G. , year 1993 . title Estimation and inference in econometrics . publisher Oxford University Press , address New York
1993
-
[11]
, author Palma, D
author Hess, S. , author Palma, D. , year 2019 . title Apollo: A flexible, powerful and customisable freeware package for choice model estimation and application . journal Journal of Choice Modelling volume 32 , pages 100170 . :https://doi.org/10.1016/j.jocm.2019.100170
2019
-
[12]
, author Vinod, H.D
author McCullough, B.D. , author Vinod, H.D. , year 2003 . title Verifying the solution from a nonlinear solver: A case study . journal American Economic Review volume 93 , pages 873–892 . https://www.aeaweb.org/articles?id=10.1257/000282803322157133, :10.1257/000282803322157133
2003 doi
-
[13]
, year 1967
author Sobol, I.M. , year 1967 . title On the distribution of points in a cube and the approximate evaluation of integrals . journal Zhurnal Vychislitelnoi Matematiki i Matematicheskoi Fiziki volume 7 , pages 784--802
1967
-
[14]
, year 2009
author Train, K. , year 2009 . title Discrete Choice Methods with Simulation . edition second edition ed., publisher Cambridge University Press , address Cambridge, MA
2009
-
[15]
, author Weeks, M
author Train, K. , author Weeks, M. , year 2005 . title Discrete choice models in preference space and willingness-to-pay space , in: editor Scarpa, R. , editor Alberini, A. (Eds.), booktitle Application of simulation methods in environmental and resource economics . publisher...
2005
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.