Pith. sign in

REVIEW 3 major objections 5 minor 15 references

Get me out of this hole: a profile likelihood approach to identifying and avoiding inferior local optima in choice models

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims that a systematic profile-likelihood search, iterated from any improved constrained solution, escapes inferior local optima in latent class and mixed logit models and yields solutions whose asymptotic-normality…

desk verdict A practical heuristic that works in its case studies, but the 'guarantee' language outstrips the unverified constrained optimization underneath it. read the letter →

arxiv 2506.02722 v1 pith:A5O7ZJCI submitted 2025-06-03 econ.EM

classification econ.EM
keywords choicemodellinglocaloptimaprofilelikelihoodlatentclassmixedlogitmaximumasymptoticnormalitywillingnesstopay
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that an apparently converged maximum-likelihood solution for a choice model can be an inferior local optimum, and that a systematic profile-likelihood sweep can find better ones. The proposed procedure fixes one parameter at a time at a grid of trial values around the initial estimate, re-estimates all other parameters, and then uses any constrained solution with higher log-likelihood as a starting point for unconstrained re-estimation, iterating until no further improvements appear. Applied to a well-known Swiss stated-choice dataset, the method identifies four new superior local optima for a two-class latent class model and seventy-six for a mixed logit model, with the best solutions improving log-likelihood from -1,578.26 to -1,552.53 and from -1,405.20 to -1,403.43 respectively. The final solutions also show log-likelihood profiles closer to the quadratic shape expected under asymptotic normality, and their implied willingness-to-pay values can differ enough from the base solution to change policy conclusions.

What carries the argument

The load-bearing object is the profile log-likelihood $\mathrm{LL}_{PL}(\tilde{\beta}_k)$ — the highest log-likelihood achievable when parameter $k$ is held fixed at $\tilde{\beta}_k$ and all other parameters are freely re-estimated. The algorithm evaluates this profile on a symmetric grid of 51 trial values per parameter, spaced at $\pm 4$ estimated robust standard errors around the base estimate, and treats any constrained solution with $\mathrm{LL} > \mathrm{LL}_{\mathrm{base}}$ as evidence of an inferior starting optimum. An outer loop (steps B1–B4) turns each such improved constrained solution into a starting point for unconstrained estimation, removes duplicate solutions, and repeats the profiling from the survivors until no further improvements and no non-quadratic profile shapes are found.

What would settle it

A decisive test is a simulated likelihood with a planted superior optimum reachable only by changing two parameters at once; if the iterative profile search, started from the inferior basin, never reaches the planted optimum, the one-parameter-grid premise is false.

Watch

Extended reading notes

Core claim

The central discovery, stated as the authors intend it, is that the maximum-likelihood solution a routine estimation run returns from reasonable starting values is often only one of many local optima, and that the 'best' of these can be reached by a structured one-parameter search rather than by hoping random starting values land in the right basin. For each parameter $k$, the paper defines the profile log-likelihood $\mathrm{LL}_{PL}(\tilde{\beta}_k) = \max_{\beta_k=\tilde{\beta}_k} \mathrm{LL}_N(\beta)$, evaluates it at 51 evenly spaced shifts $\tilde{\beta}_{k,m} = \hat{\beta}_{k,\mathrm{base}} + \gamma_m \hat{\sigma}_{k,\mathrm{base}}$ with $\gamma_m$ running from $-4$ to $4$, and re-estimates the remaining parameters at each shift using the base estimates as starting values. Any constrained model whose log-likelihood exceeds the base log-likelihood is a candidate better optimum; unconstrained re-estimation from those candidates and filtering of duplicates yields new solutions, and the whole loop repeats from each new solution. On the case-study data this recovers better optima for both latent class and mixed logit, and the improved solutions' profiles show the log-likelihood drop at $\pm 1.96\hat{\sigma}_k$ close to the asymptotic value of $1.92$, indicating the quadratic approximation is more reliable at the new solutions than at the original one.

Load-bearing premise

The method assumes that any better optimum can be reached by changing one parameter at a time along a grid of 51 points out to plus or minus four standard errors from the initial estimate, and that each constrained re-estimation returns the best fit over the remaining parameters.

Editorial extensions

If this is right

  • Estimators that report a converged local optimum can now be audited with a structured search that does not depend on the luck of random starting values.
  • For latent class models, the case study shows that solutions with nearly equal in-sample fit can imply materially different willingness-to-pay, so policy analysis should compare profile-derived solutions before relying on one run.
  • For mixed logit, the search found 76 distinct better solutions, and the best one improves fit only slightly but yields narrower confidence intervals and more stable profiles.
  • The final solutions' adherence to the expected 1.92 log-likelihood drop at ±1.96 standard errors gives analysts a practical check on whether asymptotic inference is trustworthy at the reported optimum.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the search moves one parameter at a time, it could miss optima that require simultaneous shifts in several coordinates; a two-parameter or blockwise version of the same profile idea would be a natural extension to test.
  • The same profiling procedure transfers directly to other extremum estimators such as nonlinear least squares or GMM, though the constrained re-estimations would need to be checked for their own multimodality.
  • The observation that similar log-likelihoods hide different valuations suggests reporting profile traces as standard output of choice-model estimation, letting readers assess how much of the policy result is an artifact of which optimum the optimizer found.
  • The improved asymptotic-normality behaviour of the final solutions hints that profile shape could be used as a model-selection criterion among competing local optima, not just as a search device.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a profile-likelihood-based procedure for diagnosing and escaping inferior local optima in maximum likelihood estimation of choice models. Starting from a base MLE, Step A1 defines a one-dimensional grid around each parameter using multiples of its standard error; Step A2 re-estimates the model with each parameter fixed at each grid point; Step A3 compares the constrained log-likelihoods with the base log-likelihood to detect better adjacent solutions. The paper then iterates the procedure from the best improved solutions (Steps B1–B4). Two empirical case studies on a Swiss stated-preference route-choice dataset are presented: a two-class latent class model and a multivariate mixed logit model. In both cases the procedure finds solutions with higher log-likelihood than the base solution, and a second round of profiling from the best new solution finds no further improvements. The paper also reports that the improved solutions exhibit profile shapes closer to what would be expected under asymptotic normality, and that differences in willingness-to-pay across local optima can be substantial even when log-likelihood differences are small.

Significance. If the central claim holds, the paper offers a practical, relatively easy-to-implement addition to the current toolbox for handling multimodality in choice model estimation. The topic is important: local optima are routinely acknowledged as a risk in latent class and mixed logit models, but published evidence on their prevalence and on systematic ways to escape them is scarce. The empirical demonstrations are genuine: the improved candidate solutions are obtained from constrained fits whose log-likelihoods exceed the base value, and the subsequent unconstrained re-estimation and filtering in Steps B2–B3 provide a real check that the improvements are not artifacts of a single optimization run. The paper also usefully connects its diagnostic to McCullough and Vinod's recommendation to profile the likelihood, and it reports Hessian eigenvalue and conditioning checks for the candidate solutions.

major comments (3)
  1. [§2.3, Steps A2–A3] The profile values fLL_{mk} are obtained from a single constrained optimization of the remaining K−1 parameters starting from the base MLE vector. Nothing in Step A2 ensures that this optimization reaches the global maximum of the constrained problem. If the constrained log-likelihood is itself multimodal, the reported fLL_{mk} can be too low, and the Step A3 statement that 'LLbase > fLL_{mk} for all k,m' offers some 'guarantee' of a good local optimum is therefore not supported. The same issue undermines the round-2 conclusions in Sections 3.4 and 3.5 that no further improvements exist after profiling from s*_1 and s*_22, because a missed constrained maximum could hide a better adjacent optimum. The authors should either verify each constrained maximum (for example by multi-start runs or by checking first-order and second-order conditions at every profile point) or explicitly reframe the method as a heuristic and remove the guarantee and no-further-improvement language.
  2. [§2.3, Eq. (9) and §2.4] The search space is restricted to univariate shifts of the form β_k = βhat_k,base + γ_m σhat_k,base with γ_a = −4, γ_b = 4 and M = 51. A superior optimum that requires simultaneous movement of several parameters, or that lies outside this grid range, will not be detected by Step A3. The abstract's and conclusions' phrase 'systematically analyses the parameter space' therefore overstates the coverage of the method. The authors should state this limitation prominently in the methodology section and, ideally, include a sensitivity check with a wider grid or with bivariate profiles for at least one parameter pair, so that readers can judge how likely the current grid is to miss relevant optima.
  3. [§3.4–§3.5] The counts of 99 and 210 'better local optima' reported in Step A3 refer to constrained profile points with fLL_{mk} > LLbase, not to distinct unconstrained local optima. The relevant number of distinct new solutions after the unconstrained Step B2 and filtering in Step B3 is 4 for the latent class model and 76 for the mixed logit model. The manuscript would be clearer if this distinction were made explicit, since otherwise readers may overestimate the multiplicity of genuine local maxima identified by the procedure.
minor comments (5)
  1. [§2.2, Eq. (6)] The statement that the likelihood ratio statistic is distributed chi-squared with one degree of freedom should mention that this is an asymptotic result holding under the usual regularity conditions, rather than a finite-sample property.
  2. [§3.4, Table 3] The column headers 'Profile likelihoods*1' and similar are confusingly formatted; using 's*_1' consistently, as in the text, would improve readability.
  3. [Figures 2–9] The legend label 'MLE/MLE LL' is unclear; it should be defined in the captions, for example as the base-model log-likelihood at the MLE.
  4. [§3.4, Figure 7 and §3.5, Figures 10–11] The claim that the improved solution is 'more in line with asymptotic normality' is based on visual inspection of profiles and on the single necessary condition that a ±1.96σ shift produces roughly a 1.92 drop in log-likelihood. This is suggestive but not a formal test; the paper should either add a quantitative diagnostic or soften the wording.
  5. [§4, first paragraph of the final discussion] There is a typo in 'While analyst may be concerned about the computational cost'; it should read 'While an analyst may be concerned'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the profile likelihood values and iterative improvements are computed from the data and compared directly, not derived from the paper's conclusions.

full rationale

The paper's derivation chain is self-contained. The central quantity fLL_mk is defined in Eq. (7) as the constrained maximum of LL_N(beta) with beta_k fixed, and Step A2 computes it by re-estimating the remaining parameters; Step A3 then compares these values with LL_base. A candidate point is declared better only when fLL_mk > LL_base, so the empirical finding of superior local optima is a data-dependent outcome, not an input built into the grid or the profiling procedure. The iterative algorithm (B1-B4) uses the better constrained solutions as starting values for unconstrained estimation and re-profiles each new optimum, so the final solution's status is again determined by fresh log-likelihood comparisons. The claimed asymptotic-normality benefit is also not part of the objective: the search maximizes LL, and the more quadratic profile shape of s*_1/s*_22 is an observed property reported after the fact. The use of ±4 standard-error grid (Eq. 9) is a heuristic scaling choice, and the one-parameter-at-a-time movement is an acknowledged limitation that could miss optima, but neither reduces the result to an input. Self-citations (Apollo software, Hess and Palma 2019; Bunch 2024; the Axhausen et al. 2008 dataset) supply estimation tools, optimization background, and data, not the load-bearing mathematical argument, and the profile-likelihood concept is credited to McCullough and Vinod (2003). No equation or fitted parameter was found that is equivalent by construction to the paper's conclusions.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The method relies on standard asymptotic MLE theory, the known concavity of the MNL log-likelihood, and the empirical premise that LC and MMNL likelihoods are multimodal. The only user-chosen quantities are the profile grid (width and density) and the number of simulation draws, none of which are fitted to the data.

free parameters (4)
  • profile grid width (gamma_a, gamma_b) = -4 to 4
    Analyst choice; wider range increases robustness but also computational cost (Section 3.2).
  • number of grid points M = 51
    Analyst choice; trade-off between coverage and computational cost (Section 3.2).
  • standard error type for grid scaling = robust
    Choice of robust versus classical sigma in Equation 9 affects the grid range (Section 3.2).
  • number of Sobol draws = 500
    Simulation noise in the MMNL likelihood can affect convergence and detection of optima (Section 3.2).
assumptions (3)
  • standard math Standard asymptotic MLE theory: under correct specification and a unique true parameter vector, the MLE is consistent and asymptotically normal with sandwich variance.
    Invoked in Section 2.1 to justify Wald inference and the 1.92 benchmark for a 95% confidence interval.
  • domain assumption The log-likelihood of a linear-in-parameters multinomial logit model is globally concave, so its MLE is unique.
    Used in Section 1 as the baseline case where the local optima problem does not arise.
  • domain assumption Latent class and mixed logit models can have multiple local optima in their log-likelihood functions.
    This is the motivating premise for the entire paper, stated in Section 1 and assumed throughout.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Get me out of this hole: a profile likelihood approach to identifying and avoiding inferior local optima in choice models." pith.science (2026). https://pith.science/paper/A5O7ZJCI

@misc{pith2026250602722,
  author       = {Pith},
  title        = {Pith review of: Get me out of this hole: a profile likelihood approach to identifying and avoiding inferior local optima in choice models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A5O7ZJCI}},
  note         = {Machine review of arXiv:2506.02722}
}
read the original abstract

Choice modellers routinely acknowledge the risk of convergence to inferior local optima when using structures other than a simple linear-in-parameters logit model. At the same time, there is no consensus on appropriate mechanisms for addressing this issue. Most analysts seem to ignore the problem, while others try a set of different starting values, or put their faith in what they believe to be more robust estimation approaches. This paper puts forward the use of a profile likelihood approach that systematically analyses the parameter space around an initial maximum likelihood estimate and tests for the existence of better local optima in that space. We extend this to an iterative algorithm which then progressively searches for the best local optimum under given settings for the algorithm. Using a well known stated choice dataset, we show how the approach identifies better local optima for both latent class and mixed logit, with the potential for substantially different policy implications. In the case studies we conduct, an added benefit of the approach is that the new solutions exhibit properties that more closely adhere to the property of asymptotic normality, also highlighting the benefits of the approach in analysing the statistical properties of a solution.

Figures

Figures reproduced from arXiv: 2506.02722 by the authors.

Figure 1
Figure 1. Profile likelihood results for MNL model: step A [PITH_FULL_IMAGE:figures/full_fig_p013_1.png] view at source ↗
Figure 2
Figure 2. Profile likelihood results for LC model: step A [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 3
Figure 3. Profile likelihood results for LC model: round 2 starting with [PITH_FULL_IMAGE:figures/full_fig_p018_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Profile likelihood results for LC model: round 2 starting with [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]
Figure 5
Figure 5. Figure 5: Profile likelihood results for LC model: round 2 starting with [PITH_FULL_IMAGE:figures/full_fig_p019_5.png]
Figure 6
Figure 6. Figure 6: Profile likelihood results for LC model: round 2 starting with [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: Asymptotic confidence intervals for population mean and standard deviations for base [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]
Figure 8
Figure 8. Figure 8: Profile likelihood results for MMNL model with multivariate distributions: step A [PITH_FULL_IMAGE:figures/full_fig_p024_8.png]
Figure 9
Figure 9. Figure 9: Profile likelihood results for MMNL model with multivariate distributions: round 2 [PITH_FULL_IMAGE:figures/full_fig_p024_9.png]
Figure 10
Figure 10. Figure 10: Asymptotic confidence intervals for population mean and standard deviations for base [PITH_FULL_IMAGE:figures/full_fig_p025_10.png]
Figure 11
Figure 11. Figure 11: Asymptotic confidence intervals for correlations between random terms for base solu [PITH_FULL_IMAGE:figures/full_fig_p026_11.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 14 canonical work pages

  1. [1]

    , year 1985

    author Amemiya, T. , year 1985 . title Advanced Econometrics . publisher Harvard University Press , address Cambridge, MA

  2. [2]

    , author Garrido, R.A

    author Armstrong, P. , author Garrido, R.A. , author Ort \' u zar, J. de D . , year 2001 . title Confidence interval to bound the value of time . journal Transportation Research Part E volume 37 , pages 143--161

  3. [3]

    , author Hess, S

    author Axhausen, K.W. , author Hess, S. , author K \"o nig, A. , author Abay, G. , author Bates, J.J. , author Bierlaire, M. , year 2008 . title State of the art estimates of the swiss value of travel time savings . journal Transport Policy volume 15 , pages 173--185

  4. [4]

    , author Swait, J

    author Ben-Akiva, M. , author Swait, J. , year 1986 . title The Akaike Likelihood Ratio Index . journal Transportation Science volume 20 , pages 133--136

  5. [5]

    , author Th \'e mans, M

    author Bierlaire, M. , author Th \'e mans, M. , author Zufferey, N. , year 2010 . title A heuristic for nonlinear global optimization . journal INFORMS Journal on Computing volume 22 , pages 59--70

  6. [6]

    , year 2024

    author Bunch, D.S. , year 2024 . title Numerical methods for optimization-based model estimation and inference , in: editor Hess, S. , editor Daly, A. (Eds.), booktitle Handbook of Choice Modelling, second edition . publisher Edward Elgar , p. pages 594–629

  7. [7]

    , author Gay, D.M

    author Bunch, D.S. , author Gay, D.M. , author Welsch, R.E. , year 1993 . title Algorithm 717: Subroutines for maximum likelihood and quasi-likelihood estimation of parameters in nonlinear regression models . journal ACM Trans. Math. Softw. volume 19 , pages 109–130 . https://doi.org/10.1145/151271.151279, :10.1145/151271.151279

  8. [8]

    , author Trivedi, P.K

    author Cameron, A.C. , author Trivedi, P.K. , year 2005 . title Microeconometrics : methods and applications . publisher Cambridge University Press , address Cambridge

Show all 15 references
  1. [9]

    , author Hess, S

    author Daly, A. , author Hess, S. , author de Jong, G. , year 2012 . title Calculating errors for measures derived from choice modelling estimates . journal Transportation Research Part B volume 46 , pages 333--341

  2. [10]

    , author MacKinnon, J.G

    author Davidson, R. , author MacKinnon, J.G. , year 1993 . title Estimation and inference in econometrics . publisher Oxford University Press , address New York

  3. [11]

    , author Palma, D

    author Hess, S. , author Palma, D. , year 2019 . title Apollo: A flexible, powerful and customisable freeware package for choice model estimation and application . journal Journal of Choice Modelling volume 32 , pages 100170 . :https://doi.org/10.1016/j.jocm.2019.100170

  4. [12]

    , author Vinod, H.D

    author McCullough, B.D. , author Vinod, H.D. , year 2003 . title Verifying the solution from a nonlinear solver: A case study . journal American Economic Review volume 93 , pages 873–892 . https://www.aeaweb.org/articles?id=10.1257/000282803322157133, :10.1257/000282803322157133

  5. [13]

    , year 1967

    author Sobol, I.M. , year 1967 . title On the distribution of points in a cube and the approximate evaluation of integrals . journal Zhurnal Vychislitelnoi Matematiki i Matematicheskoi Fiziki volume 7 , pages 784--802

  6. [14]

    , year 2009

    author Train, K. , year 2009 . title Discrete Choice Methods with Simulation . edition second edition ed., publisher Cambridge University Press , address Cambridge, MA

  7. [15]

    , author Weeks, M

    author Train, K. , author Weeks, M. , year 2005 . title Discrete choice models in preference space and willingness-to-pay space , in: editor Scarpa, R. , editor Alberini, A. (Eds.), booktitle Application of simulation methods in environmental and resource economics . publisher...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.