Pith. sign in

REVIEW 1 major objections 4 minor 4 references

A Multinomial Probit Model for Asymmetric Choice Responses

T0 review · 1 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A skewed multinomial probit model allows choice probabilities to respond asymmetrically to covariate changes, and the paper shows it is identified and estimable.

desk verdict A genuinely useful reparameterization for skewed multinomial probit with a workable sampler, but identification is asserted rather than proven; deserves a referee round. read the letter →

arxiv 2608.10336 v1 pith:G272IU2E submitted 2026-08-11 econ.EM stat.ME

classification econ.EMstat.ME MSC 62F1562H0562P20
keywords skew-normaldistributionmultinomialprobitasymmetricchoiceresponsesdataaugmentationBayesianinferencepriceelasticitiessubstitutionpatternsscaleidentification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes a skewed multinomial probit (SMNP) model in which the latent utilities follow a multivariate skew-normal distribution, so a positive and a negative change of the same size in a covariate do not have to move choice probabilities equally. The contribution is showing that this extension can be made identified and computationally tractable: a covariance reparameterization fixes the scale and preserves positive definiteness by construction, and a double data-augmentation scheme turns the skew-normal utilities into conditionally Gaussian ones, yielding a Metropolis–Hastings within Gibbs sampler with a single one-dimensional Metropolis step. Numerical experiments and two consumer choice applications indicate that SMNP recovers asymmetric responses when they are present, improves predictive accuracy, and changes implied price elasticities and substitution patterns relative to standard multinomial probit. The paper also shows that when the data are symmetric, SMNP behaves like the standard MNP and does not sacrifice predictive performance.

What carries the argument

The central object is the reparameterized covariance matrix of the double-augmented model, equation (7): $\Sigma_u = \begin{bmatrix} 1-\delta_1^2 & \gamma^\top \\ \gamma & \psi + \gamma\gamma^\top/(1-\delta_1^2) \end{bmatrix}$. This enforces the MNP scale normalization $\Sigma_{11}=1$ by construction because $\Sigma = \Sigma_u + \delta\delta^\top$, and it enforces positive definiteness through $|\delta_1|<1$ and $\psi \succ 0$ via a Schur complement argument. The second key piece is the stochastic representation of the skew-normal distribution, $Z_i \mid w_i \sim N(X_i\beta+\delta w_i, \Sigma_u)$ with $w_i$ a truncated standard normal, which turns the skew-normal utilities into a conditionally Gaussian regression and makes Gibbs updates available. The sampler then needs only a single Metropolis–Hastings step, for $\delta_1$.

What would settle it

For a small simulated three-alternative choice dataset generated with known skewness, compute the likelihood or posterior surface over the intercept and $\delta_1$ with $\Sigma_{11}$ fixed to 1; if a near-flat ridge of high likelihood runs along a curve intercept = $f(\delta_1)$, the reparameterization does not identify skewness separately from the intercept, and posterior recovery of $\delta$ would be prior-driven rather than data-driven.

Watch

Extended reading notes

Core claim

On the paper's terms, the discovery is that the multivariate skew-normal distribution can replace the multivariate normal error in multinomial probit without giving up identification or tractable Bayesian computation. The identification problem is that the MNP scale normalization interacts with the skewness parameters, because the skew-normal error has a nonzero mean that can be absorbed into intercepts. The paper's reparameterization writes the doubly augmented covariance $\Sigma_u$ with top-left entry $1-\delta_1^2$, so that $\Sigma=\Sigma_u+\delta\delta^\top$ automatically has $\Sigma_{11}=1$; positivity follows from $|\delta_1|<1$ and a positive-definite $\psi$. With this parameterization, the double augmentation of latent utilities $Z_i$ and truncation variables $w_i$ gives closed-form full conditionals for all parameters except $\delta_1$, which is updated by a single univariate Metropolis–Hastings step. The empirical message is that asymmetry is present in real purchase data and that ignoring it distorts own- and cross-price elasticities.

Load-bearing premise

The load-bearing premise is that fixing the first diagonal element of the latent utility covariance to one separates intercepts from skewness; the paper asserts this identification but does not prove that no continuum of (intercept, skewness) pairs produces identical choice probabilities.

Editorial extensions

If this is right

  • Standard MNP is a special case at $\delta=0$, so SMNP can be used as a drop-in generalization that reduces to the familiar model when asymmetry is absent.
  • In the paper's applications, allowing asymmetry changes own- and cross-price elasticities: the MNP understates price sensitivity for a focal detergent brand and misallocates substitution across ketchup brands.
  • The sampler is practical for the datasets studied: all full conditionals except $\delta_1$ are closed-form Gibbs steps, with a single univariate Metropolis–Hastings update.
  • When the data generating process is symmetric, SMNP shrinks toward MNP and does not lose predictive accuracy, so the added flexibility has little cost in symmetric settings.
  • Because asymmetry is introduced through the latent utility distribution rather than observed gain/loss covariates, SMNP detects asymmetric responses without requiring reference prices or predictor transformations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the identification concern is real, $\delta$ and alternative-specific intercepts are partially aliased, so posterior skewness should be checked for sensitivity to the intercept prior or to centering the skew-normal error at zero mean.
  • The distributional asymmetry is silent about mechanism: loss aversion, brand loyalty, and switching costs can all produce similar shapes; letting $\delta$ depend on covariates would make mechanisms testable.
  • The double augmentation strategy should carry over to skew-$t$ or related skew-elliptical errors because the conditionally Gaussian representation generalizes, adding tail robustness.
  • A prior-predictive check for symmetry—comparing the observed difference in response to a price increase versus a price decrease with the range implied by the $\tau_\delta$ prior—would separate prior-induced from data-induced asymmetry.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. The paper proposes a skewed multinomial probit (SMNP) model in which latent utilities follow a multivariate skew-normal distribution, allowing asymmetric choice responses. The authors introduce a reparameterization of the covariance matrix that enforces the MNP scale normalization and positive definiteness, a double data augmentation scheme yielding a Metropolis-Hastings within Gibbs sampler, and priors centered on a symmetric equicorrelated structure. Numerical experiments with symmetric and asymmetric DGPs and two empirical applications (laundry detergent and ketchup) are used to argue that SMNP recovers skewness, improves predictive accuracy, and changes implied price elasticities and substitution patterns.

Significance. The paper addresses a real gap in discrete choice modeling: the symmetry restriction of standard multinomial probit. If the identification and computational claims are correct, the SMNP model is a useful, parsimonious alternative to reference-price or random-coefficient approaches to asymmetry. The paper's numerical experiments are carefully designed with known DGPs and oracle comparisons, and the empirical evaluation uses repeated out-of-sample splits, which is a methodological strength. However, the identification proof is a necessary part of the contribution and is currently missing.

major comments (1)
  1. [Section 3.2, Eq. (7)] The claim that the reparameterization 'is key for identification' and that θ is 'the identified parameter vector' is not supported by a formal identifiability proof. The construction in Eq. (7) ensures only that every posterior draw satisfies the scale normalization Σ_11=1 and that Σ_u is positive definite; it does not establish that the map from θ to the multinomial choice probabilities is injective. Because the multivariate skew-normal error has nonzero mean √(2/π)δ, a change in δ can be compensated by an opposite change in the alternative-specific intercepts contained in X_i (Eq. 3) so that the conditional mean of Z_i is unchanged. The paper does not provide a rank condition or an argument showing that the higher-order cumulants of the skew-normal are separately identified from polychotomous choices. Without this, the MCMC sampler is operating on a parameter vector whose posterior may contain ridges, and the reported posterior densities for δ in Figures 2, 4, 5, and 7 cannot be interpreted as conclusive evidence of identification. This is load-bearing because the abstract and Section 6 list 'fully identified' as a main contribution.
minor comments (4)
  1. [Section 5.1] The sentence 'In addition, we also apply the Giacomini–White test directly to the' is incomplete and should be finished or removed.
  2. [Section 1] The word 'difficult' appears in the introduction and should be corrected to 'difficult'.
  3. [Appendix A, Eq. (21)] The expression for the expectation of ψ is written as V/((J+3)−(J−1)−1), which simplifies to V/3; the intermediate notation is confusing and could be simplified.
  4. [Table 2] In the ketchup panel, the CLS3 entries for MNP and SMNP are identical (−0.4778) but the MNP entry is flagged with a star; the note should explain the convention for ties.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: out-of-sample predictions and fixed priors make the central claims independent of fitted inputs; the identification assertion is a rigor gap, not a circular reduction.

full rationale

The paper's derivation chain is self-contained against external benchmarks. The SMNP likelihood is built from the standard multivariate skew-normal stochastic representation in equation (5), and the reparameterization in equation (7) is derived in Appendix A as a Cholesky-based restriction that enforces the scale normalization Sigma_11 = 1 and positive definiteness; it is not fitted to outcome data. The numerical experiments generate data from known DGPs and compare predictive distributions to an oracle, with out-of-sample scoring reported in Table 1, and the empirical applications use repeated 80/20 train/test splits described in Section 5. The prior on the skewness vector delta is centered at zero and fixed before seeing outcomes, so the model does not fit skewness to the quantities later reported as predictions. No fitted parameter is renamed as a prediction, and no load-bearing step is justified solely by a self-citation. The self-citations to Loaiza-Maya and Nibbering (2022, 2023) concern dataset conventions and earlier MNP estimation methods, not the SMNP identification claim. The one notable weakness is that identification of theta = (beta, delta, gamma, psi) is asserted in Section 3.2 rather than proven: equation (7) guarantees that posterior draws satisfy the normalization and positive definiteness, but the paper does not supply a rank or injectivity condition showing that the map from theta to the multinomial choice probabilities is one-to-one. That is a mathematical rigor concern, not a circular step, because no equation in the paper reduces a predicted quantity to an input fitted value or derives the central claim from a self-referential definition. Therefore no circular step can be identified from the text.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim rests on standard distributional facts and on the prior choices. The main unstated background assumptions are posterior propriety and sampler convergence; the main hand-chosen inputs are the prior hyperparameters tau_delta, tau_gamma, and c.

free parameters (3)
  • tau_delta = 0.1936
    Prior variance hyperparameter for the skewness vector delta, chosen by hand to balance shrinkage toward the MNP special case and the ability to detect asymmetry. The empirical results depend on this choice, as tau_delta controls how much prior mass is assigned to asymmetric latent utility distributions.
  • tau_gamma = 0.9 * UB_gamma (formula in Appendix A.2)
    Prior variance hyperparameter for gamma, set as 90% of the upper bound derived from positive definiteness of V. Not fitted to data, but chosen for computational stability.
  • c = 0.995
    Truncation bound for delta_1 to enforce |delta_1| < 1 and thus the scale identification constraint. Chosen for numerical stability.
assumptions (5)
  • standard math The stochastic representation of the multivariate skew-normal distribution in equation (5) holds for delta defined via alpha and Sigma.
    Standard result from Frühwirth-Schnatter & Pyne (2010); the paper relies on it for the double data augmentation.
  • standard math The multivariate skew-normal family is closed under affine transformations, so that Z_i ~ SN_J(X_i beta, Sigma, alpha).
    Used in Section 2.2 to justify the distribution of latent utilities; standard result from Azzalini (2013).
  • domain assumption The model in (1)-(2) with J differenced utilities and threshold zero is observationally equivalent to a multinomial probit with J+1 alternatives and base utility fixed at zero.
    Standard differencing in multinomial probit (Bunch 1991); the paper relies on this to justify the choice rule.
  • domain assumption The prior on (beta, delta, gamma, psi) is proper and yields a proper posterior, so that MCMC draws describe a well-defined posterior distribution.
    The paper does not prove posterior propriety; it assumes proper priors and a bounded likelihood suffice. This is typical but unstated.
  • domain assumption The proposed sampler is geometrically ergodic and produces reliable posterior summaries in the reported runs.
    The paper notes possible slow mixing for large choice sets in Section 6 but assumes the sampler is adequate in the reported settings. No convergence diagnostics are reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Multinomial Probit Model for Asymmetric Choice Responses." pith.science (2026). https://pith.science/paper/G272IU2E

@misc{pith2026260810336,
  author       = {Pith},
  title        = {Pith review of: A Multinomial Probit Model for Asymmetric Choice Responses},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/G272IU2E}},
  note         = {Machine review of arXiv:2608.10336}
}
read the original abstract

Standard multinomial probit (MNP) models specify symmetric latent utility distributions, implying that choice probabilities respond symmetrically to positive and negative covariate shifts of the same magnitude. This restriction is often implausible in empirical choice settings and can lead to misleading elasticity and substitution predictions. We propose a skewed multinomial probit (SMNP) model that captures asymmetric choice responses by specifying a multivariate skew-normal distribution for the latent utilities. The model preserves the flexible substitution patterns of the MNP framework, introduces alternative-specific skewness parameters, and nests the standard MNP model when skewness is zero. Introducing skewness creates identification and computational challenges because the skewness parameters interact with the MNP scale normalization and disrupt the conditional Gaussian updating structure used in Bayesian MNP estimation. We address these challenges through a covariance reparameterization that enforces identification and positive definiteness by construction, interpretable priors on the identified parameter space, and a double data-augmentation scheme that yields a Metropolis-Hastings within Gibbs sampler. Numerical experiments and applications to consumer choice data show that SMNP recovers asymmetric choice responses, improves probabilistic prediction, and produces economically meaningful differences in price elasticities and substitution patterns.

Figures

Figures reproduced from arXiv: 2608.10336 by the authors.

Figure 2
Figure 2. Posterior densities 𝛿1 and 𝛿2 in the SMNP model for the experiment with skewed latent utilities. The vertical lines represent true parameter values. other alternatives held constant at their mean. The choice probabilities from SMNP, MNP, and data generating process (DGP) are represented by the solid yellow, dotted black, and dashed blue lines, respectively. The SMNP model closely tracks the true probability curves. … view at source ↗
Figure 3
Figure 3. Posterior choice probabilities for the experiment with skewed latent utilities. The [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figure 4
Figure 4. Posterior densities 𝛿1 and 𝛿2 in the SMNP model for the experiment with sym￾metric latent utilities. The vertical lines represent true parameter values. the standard symmetry assumption is valid. When the latent utilities are skewed, it re￾covers the asymmetric choice responses and improves predictive accuracy relative to the MNP model. When the latent utilities are symmetric, it shrinks toward the MNP special case … view at source ↗
Figures from the paper (7 more)
Figure 5
Figure 5. Figure 5: Posterior densities for 𝛿1 to 𝛿5 , representing the skewness parameters for detergent brands Tide, EraPlus, Surf, Solo, and All, respectively, with Wisk the base category. SMNP MNP [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: Price elasticity of EraPlus and Solo given the price of EraPlus, with the prices of [PITH_FULL_IMAGE:figures/full_fig_p021_6.png]
Figure 7
Figure 7. Figure 7: Posterior densities for 𝛿1 to 𝛿3 , representing the skewness parameters for ketchup brands Hunts, Del Monte, and STB respectively, with Heinz the base category. SMNP MNP [PITH_FULL_IMAGE:figures/full_fig_p024_7.png]
Figure 8
Figure 8. Figure 8: Price elasticity given the price of Hunts indicated on the x-axis. The solid yellow [PITH_FULL_IMAGE:figures/full_fig_p024_8.png]
Figure 9
Figure 9. Figure 9: Choice probabilities given different product prices indicated on the x-axis. The [PITH_FULL_IMAGE:figures/full_fig_p040_9.png]
Figure 10
Figure 10. Figure 10: Choice probabilities given the price of EraPlus indicated on the x-axis. The [PITH_FULL_IMAGE:figures/full_fig_p041_10.png]
Figure 11
Figure 11. Figure 11: Choice probabilities given the price of Hunts indicated on the x-axis. The solid [PITH_FULL_IMAGE:figures/full_fig_p041_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

4 extracted references · 4 canonical work pages

  1. [1]

    Since 𝐿22 is a lower-triangular matrix with real and positive diagonal entries, we have 𝐿22𝐿⊤ 22 = 𝜓 is a positive-definite matrix

    Let 𝐿21𝐿11 = 𝛾 , so that 𝐿21 = 𝛾 √1−𝛿2 1 . Since 𝐿22 is a lower-triangular matrix with real and positive diagonal entries, we have 𝐿22𝐿⊤ 22 = 𝜓 is a positive-definite matrix. Under this parameterization, Σ𝑢 can be expressed as Σ𝑢 = ⎡ ⎢ ⎣ 1 − 𝛿 2 1 𝛾⊤ 𝛾 𝜓 + 𝛾𝛾 ⊤ 1−𝛿2 1 ⎤ ⎥ ⎦ . (15) 31 This parameterization allows us to place priors directly on (𝛿, 𝛾, 𝜓). S...

  2. [2]

    = 1−𝕍(𝛿 1), which always follows by the fact that 𝔼(𝛿2

  3. [3]

    Second 𝔼(𝛾) = 0.5 × 1 𝐽−1 , which can be easily imposed by setting 𝐵𝛾 = 0.5 × 1𝐽−1

    = 𝕍(𝛿 1). Second 𝔼(𝛾) = 0.5 × 1 𝐽−1 , which can be easily imposed by setting 𝐵𝛾 = 0.5 × 1𝐽−1 . Third, we must impose the condition 𝔼 [𝜓 + 𝛾𝛾 ⊤ 1 − 𝛿 2 1 ] = 0.5 × 1 𝐽−1 1⊤ 𝐽−1 + (0.5 − 𝜏𝛿)𝐼𝐽−1 , (20) where 𝜏𝛿 must be chosen such that the right-hand side of ( 20) is positive definite. It is straightforward to show that the smallest eigenvalue of this matri...

  4. [4]

    Here, the proposal density 𝑞( ̃𝛿1) is constructed via a Laplace approximation to the condi- tional posterior of ̃𝛿1

    Otherwise, the draw is rejected and we retain the previous value, ̃𝛿(𝑟) 1 = ̃𝛿(𝑟−1) 1 . Here, the proposal density 𝑞( ̃𝛿1) is constructed via a Laplace approximation to the condi- tional posterior of ̃𝛿1. Specifically, we numerically maximize the log of posterior ( 30) to obtain the Maximum A Posteriori (MAP) estimate ̃𝛿1𝑀𝐴𝑃 , which corresponds to the mod...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.