REVIEW 3 major objections 5 minor 40 references
One MCMC run now learns how many latent traits an ordinal questionnaire needs.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 17:22 UTC pith:RZYA7CIN
load-bearing objection A credible COSS-plus-Albert-Chib dimension-selection sampler for ordinal probit MGRMs, with a real gap in how the lower-triangular identifiability constraint is handled. the 3 major comments →
An efficient adaptive dimension selection algorithm for multidimensional probit graded response models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that a single overfitted probit graded response model, equipped with a cumulative ordered spike-and-slab (COSS) prior on the loading-matrix column variances, can identify the effective number of latent dimensions while estimating all parameters, so that post-hoc model comparison across candidate dimensions becomes unnecessary. The stochastic ordering property of the COSS prior guarantees that later dimensions are increasingly likely to have near-zero loadings, which the authors prove for both the exact spike and the continuous spike approximation. The adaptive truncation strategy lets the sampler discard inactive columns during MCMC, and the real-data analysis re
What carries the argument
The cumulative ordered spike-and-slab (COSS) prior assigned to the column-specific variances θ_h of the loading matrix. It is built from a non-homogeneous stick-breaking construction with nondecreasing spike probabilities π_1 ≤ ... ≤ π_K, so higher-indexed dimensions are increasingly likely to be assigned to the near-zero spike variance θ_0. The prior's stochastic ordering (Proposition 1 and Corollary 1) is what turns the loading matrix into a soft dimension-selection mechanism, and it is combined with Albert–Chib latent response augmentation to keep every full conditional Gaussian and the sampler closed-form.
Load-bearing premise
The lower-triangular identifiability constraint on the loading matrix B is stated in the text but never incorporated into the prior or the Gaussian full conditionals, so the implemented sampler may not be targeting the posterior of the model as specified.
What would settle it
Reproduce the sampler exactly as described in Section 3.2 (without lower-triangular constraints) on a dataset with a known factor structure, then run the version with the lower-triangular constraint enforced by zeroing sampled entries; if the posterior distribution of the loading matrix and the effective dimension differ substantially, the identifiability constraint is not being handled by the stated updates.
If this is right
- Analysts can obtain a posterior distribution over the number of latent dimensions instead of a single selected model, propagating dimensionality uncertainty into estimates of loadings and traits.
- The method removes the need to fit multiple fixed-dimensional models and compute AIC, BIC, DIC, WAIC, or cross-validation, which the paper shows are less stable and more biased in small samples.
- The adaptive truncation contract/expand mechanism means computation focuses on active dimensions, so a conservatively large initial truncation level does not impose a permanent computational cost.
- The approach transfers to other ordinal item response data, such as educational assessments and psychometric questionnaires, where the number of constructs is not known a priori.
Where Pith is reading between the lines
- The COSS prior could be adapted to other polytomous or mixed-format item response models, since the augmentation and stick-breaking machinery do not depend on the specific probit link; a logit link would require a different augmentation strategy.
- The loading matrix's ordered shrinkage may provide a data-driven way to sort latent dimensions by explanatory importance, potentially easing factor rotation and interpretation problems beyond dimension selection.
- A testable extension is to compare the adaptive sampler against reversible-jump MCMC or variational methods on a benchmark suite of sparse ordinal datasets, checking whether the accuracy gains persist when the true dimension is large relative to the number of items.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a Bayesian adaptive dimension-selection method for multidimensional probit graded response models. Item loadings are assigned a cumulative ordered spike-and-slab (COSS) prior on their column-specific variances, and the effective number of latent dimensions is inferred through the posterior distribution of spike/slab allocations. Inference is carried out with an Albert–Chib data-augmentation Gibbs sampler, supplemented by an adaptive truncation scheme that adds or removes latent dimensions during the run. The paper reports simulation evidence that the method recovers the true dimensionality and parameters, compares favorably with AIC/BIC/DIC/WAIC/cross-validation, and illustrates the approach on IPIP-NEO personality data, where it selects five factors in 24 of 25 runs.
Significance. If the central claims are correct, the paper offers a useful single-run Bayesian alternative to sequential model fitting for dimension selection in ordinal item response models. The COSS prior is a sensible choice for ordered shrinkage, and the stochastic-ordering propositions (Prop. 1 and Cor. 1) are correctly derived and provide the intended prior justification. The conditional updates for the unconstrained model are mostly standard and the auxiliary-variable representation for the spike-and-slab indicator is elegant. However, the manuscript currently leaves two load-bearing implementation issues unresolved: the lower-triangular identifiability constraint is asserted but never incorporated into the prior or the Gibbs conditionals, and the adaptive truncation step is not shown to preserve a well-defined target distribution. These issues affect the interpretation of every simulation and real-data result. The comparison with post-hoc criteria is also confounded by nonstandard parameter counting and different prior/identifiability settings. With substantial revision, the method could be a valuable contribution; in its present form, the central empirical claim is not yet established.
major comments (3)
- [Section 3.1, Eq. (9); Section 5] The lower-triangular identifiability constraint is stated but never implemented in the prior or the conditional updates. Eq. (9) is the full conditional for an unrestricted loading vector β_j with prior covariance Σ_B=diag(θ_1,...,θ_K). Under a lower-triangular restriction (β_jh=0 for h>j), the regression of Y*_{·j} on Z uses only the first min(j,K) columns, so the conditional covariance/mean are not (Σ_B^{-1}+Z^T Z)^{-1}(-Z^T Y*). If the constraint is enforced by zeroing sampled entries, the chain is not a Gibbs sampler for either the restricted or unrestricted posterior. If it is a hard prior, Eq. (9), the θ_h update (17), and the ρ_h update (14) must be re-derived with the restricted design matrix; they are not. Section 5 asserts the constraint is 'sufficient under our theoretical framework,' but no theoretical framework or proof is provided. Because Table 1, Figure 1, and Section 5 a
- [Section 3.3, Algorithm 1] The adaptive truncation step removes *all* inactive columns of B, wherever they occur, and re-indexes the remaining columns. This is incompatible with the ordered COSS prior: the stick-breaking spike probabilities π_h depend on the absolute index h, and after deletion the retained active dimensions have new indices and new prior weights. The paper does not specify how ρ, θ, and ω are reassigned under the re-indexing, and no stationarity correction (e.g., a Metropolis-Hastings step) is given. This is not the same as the trailing-component truncation in Bhattacharya and Dunson (2011) and Legramanti et al. (2020), and it means Algorithm 1 is not an exact Gibbs sampler for a single well-defined posterior. The claim at the end of Section 3.2 that the sampler 'exactly' targets the posterior is therefore not established.
- [Section 4.2, Fig. 1 and AIC/BIC definitions] The comparison with post-hoc criteria is confounded. The parameter count d=nH+qH+q(C-1) treats the nH latent scores as free parameters in AIC/BIC, even though the competing models are estimated with random latent traits; this penalization is not the usual marginal-likelihood effective parameter count and will systematically disfavor larger H. Additionally, the candidate models use independent N(0,1) priors and no identifiability constraint, whereas the proposed method uses the COSS prior and an (unspecified) lower-triangular constraint. Differences in Fig. 1 therefore conflate prior specification and identifiability choices with the dimension-selection mechanism. The claim of a robust accuracy advantage over AIC/BIC/CV is not cleanly supported by this comparison.
minor comments (5)
- [Section 4.1] Notation for the adaptation probability is inconsistent: the sampler is introduced with parameters α0, α1 in Section 3.3, but Section 4.1 sets 'pη0, η1q=(−1,−5×10^−4)'. Please unify.
- [Section 3.2] The text says the sampler proceeds by 'exactly' sampling from full conditionals (8)–(17), but Algorithm 1 later includes adaptive truncation steps that are not Gibbs updates. Rephrase to avoid the implication that every step is a standard full-conditional update.
- [Section 2.2.1] In Proposition 1, since θ_h is a variance, |θ_h| is redundant; stating probabilities on θ_h itself would be cleaner. This is not an error but makes the proposition read oddly.
- [Throughout] The paper uses 'W AIC' instead of the standard 'WAIC'. Also, the notation t is used both for response category and MCMC iteration; consider distinguishing them.
- [Figure 1] The panel labels in the caption (a)–(f) do not always match the settings described in the surrounding text. Please ensure the captions and in-figure labels are consistent.
Circularity Check
No circularity: prior, sampler, and validation are self-contained; the lower-triangular identifiability gap is a correctness concern, not a circular step.
full rationale
The derivation chain is self-contained. The COSS prior (Sec. 2.2), Proposition 1, Corollary 1, and the Gibbs conditional updates (8)-(17) are derived from the stated probit GRM likelihood and prior by direct algebra; no target quantity is defined in terms of the method's output, and no fitted parameter is renamed as a prediction. Dimension recovery is evaluated on synthetic data generated from an independent probit GRM with nonzero loadings drawn from U(0.25,1.25), not from the COSS posterior, so the reported Acc/MAB results are external benchmarks rather than by-construction outcomes. The IPIP-NEO application is also an external dataset. The only notable weakness is the lower-triangular identifiability constraint mentioned in Sec. 3.1 ('lower-triangular constraints are imposed on B during the sampling process') and Sec. 5 ('sufficient under our theoretical framework'), which is never incorporated into Eq. (9) or the COSS prior; this is a correctness/implementation concern about the sampler's target distribution, not a circular reduction of a prediction to a fitted input. There are no load-bearing self-citations: the cited cumulative shrinkage prior (Legramanti et al. 2020) and data-augmentation approach (Albert and Chib 1993) are external prior work.
Axiom & Free-Parameter Ledger
free parameters (5)
- κ (stick-breaking parameter for v1) =
not specified in paper
- a (stick-breaking rate for l≥2) =
2 (simulation); not specified in real data
- θ0 (spike variance) =
0.05
- aθ, bθ (slab inverse-gamma hyperparameters) =
aθ=bθ=2 in simulation
- Adaptation decay parameters α0, α1 (η0, η1 in Section 4) =
α0=-1, α1=-5e-4 (as η0/η1)
axioms (5)
- domain assumption COSS prior construction with stick-breaking weights ω_l and latent allocation ρ_h is a valid representation of the ordered spike-and-slab prior.
- standard math Albert–Chib latent data augmentation with unit-variance latent responses and ordered thresholds is valid for ordinal probit models.
- domain assumption Adaptive truncation with diminishing adaptation probability converges to the target posterior.
- ad hoc to paper Lower-triangular constraint on B is sufficient for identifiability and compatible with the COSS prior.
- standard math z_i ~ N(0, I_K) fixes the scale of the latent space.
read the original abstract
Multidimensional graded response models (MGRMs) are widely used for analyzing ordinal questionnaire data in psychological and educational assessments. A central challenge in applying these models is determining the number of latent dimensions. Conventional approaches usually fit multiple fixed-dimensional models and select among them using post-hoc criteria such as AIC, BIC, or cross-validation, which can be computationally demanding and ignore uncertainty in dimensionality during estimation. We develop an adaptive Bayesian dimension selection framework for probit MGRMs. Building on the cumulative shrinkage process, we assign a cumulative ordered spike-and-slab (COSS) prior to the column-specific variances of the item loading matrix. This prior induces increasing shrinkage across latent dimensions, allowing redundant dimensions to be shrunk toward zero while preserving flexibility for active dimensions. Albert--Chib latent response augmentation is used to handle the ordinal probit likelihood, yielding conditionally Gaussian updates for item loadings and latent traits. These updates are combined with Gibbs updates for threshold and shrinkage parameters in an efficient adaptive sampler. Simulation studies evaluate the proposed method in terms of dimension recovery, parameter estimation accuracy, and computational efficiency, with comparisons to conventional fixed-dimensional estimation and model selection procedures. The results show that the proposed approach accurately recovers the latent structure while avoiding repeated model fitting over multiple candidate dimensions. We further illustrate the method using real psychological assessment data, demonstrating its practical utility for uncovering interpretable latent structures in ordinal item responses.
Figures
Reference graph
Works this paper leans on
-
[1]
Psychometrika monograph supplement , volume=
Estimation of latent ability using a response pattern of graded scores , author=. Psychometrika monograph supplement , volume=. 1969 , publisher=
1969
-
[2]
Psychometrika , volume=
Marginal maximum likelihood estimation of item parameters: Application of an EM algorithm , author=. Psychometrika , volume=. 1981 , publisher=
1981
-
[3]
Journal of Educational and Behavioral Statistics , volume=
A straightforward approach to Markov chain Monte Carlo methods for item response models , author=. Journal of Educational and Behavioral Statistics , volume=. 1999 , publisher=
1999
-
[4]
Psychometrika , volume=
High-dimensional exploratory item factor analysis by a Metropolis-Hastings Robbins-Monro algorithm , author=. Psychometrika , volume=. 2010 , publisher=
2010
-
[5]
Psychometrika , volume=
A rationale and test for the number of factors in factor analysis , author=. Psychometrika , volume=. 1965 , publisher=
1965
-
[6]
Educational and Psychological Measurement , volume=
A comparison of factor retention methods with dichotomous data , author=. Educational and Psychological Measurement , volume=. 2011 , publisher=
2011
-
[7]
Applied Psychological Measurement , volume=
IRT model selection methods for dichotomous items , author=. Applied Psychological Measurement , volume=. 2007 , publisher=
2007
-
[8]
Journal of the American Statistical Association , volume=
Sparse Bayesian multidimensional item response theory , author=. Journal of the American Statistical Association , volume=. 2025 , publisher=
2025
-
[9]
Journal of the American Statistical Association , volume=
Fast Bayesian factor analysis via automatic rotations to sparsity , author=. Journal of the American Statistical Association , volume=. 2016 , publisher=
2016
-
[10]
arXiv preprint arXiv:1804.04231 , year=
Sparse Bayesian factor analysis when the number of factors is unknown , author=. arXiv preprint arXiv:1804.04231 , year=
-
[11]
Biometrika , volume=
Sparse Bayesian infinite factor models , author=. Biometrika , volume=. 2011 , publisher=
2011
-
[12]
Biometrika , volume=
Bayesian cumulative shrinkage for infinite factorizations , author=. Biometrika , volume=. 2020 , publisher=
2020
-
[13]
Biometrika , volume=
Generalized infinite factorization models , author=. Biometrika , volume=. 2022 , publisher=
2022
-
[14]
Journal of the American Statistical Association , volume=
Bayesian analysis of binary and polychotomous response data , author=. Journal of the American Statistical Association , volume=. 1993 , publisher=
1993
-
[15]
Applied Psychological Measurement , volume=
The past and future of multidimensional item response theory , author=. Applied Psychological Measurement , volume=. 1997 , publisher=
1997
-
[16]
2009 , publisher=
Multidimensional item response theory , author=. 2009 , publisher=
2009
-
[17]
1999 , publisher=
Test theory: A unified treatment , author=. 1999 , publisher=
1999
-
[18]
Handbook of modern item response theory , pages=
Graded response model , author=. Handbook of modern item response theory , pages=. 1997 , publisher=
1997
-
[19]
Psychometrika , volume=
A Rasch model for partial credit scoring , author=. Psychometrika , volume=. 1982 , publisher=
1982
-
[20]
Applied Psychological Measurement , volume=
A generalized partial credit model: Application of an EM algorithm , author=. Applied Psychological Measurement , volume=. 1992 , publisher=
1992
-
[21]
Structural Equation Modeling , volume=
Factor analysis with ordinal indicators: A Monte Carlo study comparing DWLS and ULS estimation , author=. Structural Equation Modeling , volume=. 2009 , publisher=
2009
-
[22]
2010 , publisher=
Bayesian item response modeling: Theory and applications , author=. 2010 , publisher=
2010
-
[23]
2016 , publisher=
Bayesian psychometric modeling , author=. 2016 , publisher=
2016
-
[24]
Psychometrika , volume=
Regularized latent class analysis with application in cognitive diagnosis , author=. Psychometrika , volume=. 2017 , publisher=
2017
-
[25]
Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume=
Asymptotic behaviour of the posterior distribution in overfitted mixture models , author=. Journal of the Royal Statistical Society: Series B (Statistical Methodology) , volume=. 2011 , publisher=
2011
-
[26]
Biometrika , pages=
A spike-and-slab prior for dimension selection in generalized linear network eigenmodels , author=. Biometrika , pages=. 2025 , publisher=
2025
-
[27]
Applied Psychological Measurement , volume=
Full-information item factor analysis , author=. Applied Psychological Measurement , volume=. 1988 , publisher=
1988
-
[28]
Journal of Educational Statistics , volume=
Bayesian estimation of normal ogive item response curves using Gibbs sampling , author=. Journal of Educational Statistics , volume=. 1992 , publisher=
1992
-
[29]
Psychometrika , volume=
MCMC estimation and some model-fit analysis of multidimensional IRT models , author=. Psychometrika , volume=. 2001 , publisher=
2001
-
[30]
Journal of Educational and Behavioral Statistics , volume=
Metropolis-Hastings Robbins-Monro algorithm for confirmatory item factor analysis , author=. Journal of Educational and Behavioral Statistics , volume=. 2010 , publisher=
2010
-
[31]
Journal of Statistical Software , volume=
mirt: A multidimensional item response theory package for the R environment , author=. Journal of Statistical Software , volume=. 2012 , publisher=
2012
-
[32]
Biometrika , volume=
Reversible jump Markov chain Monte Carlo computation and Bayesian model determination , author=. Biometrika , volume=. 1995 , publisher=
1995
-
[33]
Journal of the Royal Statistical Society: Series B (Methodological) , volume=
Bayesian model choice via Markov chain Monte Carlo methods , author=. Journal of the Royal Statistical Society: Series B (Methodological) , volume=. 1995 , publisher=
1995
-
[34]
The Annals of Statistics , volume=
Bayesian analysis of mixture models with an unknown number of components-an alternative to reversible jump methods , author=. The Annals of Statistics , volume=. 2000 , publisher=
2000
-
[35]
Journal of Computational and Graphical Statistics , volume=
Annealed importance sampling reversible jump MCMC algorithms , author=. Journal of Computational and Graphical Statistics , volume=. 2013 , publisher=
2013
-
[36]
Journal of Computational and Graphical Statistics , volume=
Non-reversible jump algorithms for Bayesian nested model selection , author=. Journal of Computational and Graphical Statistics , volume=. 2021 , publisher=
2021
-
[37]
International Conference on Artificial Intelligence and Statistics , pages=
Transport reversible jump proposals , author=. International Conference on Artificial Intelligence and Statistics , pages=. 2023 , organization=
2023
-
[38]
Journal of Applied Probability , volume=
Coupling and ergodicity of adaptive Markov chain Monte Carlo algorithms , author=. Journal of Applied Probability , volume=. 2007 , publisher=
2007
-
[39]
The NEO Personality Inventory Manual , author =
-
[40]
Journal of Research in Personality , volume =
Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120 , author =. Journal of Research in Personality , volume =. 2014 , publisher =
2014
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.