REVIEW 3 major objections 6 minor 15 references
Model selection in Gaussian SEMs can be improved by latent-aware criteria: an exact ICL from the closed-form integrated complete-data likelihood, and an importance-sampling approximation of the observed-data likelihood.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 08:34 UTC pith:PZO4ZQQZ
load-bearing objection Closed-form ICL for Gaussian SEM is a real contribution, but the importance-sampling criterion lacks Monte Carlo diagnostics and the 'robust' claim is not yet verified. the 3 major comments →
Information criteria exploiting latent structure for model selection in Structural Equation Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Within the Gaussian SEM class defined by single-loading measurement (each row of the loading matrix has exactly one nonzero entry) and an acyclic latent graph (B strictly lower triangular, Gamma = I_q), the paper establishes that the integrated complete-data likelihood p(X,Z | M) can be computed in closed form under Gaussian and inverse-gamma priors. Proposition 1 gives this formula; replacing the latent variables by their posterior expectation (the MAP estimate) yields the proposed ICL criterion (Eq. 7). Because the complete-data likelihood is tractable, the paper also approximates the integrated observed-data likelihood by importance sampling from the Gaussian conditional distribution p(Z
What carries the argument
The central object is Proposition 1's closed-form integrated complete-data likelihood. It factorizes the SEM into p independent univariate regressions (each observed variable on its single latent variable) and q independent latent regressions (each row of B), so the Gaussian-inverse-gamma priors make every parameter integral analytic. This yields the exact ICL (Integrated Completed Likelihood) criterion. The second object is the importance-sampling estimator of Eq. 8, which reuses the closed-form complete-data likelihood with proposal h(Z)=p(Z|X, theta-hat), a Gaussian whose mean and covariance are given by the conditional of the joint Gaussian model; it is the mechanism that converts the la
Load-bearing premise
The whole derivation assumes that every observed variable loads on exactly one latent variable and that the latent dependency graph is acyclic (B strictly lower triangular with Gamma=I_q); without this, the closed-form ICL and the importance-sampling approximation are not defined as written.
What would settle it
Simulate a Gaussian SEM whose true model has a cross-loading (one observed variable with two nonzero loadings) or a feedback loop in the latent graph, fit the candidate models as in Section 4, and compute the proposed ICL and logIL.IS: if selection rates collapse or the criteria cannot be evaluated, that shows the method does not extend beyond the single-loading acyclic class.
If this is right
- For Gaussian SEMs with single-loading and acyclic latent graphs, the ICL criterion can be evaluated in closed form, avoiding Laplace approximations.
- logIL.IS provides a latent-aware approximation of the integrated observed-data likelihood that is competitive across weak and strong signals and improves on BIC when direct effects are weak.
- ICL is reliable only when the latent variables are estimated accurately; at measurement error sigma^2=0.3 it systematically selects too-simple models even at n=1000.
- The oracle comparisons imply the ICL formula itself is sound: the bottleneck is estimation of continuous latent variables, not the complete-data formulation.
- No single criterion dominates; the simulations suggest the best choice depends on sample size, signal strength, and measurement quality.
Where Pith is reading between the lines
- The paper's oracle experiments suggest a practical fix: instead of plugging one MAP estimate of Z into ICL, one could average the closed-form complete-data likelihood over a posterior sample of Z; this is a natural follow-up that may recover the oracle's performance under high measurement error.
- The closed-form derivation depends on the single-loading measurement model and the acyclic structural model, so the criteria are not immediately usable for cross-loadings or feedback loops; extending Proposition 1 to those cases would be a direct test of the approach's generality.
- LogIL.IS incurs substantial computational cost with R=n samples; a variance-reduced importance sampler or a deterministic quadrature version could make latent-aware selection practical for large model spaces while retaining its weak-signal advantage.
- Since ICL and logIL.IS respond differently to weak direct effects, an adaptive strategy that uses logIL.IS for screening and ICL for confirmation might be better than either criterion alone; this is an editorial suggestion, not a paper claim.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes two model selection criteria for a restricted class of Gaussian structural equation models (SEMs) in which each observed variable loads on exactly one latent variable, the structural coefficient matrix B is strictly lower triangular, and the innovation covariance is the identity. The first criterion, ICL (Eq. 7), plugs maximum-a-posteriori latent estimates into the closed-form integrated complete-data likelihood derived in Proposition 1. The second, logIL.IS (Eq. 8), approximates the integrated observed-data likelihood by importance sampling from the plug-in Gaussian posterior p(Z|X, θ̂). The criteria are compared with AIC, BIC, CAIC, ABIC, HBIC, and IBIC in an extensive simulation study covering four latent structures (null, direct, indirect, complete), sample sizes n=70, 250, 1000, varying signal strengths, and two residual variances. The paper concludes that logIL.IS is robust and competitive across scenarios and that ICL is particularly effective when the latent variables are accurately estimated, a claim supported by an oracle comparison using true latent values.
Significance. If the results hold, the paper contributes a tractable exact integrated complete-data likelihood for an SEM subclass, enabling a direct ICL criterion and a latent-structure-aware approximation of the integrated observed-data likelihood. The proof of Proposition 1 in Appendix A is detailed and appears algebraically sound, and the oracle comparison in Tables 1–3 is a valuable internal control that isolates the effect of latent-variable estimation error. However, the comparative claims rest on simulation evidence whose uncertainty is not quantified, and the importance-sampling estimator is used without any Monte Carlo diagnostics. These issues currently prevent the paper from fully establishing its stated claims.
major comments (3)
- [§3.3, Eq. (8)] The logIL.IS estimator uses the proposal h(Z)=p(Z|X,θ̂), a single Gaussian. In the simulated settings with strong measurement (σ²=0.1), the conditional covariance Υ is typically smaller than the identity, while the target p(X,Z|M) in Proposition 1 decays as exp(−½Σz²) times polynomial factors (e.g., s_j^{−1/2}, det(Š_h)^{−1/2}). The importance ratio is therefore unbounded as ||Z||→∞, and the estimator can have infinite variance. No effective sample size, Monte Carlo error, or weight diagnostics are reported for Eq. (8), and R=n is an arbitrary choice. Without such diagnostics, the claim that logIL.IS is 'robust and competitive' is not substantiated. Please report ESS/weight distributions and a sensitivity analysis in R, or replace the proposal with a heavier-tailed distribution (e.g., Student-t) to ensure finite variance.
- [§4.3, Figures 2–5 and Tables 1–3] All selection rates are proportions out of 100 datasets, yet no standard errors, confidence intervals, or error bars are provided. Differences of 5–10 percentage points—on which several comparative statements rely, such as logIL.IS being 'best' for the complete model with n=70 and b31=0.1 in Figure 3—are within binomial sampling noise. The paper should report binomial confidence intervals or standard errors, and ideally the number of replications should be increased or justified. This is necessary to support the comparative conclusions about relative performance.
- [§4.2] The proposed criteria depend on hand-set hyperparameters (α_j=1, β_j²=1, δ_j=2, κ_h=2I) and the importance-sampling size R=n. No sensitivity analysis is presented, so it is unclear whether the observed behavior of ICL and logIL.IS is robust to reasonable prior changes. At a minimum, the authors should vary δ_j and κ_h (and possibly R) and report the resulting selection rates. Without this, the simulation conclusions may be an artifact of the particular hyperparameter choices.
minor comments (6)
- [§3.2.4] The estimator Ẑ is called the MAP, but it is actually the conditional expectation under p(Z|X,θ̂); the mode of the integrated posterior p(Z|X,M) is generally different. Please clarify the terminology.
- [§4.2] The choice R=n is not justified. Provide a rationale or a small experiment showing stability of logIL.IS as R increases.
- [Appendix A, Lemma 1] The notation ∆_j = diag(δ_1,...,δ_r) conflicts with the scalar δ_j used later. Unify notation between the lemma and its application.
- [Appendix A, proof of Proposition 1] The line 't_j = ... = t_j' repeats the symbol on both sides of the equality and is confusing. Please clean up the derivation.
- [Figures 2–5] The color scheme for methods is described only in the text; add a legend or explicit labels within each figure to improve readability.
- [§2.1 and §5] The paper restricts to no cross-loadings, recursive B, and Γ=I, but the title and several statements refer to SEMs generally. Add a paragraph in the conclusion explicitly stating the scope and the lack of robustness evidence for departures from these assumptions.
Circularity Check
No significant circularity; the criteria are derived analytically and evaluated on independent simulations.
full rationale
The derivation chain is self-contained. Proposition 1 (Section 3.2.4) is an exact analytic integration of the complete-data likelihood under the stated conjugate priors, and the proof in Appendix A is carried out explicitly without invoking any external or self-cited uniqueness theorem. The ICL criterion (Eq. 7) substitutes MAP latent estimates, which is a standard and explicitly acknowledged approximation following Biernacki et al.; the paper even isolates the effect of latent estimation through oracle comparisons in Tables 1-3, concluding that practical ICL's weakness is due to latent-variable estimation rather than to the criterion itself. The logIL.IS criterion (Eq. 8) is a standard importance-sampling identity: it approximates the integrated observed-data likelihood by sampling from h(Z)=p(Z|X,theta-hat), while the target p(X,Z|M) from Proposition 1 does not depend on theta-hat, so the data-dependent proposal does not make the estimator circular. Hyperparameters and R=n are fixed by hand, not fitted to selection outcomes, and the method is benchmarked against external classical criteria on simulated data. The only flagged concerns are technical robustness issues (potential unbounded importance weights, lack of ESS diagnostics) and scope limitations (restrictive identifying assumptions), neither of which is a circularity. No load-bearing self-citations are present; the cited ICL literature is external.
Axiom & Free-Parameter Ledger
free parameters (2)
- Prior hyperparameters (alpha_j, beta_j^2, nu_j, delta_j, mu_h, kappa_h) =
alpha=1, beta^2=1, nu=0, delta=2, mu=0, kappa=2I
- Importance-sampling sample size R =
R = n
axioms (6)
- domain assumption Gaussian linear SEM: X_i = Lambda Z_i + epsilon_i, Z_i = (I-B)^{-1} xi_i, with Gaussian errors and Gamma = I_q
- domain assumption No cross-loadings: each row of Lambda has exactly one non-zero coefficient
- domain assumption Identifiability: B is strictly lower triangular and Gamma = I_q
- ad hoc to paper Conjugate priors with hyperparameters fixed by hand
- ad hoc to paper MAP plug-in for latent variables in ICL
- ad hoc to paper Importance-sampling proposal h(Z) = p(Z|X, theta-hat) with R = n draws
read the original abstract
Structural equation models (SEM) are widely used to describe dependency structures between latent variables, making model selection a key issue in many applications. Existing information criteria are generally based on the integrated observed-data likelihood and therefore do not explicitly account for the latent structure of the model. In this paper, we propose two new information criteria derived from the integrated complete-data likelihood. The first adapts the Integrated Completed Likelihood criterion to Gaussian SEM, while the second proposes an alternative approach to approximating the integrated observed-data log-likelihood by incorporating latent structural information and using an importance sampling strategy. Their performance is assessed through an extensive simulation study covering null, direct, indirect and complete latent structures under different sample sizes and signal strengths. The results show that the proposed importance sampling strategy provides robust and competitive model selection across a wide range of scenarios, whereas the proposed ICL criterion is particularly effective for recovering latent dependency structures when the latent variables are accurately estimated. These findings demonstrate the potential benefits of explicitly exploiting the latent structure when developing information criteria for structural equation models.
Figures
Reference graph
Works this paper leans on
-
[1]
Akaike, H. (1974). A new look at the statistical model identification.IEEE transactions on automatic control, 19(6):716–723
1974
-
[2]
Biernacki, C., Celeux, G., and Govaert, G. (2000). Assessing a mixture model for clustering with the integrated completed likelihood. IEEE transactions on pattern analysis and machine intelligence, 22(7):719–725
2000
-
[3]
Biernacki, C., Celeux, G., and Govaert, G. (2010). Exact and monte carlo calculations of integrated likelihoods for the latent class model.Journal of Statistical Planning and Inference, 140(11):2991–3002
2010
-
[4]
A., Harden, J
Bollen, K. A., Harden, J. J., Ray, S., and Zavisca, J. (2014). Bic and alternative bayesian information criteria in the selection of structural equation models.Structural equation modeling: a multidisciplinary journal, 21(1):1–19
2014
-
[5]
A., Ray, S., Zavisca, J., and Harden, J
Bollen, K. A., Ray, S., Zavisca, J., and Harden, J. J. (2012). A comparison of bayes factor approximation methods including two new methods.Sociological Methods & Research, 41(2):294–324
2012
-
[6]
Bozdogan, H. (1987). Model selection and akaike’s information criterion (aic): The general theory and its analytical extensions. Psychometrika, 52(3):345–370. 15 Model selection in Structural Equation Models (SEM)
1987
-
[7]
J., Coffman, D
Dziak, J. J., Coffman, D. L., Lanza, S. T., Li, R., and Jermiin, L. S. (2020). Sensitivity and specificity of information criteria. Briefings in bioinformatics, 21(2):553–565
2020
-
[8]
Haughton, D. M. (1988). On the choice of a model to fit data from an exponential family.The annals of statistics, pages 342–355
1988
-
[9]
M., Oud, J
Haughton, D. M., Oud, J. H., and Jansen, R. A. (1997). Information and other criteria in structural equation model selection. Communications in Statistics-Simulation and Computation, 26(4):1477–1516. Jöreskog, K. G. (1970). A general method for estimating a linear structural equation system.ETS Research Bulletin Series, 1970(2):i–41
1997
-
[10]
Lin, L.-C., Huang, P.-H., and Weng, L.-J. (2017). Selecting Path Models in SEM: A Comparison of Model Selection Criteria. Structural Equation Modeling: A Multidisciplinary Journal, 24(6):855–869
2017
-
[11]
Preacher, K. J. and Yaremych, H. E. (2023). Model selection in structural equation modeling.Handbook of structural equation modeling, pages 206–222
2023
-
[12]
Raftery, A. E. (1995). Bayesian model selection in social research.Sociological methodology, pages 111–163
1995
-
[13]
Schwarz, G. (1978). Estimating the dimension of a model.The annals of statistics, pages 461–464
1978
-
[14]
Sclove, S. L. (1987). Application of model-selection criteria to some problems in multivariate analysis. Psychometrika, 52(3):333–343
1987
-
[15]
Yang, C.-C. (2006). Evaluating latent class analysis models in qualitative phenotype identification. Computational statistics & data analysis, 50(4):1090–1104. 16 Model selection in Structural Equation Models (SEM) Appendices Appendix A Proof of Proposition 1 Proof of Proposition 1.Note that p(X,Z|ϑ,f,ω) = ∫ p(X,Z|θ;f,ω)p(θ|ϑ;f,ω)dθ = ∫ p(X|Z,Λ,Σ)p(Λ,Σ|ϑ,...
2006
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.