{"id":"cc754c53-70f8-4af0-b8d8-cb751391f938","arxiv_id":"2607.21053","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Two new criteria based on integrated complete-data likelihood are introduced for Gaussian SEM model selection; simulations show one is broadly competitive while the other succeeds only when latent variables are well estimated.","lead":"This paper proposes two new scoring rules for choosing between structural equation models, in which hidden variables connect observed measurements. In computer simulations, one rule makes reliable choices across many settings, while the other works well only when the hidden variables are estimated precisely.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Importance-sampling estimator in Eq. (8) may be unstable: plug-in Gaussian proposal has lighter tails than the integrated posterior, so weights can be unbounded; no ESS/Monte Carlo diagnostics are reported, leaving logIL.IS's 'robust' claim unverified.","rationale":"The reader's weakest assumption is about the restrictiveness of the SEM framework (single loading per observed variable, triangular B, Γ=I) and the lack of evidence for extensions. That is a scope limitation, not an internal flaw. The more load-bearing concern is internal to the proposed logIL.IS criterion: the importance sampling target and proposal are mismatched in tail behavior, which can make the estimator have infinite variance. This threatens the paper's central claim that logIL.IS is robust and competitive even within the assumed model class. The paper reports no Monte Carlo diagnostics (ESS, variance of log-likelihood estimates, or comparisons against exact integration), so the simulation results do not currently establish the reliability of logIL.IS. The ICL criterion is somewhat separate and the oracle experiment gives it support, but the logIL.IS claim is the principal novelty. The concern is concrete and testable; until the test is run, the appropriate verdict remains conditional, as the reader concluded, but for a different reason.","tokens_in":17615,"tokens_out":12977,"duration_ms":133848,"concrete_test":"In the toy model with q=p=1 and B=0, compute the exact p(X|M) by numerical quadrature over Z (or closed-form t density), then run logIL.IS with R=n on 100 datasets and compare the Monte Carlo estimates to the exact values. Also compute the effective sample size ESS=(Σw)²/Σw² for each run. If the logIL.IS estimate has bias larger than 0.5 log-units or if ESS is often below 1% of R, the unbounded-weight concern is confirmed and the simulation results require re-examination.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that logIL.IS is robust rests on the accuracy of the importance sampling approximation in Eq. (8). The proposal is h(Z)=p(Z|X,θ̂), a single Gaussian. The target integrand p(X,Z|M) is the closed-form integrated complete-data likelihood from Proposition 1. After normalization in Z, the induced posterior p(Z|X,M) is a scale mixture of Gaussians (Student-t-like) because Proposition 1 integrates over the variance parameters and regression coefficients. Its tails are heavier than those of the plug-in Gaussian: for large ||Z||, p(X,Z|M) decays as exp(-½||Z||²)·O(||Z||^{-m}) from the s_j^{-1/2} factors, while p(Z|X,θ̂) decays as exp(-½ Z^T Υ̂^{-1} Z) with Υ̂^{-1} > I in the usual identified SEM (conditional covariance is smaller than marginal identity). Thus the importance ratio w(Z)=p(X,Z|M)/h(Z) is unbounded as ||Z||→∞, and the estimator can have infinite variance. No effective sample size, Monte Carlo error, or comparison with an independent integration is reported; R=n is an arbitrary choice. If the weights are dominated by rare extreme draws, the logIL.IS values used in Figures 2–5 may be noisy or biased, so the reported selection rates are not reliable evidence for the 'robust and competitive' claim.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two model selection criteria for a restricted class of Gaussian structural equation models (SEMs) in which each observed variable loads on exactly one latent variable, the structural coefficient matrix B is strictly lower triangular, and the innovation covariance is the identity. The first criterion, ICL (Eq. 7), plugs maximum-a-posteriori latent estimates into the closed-form integrated complete-data likelihood derived in Proposition 1. The second, logIL.IS (Eq. 8), approximates the integrated observed-data likelihood by importance sampling from the plug-in Gaussian posterior p(Z|X, θ̂). The criteria are compared with AIC, BIC, CAIC, ABIC, HBIC, and IBIC in an extensive simulation study covering four latent structures (null, direct, indirect, complete), sample sizes n=70, 250, 1000, varying signal strengths, and two residual variances. The paper concludes that logIL.IS is robust and competitive across scenarios and that ICL is particularly effective when the latent variables are accurately estimated, a claim supported by an oracle comparison using true latent values.","tokens_in":18068,"tokens_out":7293,"duration_ms":80460,"significance":"If the results hold, the paper contributes a tractable exact integrated complete-data likelihood for an SEM subclass, enabling a direct ICL criterion and a latent-structure-aware approximation of the integrated observed-data likelihood. The proof of Proposition 1 in Appendix A is detailed and appears algebraically sound, and the oracle comparison in Tables 1–3 is a valuable internal control that isolates the effect of latent-variable estimation error. However, the comparative claims rest on simulation evidence whose uncertainty is not quantified, and the importance-sampling estimator is used without any Monte Carlo diagnostics. These issues currently prevent the paper from fully establishing its stated claims.","major_comments":[{"comment":"The logIL.IS estimator uses the proposal h(Z)=p(Z|X,θ̂), a single Gaussian. In the simulated settings with strong measurement (σ²=0.1), the conditional covariance Υ is typically smaller than the identity, while the target p(X,Z|M) in Proposition 1 decays as exp(−½Σz²) times polynomial factors (e.g., s_j^{−1/2}, det(Š_h)^{−1/2}). The importance ratio is therefore unbounded as ||Z||→∞, and the estimator can have infinite variance. No effective sample size, Monte Carlo error, or weight diagnostics are reported for Eq. (8), and R=n is an arbitrary choice. Without such diagnostics, the claim that logIL.IS is 'robust and competitive' is not substantiated. Please report ESS/weight distributions and a sensitivity analysis in R, or replace the proposal with a heavier-tailed distribution (e.g., Student-t) to ensure finite variance.","section":"§3.3, Eq. (8)"},{"comment":"All selection rates are proportions out of 100 datasets, yet no standard errors, confidence intervals, or error bars are provided. Differences of 5–10 percentage points—on which several comparative statements rely, such as logIL.IS being 'best' for the complete model with n=70 and b31=0.1 in Figure 3—are within binomial sampling noise. The paper should report binomial confidence intervals or standard errors, and ideally the number of replications should be increased or justified. This is necessary to support the comparative conclusions about relative performance.","section":"§4.3, Figures 2–5 and Tables 1–3"},{"comment":"The proposed criteria depend on hand-set hyperparameters (α_j=1, β_j²=1, δ_j=2, κ_h=2I) and the importance-sampling size R=n. No sensitivity analysis is presented, so it is unclear whether the observed behavior of ICL and logIL.IS is robust to reasonable prior changes. At a minimum, the authors should vary δ_j and κ_h (and possibly R) and report the resulting selection rates. Without this, the simulation conclusions may be an artifact of the particular hyperparameter choices.","section":"§4.2"}],"minor_comments":[{"comment":"The estimator Ẑ is called the MAP, but it is actually the conditional expectation under p(Z|X,θ̂); the mode of the integrated posterior p(Z|X,M) is generally different. Please clarify the terminology.","section":"§3.2.4"},{"comment":"The choice R=n is not justified. Provide a rationale or a small experiment showing stability of logIL.IS as R increases.","section":"§4.2"},{"comment":"The notation ∆_j = diag(δ_1,...,δ_r) conflicts with the scalar δ_j used later. Unify notation between the lemma and its application.","section":"Appendix A, Lemma 1"},{"comment":"The line 't_j = ... = t_j' repeats the symbol on both sides of the equality and is confusing. Please clean up the derivation.","section":"Appendix A, proof of Proposition 1"},{"comment":"The color scheme for methods is described only in the text; add a legend or explicit labels within each figure to improve readability.","section":"Figures 2–5"},{"comment":"The paper restricts to no cross-loadings, recursive B, and Γ=I, but the title and several statements refer to SEMs generally. Add a paragraph in the conclusion explicitly stating the scope and the lack of robustness evidence for departures from these assumptions.","section":"§2.1 and §5"}],"recommendation":"major_revision","confidential_remarks":"The importance-sampling tail concern raised in the stress test is legitimate and should be addressed with diagnostics rather than dismissed. The central derivation in Proposition 1 appears sound, and the oracle experiment is a strong feature. The main obstacles to publication are the missing uncertainty quantification for the simulation results and the unexamined stability of Eq. (8); both are fixable within the manuscript's scope, so major revision rather than rejection seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jan,\n\nHere's my read of arXiv:2607.21053. The main contribution is real: Proposition 1 gives a closed-form integrated complete-data likelihood for Gaussian SEMs with conjugate priors and a single-loading structure, and the resulting exact ICL is new for this setting. The importance-sampling approximation of the integrated observed likelihood is a sensible idea, and the oracle experiment—comparing ICL using estimated latents against ICL using true latents—is a good internal control that isolates the effect of latent estimation. The paper is honest about the cost: ICL degrades when latent estimates are poor, and the simulations show this clearly.\n\nThe soft spots are mostly empirical. The importance-sampling estimator in Eq. (8) has a potential tail problem: the target p(X,Z|M) has heavier-than-Gaussian tails because Proposition 1 integrates over variance and regression parameters, while the proposal h(Z) is a single Gaussian fit at the MLE. If the proposal is lighter-tailed than the target in any direction, the importance weights can be unbounded and the estimator may have infinite variance. The paper doesn't report effective sample size, Monte Carlo error, or any comparison against an independent integration, so the 'robust and competitive' claim for logIL.IS is not yet supported. R=n is chosen without justification. These are fixable by adding diagnostics and maybe a fatter-tailed proposal.\n\nOther gaps: the simulations cover only the single-loading, DAG case with diagonal Gamma; selection rates have no error bars; the hyperparameters are fixed without sensitivity analysis; and there's no real-data application or code/data. None of these are load-bearing flaws in the math—the derivation in Appendix A checks out at a line-by-line level—but they limit the strength of the empirical conclusions.\n\nWho benefits: anyone working on SEM model selection, especially in psychometrics or ecology, where latent DAGs are natural. The paper is for a methods audience and would benefit from being read carefully by a referee who knows importance sampling.\n\nMy verdict: send it to peer review, but with a requirement that the authors report ESS and Monte Carlo error for logIL.IS, and better justify R. The ICL result alone is worth publishing; the logIL.IS claim needs more support.","headline":"Closed-form ICL for Gaussian SEM is a real contribution, but the importance-sampling criterion lacks Monte Carlo diagnostics and the 'robust' claim is not yet verified.","tokens_in":18478,"tokens_out":3223,"would_cite":true,"duration_ms":33508,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Model selection in Gaussian SEMs can be improved by latent-aware criteria: an exact ICL from the closed-form integrated complete-data likelihood, and an importance-sampling approximation of the observed-data likelihood.","keywords":["structural equation models","model selection","latent variables","integrated complete-data likelihood","ICL criterion","importance sampling","Gaussian SEM","latent dependency structure"],"falsifier":"Simulate a Gaussian SEM whose true model has a cross-loading (one observed variable with two nonzero loadings) or a feedback loop in the latent graph, fit the candidate models as in Section 4, and compute the proposed ICL and logIL.IS: if selection rates collapse or the criteria cannot be evaluated, that shows the method does not extend beyond the single-loading acyclic class.","tokens_in":17562,"feed_emoji":"📊","tokens_out":7477,"duration_ms":73787,"temperature":0.7,"pith_summary":"The paper asks whether model selection in Gaussian structural equation models (SEMs) can be improved by considering the integrated complete-data likelihood—the joint probability of the observed data and the latent variables—instead of the integrated observed-data likelihood alone. It derives a closed-form expression for this integrated complete-data likelihood under conjugate priors (Proposition 1), giving an exact ICL criterion. It then uses that closed form in an importance-sampling scheme to approximate the integrated observed-data likelihood, producing a second criterion, logIL.IS, that explicitly exploits the latent structure. Simulations across null, direct, indirect, and complete latent graphs show logIL.IS is consistently competitive across weak and strong signals, while ICL recovers the correct structure when latent variables are estimated accurately but underselects complex models when measurement error is high. The paper concludes that explicitly accounting for latent structure is a promising route for SEM model selection, with no single criterion dominating.","feed_headline":"Exact ICL criterion recovers latent dependency graphs","feed_subtitle":"Two latent-aware criteria—one exact, one sampled—stay competitive with AIC/BIC across weak and strong signals.","key_machinery":"The central object is Proposition 1's closed-form integrated complete-data likelihood. It factorizes the SEM into p independent univariate regressions (each observed variable on its single latent variable) and q independent latent regressions (each row of B), so the Gaussian-inverse-gamma priors make every parameter integral analytic. This yields the exact ICL (Integrated Completed Likelihood) criterion. The second object is the importance-sampling estimator of Eq. 8, which reuses the closed-form complete-data likelihood with proposal h(Z)=p(Z|X, theta-hat), a Gaussian whose mean and covariance are given by the conditional of the joint Gaussian model; it is the mechanism that converts the la","core_discovery":"Within the Gaussian SEM class defined by single-loading measurement (each row of the loading matrix has exactly one nonzero entry) and an acyclic latent graph (B strictly lower triangular, Gamma = I_q), the paper establishes that the integrated complete-data likelihood p(X,Z | M) can be computed in closed form under Gaussian and inverse-gamma priors. Proposition 1 gives this formula; replacing the latent variables by their posterior expectation (the MAP estimate) yields the proposed ICL criterion (Eq. 7). Because the complete-data likelihood is tractable, the paper also approximates the integrated observed-data likelihood by importance sampling from the Gaussian conditional distribution p(Z","pith_inferences":["The paper's oracle experiments suggest a practical fix: instead of plugging one MAP estimate of Z into ICL, one could average the closed-form complete-data likelihood over a posterior sample of Z; this is a natural follow-up that may recover the oracle's performance under high measurement error.","The closed-form derivation depends on the single-loading measurement model and the acyclic structural model, so the criteria are not immediately usable for cross-loadings or feedback loops; extending Proposition 1 to those cases would be a direct test of the approach's generality.","LogIL.IS incurs substantial computational cost with R=n samples; a variance-reduced importance sampler or a deterministic quadrature version could make latent-aware selection practical for large model spaces while retaining its weak-signal advantage.","Since ICL and logIL.IS respond differently to weak direct effects, an adaptive strategy that uses logIL.IS for screening and ICL for confirmation might be better than either criterion alone; this is an editorial suggestion, not a paper claim."],"forward_implications":["For Gaussian SEMs with single-loading and acyclic latent graphs, the ICL criterion can be evaluated in closed form, avoiding Laplace approximations.","logIL.IS provides a latent-aware approximation of the integrated observed-data likelihood that is competitive across weak and strong signals and improves on BIC when direct effects are weak.","ICL is reliable only when the latent variables are estimated accurately; at measurement error sigma^2=0.3 it systematically selects too-simple models even at n=1000.","The oracle comparisons imply the ICL formula itself is sound: the bottleneck is estimation of continuous latent variables, not the complete-data formulation.","No single criterion dominates; the simulations suggest the best choice depends on sample size, signal strength, and measurement quality."],"fun_headline_variants":["Exact ICL recovers latent graphs in SEM selection","Importance sampling with latent info rivals AIC/BIC","New criteria harness latent structure for SEM choice","Two latent-aware criteria for robust SEM model selection"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole derivation assumes that every observed variable loads on exactly one latent variable and that the latent dependency graph is acyclic (B strictly lower triangular with Gamma=I_q); without this, the closed-form ICL and the importance-sampling approximation are not defined as written.","fun_headline_variants_meta":{"raw":{"variants":["Exact ICL recovers latent graphs in SEM selection","Importance sampling with latent info rivals AIC/BIC","New criteria harness latent structure for SEM choice","Two latent-aware criteria for robust SEM model selection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000739,"raw_usage":{"total_tokens":3112,"prompt_tokens":696,"completion_tokens":2416,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":440,"completion_tokens_details":{"reasoning_tokens":2354}},"tokens_in":440,"tokens_out":2416,"duration_ms":19525,"temperature":1.0,"reasoning_tokens":2354,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T08:34:59.016531+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a Gaussian SEM whose true model has a cross-loading (one observed variable with two nonzero loadings) or a feedback loop in the latent graph, fit the candidate models as in Section 4, and compute the proposed ICL and logIL.IS: if selection rates collapse or the criteria cannot be evaluated, that shows the method does not extend beyond the single-loading acyclic class.","supporting_citations":[],"review_version":1}