{"id":"98fa7b1f-a836-4379-8eda-5a248932966b","arxiv_id":"2608.10328","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A single set-theoretic condition characterizes identifiability of block structured latent variable models, and under it the constrained maximum likelihood estimator attains oracle rates and oracle Cramer-Rao efficiency, with a linear-convergence algorithm.","lead":"This paper builds a general statistical theory for latent variable models in which observed items are grouped into blocks, each tied to a subset of hidden factors. It states exactly when such models are identifiable, shows the maximum likelihood estimator is consistent and efficient, and supplies a provably convergent algorithm.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3's exact equality between the Lagrangian minimizer and the constrained MLE is the least secure link: the penalty is scaled by 1/(NJ), so the proof must show a zero-penalty representative exists for the finite-sample global minimizer, and the core constructive step is only sketched.","rationale":"The reader's weakest_assumption focused on Assumption 1(iv) and Assumption 3 (J* proportional to J), which are genuine coverage restrictions. My concern is adjacent but slightly different: the most load-bearing step for the paper's central claims is the exact global-optimality equivalence in Theorem 3, whose proof is only sketched and whose key mechanisms live in the missing supplement. Given the coherent internal structure, the plausible M-Q Condition, and supportive simulations, I do not see a demonstrated error; but the global-optimality step is where a hidden assumption would most damage the paper. This supports the reader's CONDITIONAL verdict rather than a full ACCEPT, and it does not require moving to REJECT unless the supplement or the numerical check reveals an actual counterexample. The paper would be materially strengthened by releasing the supplement, code, and the numerical verification described above.","tokens_in":25191,"tokens_out":31860,"duration_ms":348345,"concrete_test":"Obtain the supplement and independently verify Lemma 3 and the constructive contradiction for P(~F) != 0; in parallel, run a certified global optimization check on a minimal M-Q design (e.g., Figure 1(d), D=5, logistic model) for N=J=50,100,200, solving both (6) and (8). If any reported global minimizer of (8) has P(~F) > 0, or if the optima of (6) and (8) differ by more than numerical tolerance, Theorem 3 fails as stated; if P(~F)=0 and the optima coincide in all runs, the concern is resolved for those configurations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central inference results (Theorems 4-5) all flow through Theorem 3, which asserts that the global minimizer of L_nu over the box (8) exactly equals the constrained MLE (6). This is a strong conclusion: L_nu = L + nu P, with P(F) = (1/(2NJ)) times squared constraint violations, so P is O(1/(NJ)) while the negative log-likelihood L has scale NJ. A finite-sample overfitting gain of O(1) in L therefore dwarfs any O(1/(NJ)) penalty, and the claimed equality can hold only if every near-optimal point can be replaced, with no increase in L, by a point satisfying the constraints exactly. The main text does not prove this; it states that a reference parameter set in Xi* has nearly minimal L, that the global minimizer therefore has small P, and that a constructive argument converts P != 0 into a strictly smaller L_nu. The key lemmas (Lemma 3, the Hv construction, and the P(~F) != 0 contradiction) are all deferred to the unavailable supplement. In particular, 'small P' is not 'zero P', and the proof that the finite-sample global minimizer lies in the identifiable constrained family is exactly the step that needs the full supplement. Without that step, Theorem 3 is an unverified lynchpin rather than an established equivalence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies a general class of block structured latent variable models, where the observed variables are partitioned into blocks and each block loads on a specified subset of latent factors, with partial orthogonality constraints encoded by a matrix M. The main contributions are: (i) a combinatorial condition, called the M-Q Condition, claimed to be necessary and sufficient for structural identifiability of the model parameters (Theorem 1); (ii) a Lagrangian-type formulation with a quadratic penalty P(F), whose global optimum is claimed to coincide exactly with the constrained maximum likelihood estimator (Theorem 3); (iii) non-asymptotic l2 and l-infinity error bounds and asymptotic normality results with oracle Cramer-Rao variances (Theorems 4 and 5); and (iv) a blockwise first-order algorithm with linear convergence and statistical equivalence to the exact estimator (Proposition 1 and Corollary 1). The theory is illustrated by simulations over four block designs and by an application to PISA 2022 data with a bifactor model. All technical proofs are deferred to a separate Supplementary Material file that is not included in the arXiv submission.","tokens_in":25589,"tokens_out":7237,"duration_ms":68039,"significance":"If the results are correct, the paper would provide the first general, checkable characterization of identifiability for block structured latent variable models, and would unify many existing special-case results. The M-Q Condition is a genuinely new combinatorial device, and the claimed rates (N^{-1}+J^{-2} for loadings, N^{-2}+J^{-1} for factors) together with oracle-efficiency asymptotic distributions would be substantial advances over the existing O_p(N^{-1}+J^{-1}) bounds. The simulation coverage rates and the empirical analysis are consistent with the stated theorems and add credibility. However, the central equivalence in Theorem 3 is only sketched, and the entire proof machinery is in an unavailable supplement, so the paper as submitted cannot be fully verified.","major_comments":[{"comment":"The claimed exact equality between the Lagrangian minimizer and the constrained MLE is the lynchpin of the paper, but it is not established in the main text. Since P(F) is scaled by 1/(NJ) while the negative log-likelihood L has scale NJ, a finite-sample improvement of O(1) in L can dominate any penalty contribution; the proof must show that a zero-penalty representative exists with no increase in L. The text only states that a reference parameter set in Xi* has nearly minimal L, that the global minimizer therefore has small P, and that a constructive argument converts P != 0 into a strictly smaller L_nu. These are the exact steps that require Lemma 3, the Hv construction, and the P(~F) != 0 contradiction, all of which are deferred to the unavailable supplement. Because Theorems 4, 5, Proposition 1, and Corollary 1 all flow through Theorem 3, this is a load-bearing gap.","section":"Section 3.2, Theorem 3 and Eq. (14)"},{"comment":"The theory requires J* to be proportional to J (Assumption 3), and the strong convexity lower bound in Eq. (12) degrades with J*/J. But J* is defined as max over S of min_{k in S} |J_k|, which can be O(1) while the M-Q Condition still holds, because a single small block can supply the identifying restrictions. Thus the paper excludes designs whose identifying block is vanishingly small, and its claims in the introduction and Section 4 about covering a 'wide range of asymptotic regimes' and 'various block designs' are considerably stronger than what is actually proved. The statement should be explicitly qualified to the J*/J asymptotically non-vanishing regime.","section":"Assumption 3 and Theorem 2, Eq. (12)"},{"comment":"The arXiv version submitted for review does not include the Supplementary Material, yet the proofs of all main results are in it: Lemma 1 and Lemma 3 are used in Section 3, the proof details of Theorem 3 are in Section E, Proposition 1 states 'Proof see Section E.5', and the initialization of the algorithm is given as Algorithm S2 in the supplement. Consequently, none of the theorem statements can be independently checked from the submitted document. This is not itself a mathematical error, but it makes the manuscript incomplete as a standalone submission; the supplement must be provided to the referees.","section":"Throughout; Supplement references"}],"minor_comments":[{"comment":"The second closure rule of the M-extended pi-system is terse: it should state explicitly that n is a positive integer and that the sets S_0 \\ S_r are relative complements; the current wording leaves the reader to guess the intended meaning for n=1 and for empty intersections.","section":"Definition 2"},{"comment":"The penalty P(F) is written with notation such as diag(M_ff) and M_ff - M_ff∘M; the dimensions and the use of the Hadamard product are clear only in context, so a brief sentence defining each term would improve readability.","section":"Section 3, P(F) definition"},{"comment":"The truncation step enforces only the box constraint; the orthogonality constraints and the normalization (2) are not enforced by projection during the iterations. The text should state explicitly how the algorithm output is related to the constrained MLE given Theorem 3 and Proposition 1, rather than leaving the reader to infer this.","section":"Algorithm 1, Step 11"},{"comment":"The claim that the Lagrangian-type formulation 'coincides exactly' with the original problem omits the sign indeterminacy caveat that is correctly handled later in the paper; adding a short qualification would avoid overstatement.","section":"Abstract and Section 3 summary"},{"comment":"The PISA analysis assumes that booklet assignment is conditionally ignorable, but the model and likelihood used for the observed-response data are described in only one sentence; a fuller statement of the assumed missing-data mechanism and its justification would help practitioners assess the analysis.","section":"Section 7"}],"recommendation":"major_revision","confidential_remarks":"The submitted arXiv file does not contain the Supplementary Material despite numerous references to it, including the proofs of Lemma 1, Lemma 3, Proposition A2, and Proposition 1. Please ensure the supplement is shared with referees before any further evaluation. The self-citations Cui & Xu (2026a, 2026b) are used as foundations, and the novelty over those works is stated only in general terms; if the supplement resolves the constructive step behind Theorem 3, the paper would likely be a substantial contribution to the statistical analysis of structured latent variable models."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The identifiability theorem is the real contribution: the M-extended pi-system is a clean combinatorial device, and the if-and-only-if characterization genuinely subsumes the special cases in Bing et al., Fang et al., and Qiao et al. The rest of the paper—oracle rates, Cramer-Rao efficiency, the algorithm guarantee—flows through Theorem 3, and that theorem is not actually proved in the manuscript. The supplement is promised but absent, and the main-text sketch is exactly where the subtle step lives. The stress-test note lands: with the penalty scaled by 1/(NJ), small P does not imply zero P, and the claim that any infeasible global minimizer can be improved is precisely what needs the deferred construction.\n\nWhat the paper does well, in addition to Theorem 1: Theorem 2's identification of J* as the critical block size, with a matching upper bound, is a solid contribution. The examples are genuinely illustrative, especially Example 5, where relaxing one orthogonality constraint breaks identifiability. The rates in Theorem 4 are stated cleanly, and the simulation coverage tracks the asymptotics. The statements look internally consistent; the gap is in what we can verify from the arXiv source, not in what the statements assert.\n\nSoft spots, in order: (1) Theorem 3 is an unverified lynchpin; no referee should sign off on Theorems 4 and 5 until the supplement is available and the constructive step is checked. (2) Assumption 3 requires J* proportional to J, so designs with tiny critical blocks are outside the theory; the authors acknowledge the curvature degenerates, which is a real limitation but not an error. (3) The truncation bound M in the definition of the feasible region is load-bearing for the global optimality step; it needs proof that the bound does not bite. (4) There is no code or data, and the PISA analysis is mostly illustrative. The self-citations are not a problem; they cite companion work for machinery, not for the central results.\n\nThis paper is for people working on high-dimensional factor models, item response theory, and multi-source data integration. It deserves a serious referee and a conditional accept shape: send it out, require the supplement to be part of the review package, and ask the authors to state what happens as J*/J goes to zero. I would bring it to the reading group.","headline":"A substantial theory paper with a real identifiability contribution, but the inference results rest on a global-optimality theorem that is deferred to an unavailable supplement.","tokens_in":26184,"tokens_out":2646,"would_cite":true,"duration_ms":20664,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H25","62F12","62E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that block structured latent variable models are structurally identifiable exactly when the M-Q Condition holds, and that the constrained maximum likelihood estimator then achieves oracle rates and efficient inference.","keywords":["block structured latent variable model","identifiability","maximum likelihood estimation","non-convex optimization","Lagrangian formulation","factor analysis","oracle rates","asymptotic normality"],"falsifier":"Construct a design where two latent factors appear together in every block (both in or both out of each $A_k$) and no orthogonality constraint separates them; the M-Q Condition fails, so Theorem 1 predicts non-identifiability, and the likelihood should admit a continuum of equivalent parameterizations. The same design, simulated with $N$ and $J$ growing, should show the smallest eigenvalue of the scaled Hessian at the truth converging to zero, contradicting the strong-convexity conclusion.","tokens_in":26718,"feed_emoji":"🧩","tokens_out":7784,"duration_ms":69078,"temperature":0.7,"pith_summary":"The paper studies latent variable models in which observed variables are partitioned into blocks, each block loading on a designated subset of latent factors, with optional orthogonality constraints between factors. It aims to replace case-by-case analyses of special designs with a single statistical theory: necessary and sufficient conditions for identifiability, consistency of the constrained maximum likelihood estimator, and efficient inference. The central claim is that identifiability holds exactly when the M-Q Condition is satisfied, meaning every single latent factor can be isolated by set intersections and differences of the block factor sets, aided by the orthogonality constraints. To make estimation tractable, the paper introduces a Lagrangian-type objective whose global minimizer coincides with the constrained maximum likelihood estimator, and a block-wise gradient algorithm that converges linearly to it. If the theory is correct, the practical question \"which block design identifies my model?\" becomes a checkable combinatorial test, and the resulting estimators inherit oracle error rates and Cramér-Rao optimal variances.","feed_headline":"Identifiability of block latent models: one checkable condition","feed_subtitle":"Theorem 1 ties identifiability to the M-Q Condition; estimation then achieves oracle rates and efficient inference.","key_machinery":"The M-extended $\\pi$-system $\\pi_M(A_Q)$ is the smallest collection of factor-index subsets that contains each block's factor set $A_k$ and is closed under intersections and under a difference rule triggered by orthogonality entries of $M$; the M-Q Condition requires every singleton $\\{r\\}$ to belong to it. This closure is the combinatorial certificate of identifiability: it records exactly which contrasts among latent factors the block design and orthogonality constraints can distinguish. The second workhorse is the Lagrangian-type objective $L_\\nu(F, \\Lambda_Q, \\beta) = L(F, \\Lambda_Q, \\beta) + \\nu P(F)$, where $P(F)$ is a quadratic penalty vanishing exactly on the normalization and orthogonality constraints. Its scaled Hessian's smallest eigenvalue is shown to be governed by $J^*/J$, the relative size of the smallest block that cannot be removed while preserving the M-Q Condition, and this curvature bound drives the global-optimality, error-rate, and algorithm-convergence results.","core_discovery":"Theorem 1 states that the parameters $(F, \\Lambda_Q, \\beta)$ of a block structured latent variable model are structurally identifiable if and only if every singleton $\\{r\\}$ lies in the M-extended $\\pi$-system generated by the block sets $A_Q$. Theorem 3 shows that, under the same condition together with regularity assumptions, the constrained maximum likelihood estimator equals the global minimizer of the Lagrangian-type objective $L_\\nu$, and also equals the unique minimizer of $L_\\nu$ in a small neighborhood of the true parameters. Theorems 4 and 5 give non-asymptotic $\\ell_2$ and $\\ell_\\infty$ error bounds and asymptotic normality: loadings and intercepts estimate at rate $N^{-1} + J^{-2}$, latent factors at $N^{-2} + J^{-1}$, and the asymptotic variances reach the oracle Cramér-Rao lower bounds. The paper's claim, taken as a whole, is that identifiability, consistency, and efficiency for this model class are fully characterized by the combinatorial M-Q Condition and the size $J^*$ of the smallest non-removable block.","pith_inferences":["The closure criterion could be used prospectively as a design tool for multi-source data integration: arrange source groupings so each latent dimension is separable in $\\pi_M(A_Q)$, avoiding identifiability failures before data collection.","The theory predicts a phase transition as $J^*/J \\to 0$: curvature and error bounds degrade, so designs whose identifying block is vanishingly small should exhibit markedly worse estimation, which is testable in simulation.","The Lagrangian equivalence is a general template: any constrained non-convex estimator whose constraint set can be encoded as the zero set of a quadratic penalty may inherit the same global-optimum coincidence, provided a curvature certificate analogous to $J^*/J$ exists.","The paper treats the link function and the number of latent factors as known; extending the closure condition to unknown link or unknown dimension is a natural next step that the current framework does not yet cover."],"forward_implications":["A practitioner can decide whether a proposed block design and orthogonality constraint identify the model by computing the closure $\\pi_M(A_Q)$ and checking whether every singleton appears, instead of verifying analytic conditions design by design.","Given a fixed block structure, the minimal orthogonality constraint needed for identifiability can be found by relaxing entries of $M$ until the M-Q Condition fails.","The constrained maximum likelihood estimator can be computed by a first-order block-wise gradient algorithm with linear convergence, and after $T \\gg \\log(N \\vee J)$ iterations the algorithm's output inherits the exact estimator's asymptotic distribution.","Across scaling regimes such as $N = o(J^2)$ and $J = o(N^2)$, loading and factor estimators reach oracle rates and asymptotically efficient variances, so uncertainty quantification for factor scores and loadings is available in broad settings."],"supporting_citations":[{"why":"Supplies the baseline structured latent factor identifiability analysis under restrictive assumptions that the block theory generalizes.","marker":"Chen et al. 2020"},{"why":"Gives identifiability conditions for bifactor models, a special case of the block designs covered by the M-Q Condition.","marker":"Fang et al. 2021"},{"why":"Establishes identifiability for hierarchical latent factor models, another special case folded into the general block framework.","marker":"Qiao et al. 2025"},{"why":"Provides the high-dimensional factor model asymptotic inference theory that the present results refine for block constraints.","marker":"Bai & Li 2012"},{"why":"Treats maximum likelihood estimation in nonlinear factor models and is a previous benchmark for the constrained non-convex problem.","marker":"Wang 2022"},{"why":"Develops inference for generalized latent factor models under minimal sparsity, a setting distinct from block-structured orthogonality constraints.","marker":"Cui & Xu 2026b"},{"why":"Uses quadratic-penalty Lagrangian formulations for nonlinear panel factor models, which the paper adapts to blockwise orthogonality constraints.","marker":"Chen, Fernández-Val & Weidner 2021"}],"fun_headline_variants":["M-Q condition: identifiability and oracle efficiency","Block latent models: one condition, full theory","Identifiability, consistency, efficiency: same condition","M-Q condition ties identifiability to optimal estimation","Sharp rates for block latent models via M-Q condition"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The theory assumes every block is informative enough to identify all latent factors it loads on (full-rank loadings with eigenvalues bounded away from zero) and that the smallest non-removable block grows proportionally with the total number of items, so $J^*/J$ stays bounded away from zero.","fun_headline_variants_meta":{"raw":{"variants":["M-Q condition: identifiability and oracle efficiency","Block latent models: one condition, full theory","Identifiability, consistency, efficiency: same condition","M-Q condition ties identifiability to optimal estimation","Sharp rates for block latent models via M-Q condition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000521,"raw_usage":{"total_tokens":2345,"prompt_tokens":972,"completion_tokens":1373,"prompt_tokens_details":{"cached_tokens":0},"prompt_cache_hit_tokens":0,"prompt_cache_miss_tokens":972,"completion_tokens_details":{"reasoning_tokens":1312}},"tokens_in":972,"tokens_out":1373,"duration_ms":11621,"temperature":1.0,"reasoning_tokens":1312,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T03:09:22.336650+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a design where two latent factors appear together in every block (both in or both out of each $A_k$) and no orthogonality constraint separates them; the M-Q Condition fails, so Theorem 1 predicts non-identifiability, and the likelihood should admit a continuum of equivalent parameterizations. The same design, simulated with $N$ and $J$ growing, should show the smallest eigenvalue of the scaled Hessian at the truth converging to zero, contradicting the strong-convexity conclusion.","supporting_citations":[],"review_version":1}