{"id":"a5d49550-d155-4c58-bbb2-849c8b0db3e1","arxiv_id":"2608.10731","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A general factor is distinguishable from correlated factors when within-cluster loadings are non-proportional; the paper proves a sufficient condition, maps the boundary numerically, and tests a two-step decision procedure.","lead":"This paper proves when a general factor in a bifactor model can be distinguished from correlated group factors, and offers a two-step procedure to make the decision honestly. It matters because many psychological studies claim a general factor from model fit alone, which this work shows can be misleading.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1 is sound under its stated diagonal-uniqueness assumption; the genuine soft spot is the unquantified gap between that assumption and the adversarial doublet population, where the operational comparison can be wrong with growing N.","rationale":"The reader correctly identifies the diagonal-uniqueness assumption as the weakest point. My independent reading confirms that the proof of Theorem 1 is internally coherent: Lemma A3 is correct, the rank-budget argument in Appendix A is valid, and no algebraic gap appears in the derivation. The soft spot is not the internal logic but the external validity of the population-level claim in the presence of residual dependence. Finding 9 demonstrates an operational hazard, but the paper does not report D_5 for the adversarial population, which is the quantity needed to decide whether the hazard reflects a genuine population-level distinction or a mismatch between the anchored comparison and the unrestricted K-factor class. The paper explicitly acknowledges that the step-2 comparison is not an estimator of D_K, so this gap is disclosed rather than hidden. Because the central theorem is correct under its stated conditions and the limitations are transparent, the reader's ACCEPT verdict should stand; the proposed D_5 computation would sharpen the scope statement without overturning the paper's main contribution.","tokens_in":27895,"tokens_out":17209,"duration_ms":178177,"concrete_test":"For the P5+doublet population (c = 0.20), compute D_5 as defined in Section 2.2.4 by population-level ML with at least 40 random starts and tolerance 1e-8, as in Table 11, and record the max-entry discrepancy of Sigma from the best rank-5 Lambda Lambda' + Psi. If D_5 is zero to machine precision, the doublet is exactly absorbable by the K-factor class, showing Finding 9 is an artifact of the anchored design rather than a violation of the theorem's domain. If D_5 > 0, the diagonal-uniqueness violation is population-level and the scope limitation is exactly as the paper describes. Repeating the computation at c = 0.24 would show where the doublet stops being absorbable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 1 is conditional on Sigma = C + Psi with Psi diagonal. Under that condition, the proof is sound: Lemma A3 forces any cluster with w_k != 0 and at least two nonzero group loadings to contribute at least one dimension of rank, and the exhaustion argument rules out an exact K-factor rival. The concern is not with the theorem but with what it licenses. The paper's own adversarial population (P5+doublet, Table 3) violates the diagonal-uniqueness condition by construction, and Finding 9 shows the step-2 anchored comparison increasingly prefers the bifactor as N grows. The paper does not report the unrestricted population distance D_5 for this population. If D_5 = 0, the doublet is absorbable by some rank-5 Lambda with diagonal uniquenesses; then the population is in the K-factor class, the general factor is not distinguishable at the population level, and Finding 9 is an artifact of the anchored comparison design rather than a consequence of violating the theorem's assumption. If D_5 > 0, the violation genuinely creates a population-level pull toward an added dimension, and the error is one of substantive interpretation. Either way, the operational decision is not protected by Theorem 1, and the paper's own concession that step 2 is 'evidence about whether an added general column improves on the delivered anchored structure, not an exact sample estimator or formal test of unrestricted D_K' leaves the central practical claim dependent on an unquantified gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a covariance-level theory of when an additional general (bifactor) dimension is distinguishable from a set of correlated first-order factors. Proposition 1 states that if general and group loadings are proportional within every cluster, the bifactor and correlated-factor classes are covariance-equivalent; Theorem 1 states that if every cluster is non-proportional, no K-factor model with diagonal uniquenesses can reproduce the bifactor covariance. The paper defines a population distance D_K to the K-factor class, describes a mixed boundary that is characterized numerically, and proposes a two-step partially exploratory factor analysis procedure: an oblique anchored sweep to identify a stable first-order structure, followed by an anchored variational-BIC comparison of oblique versus bifactor models. Simulation studies calibrate the procedure's thresholds, identify hazards including count instability and absorbed residual dependence, and four empirical datasets illustrate delivery, non-delivery, and conditional comparisons. The manuscript is explicit that step 2 is a conditional comparison at a selected structure, not a formal estimator or test of the unrestricted population distance.","tokens_in":28187,"tokens_out":8723,"duration_ms":100432,"significance":"The paper's central theoretical contribution is substantial. The proof of Theorem 1 in Appendix A is internally coherent, and the key Lemma A3 algebra checks: a diagonal difference between two rank-one within-cluster terms forces either proportionality or a single-indicator spike. This gives a sharp sufficient condition for distinguishability that does not depend on estimation details. The paper also ships a reproducible OSF archive, calibrates its delivery threshold on a prespecified half of replications and checks it on the untouched half, and treats explicit non-delivery as a legitimate output. If the distinction between identifiability and distinguishability becomes standard in the bifactor literature, this paper will have done useful conceptual work. The principal weaknesses are operational: the step-2 criterion is nonstandard and not fully specified, and the paper's headline hazard of 'absorbed local dependence imitating a general factor' is not covered by Theorem 1 and is not quantified by the population distance D_5 for the adversarial population.","major_comments":[{"comment":"The adversarial population P5+doublet is constructed to violate the diagonal-uniqueness assumption on which Theorem 1 rests: the residual doublet is nonzero off-diagonal covariance. The manuscript never reports D5(Σ) for this population. If D5 = 0, the population is actually in the K-factor class, and the growing bifactor preference in Table 5 is an artifact of the anchored comparison design rather than evidence about the population property; if D5 > 0, the bifactor preference reflects a genuine population-level distance and the error is only in the substantive interpretation. These two readings have different implications for the claimed hazard. Please compute D5 for P5+doublet with the same population-level ML procedure used in Table 11, and adjust the Finding 9 wording to state which reading holds.","section":"Section 3.2.1, Table 5, and Section 5 (Finding 9)"},{"comment":"The operational criterion of step 2, 'variational BIC evaluated at the variational estimate under hard selection with the effective parameter count defined as the Jacobian rank of the implied covariance', is not defined precisely enough to be audited. The Jacobian of which map, evaluated at which point, and how the hard-selected zeros and boundary configurations are counted should be stated formally, or the defining equation from Chen and Jin (2026a) should be reproduced in the text. Every step-2 margin in Tables 5 and 10, and the calibration guidance in Table 6, depends on this definition, so the criterion cannot remain a verbal gloss.","section":"Section 2.3 and Section 2.4 (operational variational BIC)"}],"minor_comments":[{"comment":"The sentence 'the choice Finding 7 revises once the two are compared end to end' is confusing because Finding 7 is introduced later in Study 3, not among the six preliminary findings listed in Section 3.1; please rephrase to say that the choice is revised by Finding 7.","section":"Section 3.1"},{"comment":"The text repeatedly renders 'sufficient' as 'suﬀicient'; this typographical error should be corrected in the final version.","section":"Throughout"},{"comment":"The Holzinger 24 step-2 BIC margin is only -6.8, which is conventionally negligible; the text says it 'follows the same direction with smaller margins,' but a sensitivity note stating that this margin is weak would help readers avoid over-interpreting a single dataset.","section":"Section 4.4, Table 10"},{"comment":"The paper notes that in simulations the L1/L2 boundary is drawn by backbone-specification agreement while in applications it is drawn by whether the delivered count is also the persisting count; this difference should be stated as a limitation in the Discussion rather than only as a parenthetical remark.","section":"Section 4.3 and Table 9"}],"recommendation":"major_revision","confidential_remarks":"The paper leans heavily on the author's own vbpm/PEFA machinery and cites it extensively; this is not by itself a problem because the code is public and the threshold calibration is internal, but the editor may wish to confirm that the OSF archive is accessible to reviewers before final acceptance. The two requested clarifications, the D5 value for the adversarial population and a formal definition of the operational BIC, are within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful news is Theorem 1. It gives a sharp sufficient condition for when an added general factor is distinguishable from correlated factors at the covariance level, and the proof in Appendix A checks out. Lemma A3 is the right workhorse, and the rank-counting argument is clean. Proposition 1 is the known Schmid-Leiman/higher-order equivalence, correctly attributed; the novelty is the converse direction and the graded distance D_K. The mixed-boundary analysis is genuinely new but only numerical, which the paper admits.\n\nThe simulation work is more disciplined than most in this area. Calibrating the delivery threshold on a prespecified half and checking once on the untouched half is good practice. Disclosing the Finding 4 pipeline bug openly is to the authors' credit. And the adversarial doublet population is the right way to stress-test the procedure's assumptions.\n\nThe soft spot is the diagonal-uniqueness assumption. Theorem 1 lives in a world where residuals are uncorrelated. Once you allow residual dependence—as the paper's own P5+doublet does—the theorem's separation guarantee no longer applies. The paper is honest about this, but the consequence is that the operational step-2 comparison is not protected by the theorem. The stress-test note makes a fair point: the paper never reports the unrestricted D_5 for the doublet population. If D_5 = 0, then the doublet is absorbable by some rank-5 model with diagonal uniquenesses, and Finding 9's growing preference for the bifactor is an artifact of the anchored comparison design, not a population-level fact. If D_5 > 0, it is a real pull toward an added dimension. Either way, the operational decision is narrower than the abstract's opening sentence implies. That is a real gap, but it is a gap in the scope of the operational claim, not in the theorem itself.\n\nThe mixed boundary is also left open—a formal characterization is explicitly deferred. That is a limitation, but a proportionate one.\n\nWho is this for? Psychometricians who care about the bifactor versus correlated-factors decision, and methodologists working on factor model selection. The paper deserves a serious referee. My recommendation: send it out, but ask the authors to tighten the abstract, report D_5 for the doublet population, and be explicit in the discussion that the operational procedure is a heuristic that can be misled by residual dependence, not a formal test of the population property.","headline":"Theorem 1 is a genuine, sound result with a clean proof; the operational procedure is honestly bounded, and the soft spot is the unquantified gap between the theorem's diagonal-uniqueness world and residual dependence.","tokens_in":28737,"tokens_out":2503,"would_cite":true,"duration_ms":25574,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H25","62P15"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that whether a general factor is distinguishable from correlated group factors is a covariance-level property, decides the reducible and non-proportional cases, and makes the middle case a graded distance to the…","keywords":["bifactor model","distinguishability","covariance equivalence","higher-order model","partially exploratory factor analysis","stable structure","residual dependence","proportionality"],"falsifier":"Run the paper's adversarial population: a higher-order five-cluster base with a residual covariance of 0.20 between two non-anchor items of cluster 1, then fit the two-step process at sample sizes 500, 1000, and 2000. If the step-1 stability indicators become non-clean, or if the bifactor is not preferred increasingly often as the sample size grows, Finding 9's claim that absorbed local dependence imitates a general factor with clean stability indicators would be refuted.","tokens_in":27643,"feed_emoji":"📊","tokens_out":11216,"duration_ms":104983,"temperature":0.7,"pith_summary":"The paper tries to establish exactly when an added general dimension, a factor that loads on every item, is distinguishable from correlated first-order factors, and it locates the answer in the population covariance matrix rather than in estimators or sample size. Its two anchors are Proposition 1, showing that proportional general and group loadings in every cluster make a bifactor covariance-equivalent to a K-factor correlated model, and Theorem 1, showing that if every cluster is non-proportional, no K-factor model with diagonal uniquenesses can reproduce the bifactor covariance, so the general dimension is required. Between these poles sits a mixed boundary that depends on the pattern of within-cluster loading ratios and is quantified by the population distance $D_K$ to the K-factor class. The paper then builds a two-step partially exploratory procedure that delivers a stable first-order structure only when it reproduces across adjacent counts and treats non-delivery as a legitimate outcome; simulation shows that a unanimous count can still fail to reproduce and that absorbed local dependence can imitate a general factor, with the error growing in sample size while stability indicators stay clean. A reader would care because this recasts the recurring bifactor-versus-correlated-factors debate as a question that can be posed at the covariance level and answered conditionally, rather than as a model-selection lottery.","feed_headline":"Misaligned loadings are what make a general factor detectable","feed_subtitle":"Proportional clusters make the two models identical; misaligned clusters make the general factor detectable","key_machinery":"The carrying object is the within-cluster decomposition $b_{g|k} = \\alpha_k u_k + w_k$, with $u_k$ the normalized group-loading column and $w_k$ its orthogonal residual; $w_k = 0$ defines a silent cluster. Theorem 1's proof turns on Lemma A3, an elementary rank-one fact: for $n \\geq 3$, if $\\tau bb' - ss'$ is diagonal, then $s$ has at most one nonzero entry or $s$ is proportional to $b$, which forces each non-proportional cluster to push one dimension of rank through its block and exhausts a rank-$K$ budget. The second piece of machinery is the population distance $D_K(\\Sigma)$ to the class of K-factor covariance matrices with nonnegative diagonal uniquenesses, which grades distinguishability and defines the hypothesis $H_0: D_K = 0$; Proposition 1 gives $D_K = 0$ on reducible structures, and Theorem 1 gives $D_K > 0$ under its conditions. A resistant cluster, defined as having at least four items and two disjoint unequal within-cluster ratio pairs, is the pattern that numerical analysis identifies as keeping a positive distance at the mixed boundary.","core_discovery":"On the paper's own terms, the central discovery is that 'is the bifactor distinguishable from correlated factors?' has a covariance-level answer. Decompose each cluster's general loadings into the part aligned with the cluster's group loadings and a residual $w_k$; clusters with $w_k = 0$ are silent and contribute nothing to the distinguishability question at any sample size. If all clusters are silent, Proposition 1 states that the bifactor covariance is exactly a K-factor covariance with the same uniquenesses, so no data can separate the two representations. If every cluster has $w_k \\neq 0$, Theorem 1 states that $\\operatorname{rank}(C + \\Delta) \\geq K+1$ for every diagonal $\\Delta$, so the general dimension cannot be absorbed into uniquenesses and no K-factor model fits the population covariance. The intermediate case is mixed and graded: distinguishability is measured by $D_K(\\Sigma)$, the minimum likelihood discrepancy from $\\Sigma$ to the best K-factor model, and whether a lone non-proportional cluster survives depends on its within-cluster ratio pattern, with resistant clusters retaining a strictly positive distance.","pith_inferences":["The within-cluster ratio pattern could be used prospectively as a design diagnostic: before fitting a bifactor, inspect estimated ratios $\\rho_i = b_{s,i}/b_{g,i}$ within each cluster, and flag clusters whose ratios are constant except for one entry as weak carriers of the general-factor decision.","$D_K$ could be translated into a practical equivalence bound: if the population distance falls below a sample-size-dependent threshold, the added general column is arguably negligible even if formally required, analogous to a 'close enough' criterion for model comparisons.","The doublet hazard implies a testable prediction for applied data: in batteries where a general factor is preferred only after residual dependence is absorbed, the general-factor loadings should fail to replicate in a new sample or show weak external validity, because they are absorbing covariance rather than a stable dimension.","The mixed-boundary analysis leaves a concrete research target: a complete characterization of which mixed configurations admit exact K-factor rivals would let practitioners know exactly which silent clusters are fatal to the general-factor decision."],"forward_implications":["When loadings are proportional in every cluster, collecting more data or using a better estimator cannot decide between bifactor and correlated factors: the covariance matrices are identical, so the choice is a parsimony or substantive decision, not an empirical one.","When every cluster is non-proportional, the population covariance cannot be reproduced by any K-factor model with diagonal uniquenesses, so a general factor is required in the population; statistical power to detect it scales with $N \\cdot D_K$ rather than with eigenvalue gaps.","Distinguishability is graded, so reporting only a binary 'general factor yes or no' loses information; the population distance $D_K$ and the per-cluster residuals $w_k$ describe how close the data are to the reducible boundary.","The two-step procedure makes stable structure a precondition for the bifactor comparison: if the first-order structure does not reproduce across adjacent counts, the appropriate output is non-delivery, not a comparison at an unsupported count.","Absorbed local dependence can manufacture a general factor even at the correct count, with the error growing in sample size, so residual diagnostics should accompany the comparison wherever residual dependence is plausible."],"supporting_citations":[{"why":"Supplies the Schmid–Leiman transformation and the constant within-cluster ratio that defines the reducible bifactor, the starting point of Proposition 1.","marker":"Schmid & Leiman (1957)"},{"why":"Formalizes the relationship between higher-order and hierarchical factor models and gives the second-order loading formula used in Proposition 1's equivalence.","marker":"Yung et al. (1999)"},{"why":"Settles bifactor parameter identifiability and provides sufficient conditions in terms of within-cluster linear independence, which the paper separates from distinguishability.","marker":"Fang et al. (2021)"},{"why":"States that the higher-order model imposes a proportionality constraint on the bifactor, explaining why the bifactor tends to fit better and motivating the population-level criterion.","marker":"Gignac (2016)"},{"why":"Prior analysis of when second-order and bifactor models are distinguishable, framing the question that Theorem 1 answers.","marker":"Mansolf & Reise (2017)"},{"why":"Supplies the partially exploratory factor analysis machinery, gain rule, and factor-number selection used in step 1 of the two-step procedure.","marker":"Chen & Jin (2026a)"},{"why":"Defines the design-matrix (PCFA) specification with anchor entries that the paper uses in all simulation and empirical fits.","marker":"Chen (2021)"},{"why":"Provides the constraint-based recovery guarantees that require pairwise non-proportionality, against which the paper positions its weaker cluster-level condition.","marker":"Qiao et al. (2025a)"}],"fun_headline_variants":["General factor detectable only when loadings misalign","Proportional loadings make bifactor and K-factor identical","No way to separate bifactor if loadings are proportional","General factor visibility hinges on non-proportional loadings","Silent clusters make the general factor undetectable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole separation argument assumes residuals are uncorrelated, so that a rival K-factor model can differ only by a diagonal uniqueness shift; if real data carry local residual dependence, that difference is no longer purely diagonal and the guarantee that a non-proportional cluster cannot be traded away no longer holds.","fun_headline_variants_meta":{"raw":{"variants":["General factor detectable only when loadings misalign","Proportional loadings make bifactor and K-factor identical","No way to separate bifactor if loadings are proportional","General factor visibility hinges on non-proportional loadings","Silent clusters make the general factor undetectable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000411,"raw_usage":{"total_tokens":2152,"prompt_tokens":994,"completion_tokens":1158,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":1081}},"tokens_in":610,"tokens_out":1158,"duration_ms":9918,"temperature":1.0,"reasoning_tokens":1081,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:32:25.036694+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's adversarial population: a higher-order five-cluster base with a residual covariance of 0.20 between two non-anchor items of cluster 1, then fit the two-step process at sample sizes 500, 1000, and 2000. If the step-1 stability indicators become non-clean, or if the bifactor is not preferred increasingly often as the sample size grows, Finding 9's claim that absorbed local dependence imitates a general factor with clean stability indicators would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Schmid–Leiman transformation and the constant within-cluster ratio that defines the reducible bifactor, the starting point of Proposition 1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Formalizes the relationship between higher-order and hierarchical factor models and gives the second-order loading formula used in Proposition 1's equivalence."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Settles bifactor parameter identifiability and provides sufficient conditions in terms of within-cluster linear independence, which the paper separates from distinguishability."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"States that the higher-order model imposes a proportionality constraint on the bifactor, explaining why the bifactor tends to fit better and motivating the population-level criterion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior analysis of when second-order and bifactor models are distinguishable, framing the question that Theorem 1 answers."}],"review_version":1}