REVIEW 2 major objections 4 minor 40 references
When Is a General Factor Distinguishable? Non-Proportionality, Stable Structure, and the Bifactor Decision
T0 review · 2 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper establishes that whether a general factor is distinguishable from correlated group factors is a covariance-level property, decides the reducible and non-proportional cases, and makes the middle case a graded distance to the…
desk verdict Theorem 1 is a genuine, sound result with a clean proof; the operational procedure is honestly bounded, and the soft spot is the unquantified gap between the theorem's diagonal-uniqueness world and residual dependence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the within-cluster decomposition $b_{g|k} = \alpha_k u_k + w_k$, with $u_k$ the normalized group-loading column and $w_k$ its orthogonal residual; $w_k = 0$ defines a silent cluster. Theorem 1's proof turns on Lemma A3, an elementary rank-one fact: for $n \geq 3$, if $\tau bb' - ss'$ is diagonal, then $s$ has at most one nonzero entry or $s$ is proportional to $b$, which forces each non-proportional cluster to push one dimension of rank through its block and exhausts a rank-$K$ budget. The second piece of machinery is the population distance $D_K(\Sigma)$ to the class of K-factor covariance matrices with nonnegative diagonal uniquenesses, which grades distinguishability and defines the hypothesis $H_0: D_K = 0$; Proposition 1 gives $D_K = 0$ on reducible structures, and Theorem 1 gives $D_K > 0$ under its conditions. A resistant cluster, defined as having at least four items and two disjoint unequal within-cluster ratio pairs, is the pattern that numerical analysis identifies as keeping a positive distance at the mixed boundary.
What would settle it
Run the paper's adversarial population: a higher-order five-cluster base with a residual covariance of 0.20 between two non-anchor items of cluster 1, then fit the two-step process at sample sizes 500, 1000, and 2000. If the step-1 stability indicators become non-clean, or if the bifactor is not preferred increasingly often as the sample size grows, Finding 9's claim that absorbed local dependence imitates a general factor with clean stability indicators would be refuted.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that 'is the bifactor distinguishable from correlated factors?' has a covariance-level answer. Decompose each cluster's general loadings into the part aligned with the cluster's group loadings and a residual $w_k$; clusters with $w_k = 0$ are silent and contribute nothing to the distinguishability question at any sample size. If all clusters are silent, Proposition 1 states that the bifactor covariance is exactly a K-factor covariance with the same uniquenesses, so no data can separate the two representations. If every cluster has $w_k \neq 0$, Theorem 1 states that $\operatorname{rank}(C + \Delta) \geq K+1$ for every diagonal $\Delta$, so the general dimension cannot be absorbed into uniquenesses and no K-factor model fits the population covariance. The intermediate case is mixed and graded: distinguishability is measured by $D_K(\Sigma)$, the minimum likelihood discrepancy from $\Sigma$ to the best K-factor model, and whether a lone non-proportional cluster survives depends on its within-cluster ratio pattern, with resistant clusters retaining a strictly positive distance.
Load-bearing premise
The whole separation argument assumes residuals are uncorrelated, so that a rival K-factor model can differ only by a diagonal uniqueness shift; if real data carry local residual dependence, that difference is no longer purely diagonal and the guarantee that a non-proportional cluster cannot be traded away no longer holds.
Editorial extensions
If this is right
- When loadings are proportional in every cluster, collecting more data or using a better estimator cannot decide between bifactor and correlated factors: the covariance matrices are identical, so the choice is a parsimony or substantive decision, not an empirical one.
- When every cluster is non-proportional, the population covariance cannot be reproduced by any K-factor model with diagonal uniquenesses, so a general factor is required in the population; statistical power to detect it scales with $N \cdot D_K$ rather than with eigenvalue gaps.
- Distinguishability is graded, so reporting only a binary 'general factor yes or no' loses information; the population distance $D_K$ and the per-cluster residuals $w_k$ describe how close the data are to the reducible boundary.
- The two-step procedure makes stable structure a precondition for the bifactor comparison: if the first-order structure does not reproduce across adjacent counts, the appropriate output is non-delivery, not a comparison at an unsupported count.
- Absorbed local dependence can manufacture a general factor even at the correct count, with the error growing in sample size, so residual diagnostics should accompany the comparison wherever residual dependence is plausible.
Reading between the lines
- The within-cluster ratio pattern could be used prospectively as a design diagnostic: before fitting a bifactor, inspect estimated ratios $\rho_i = b_{s,i}/b_{g,i}$ within each cluster, and flag clusters whose ratios are constant except for one entry as weak carriers of the general-factor decision.
- $D_K$ could be translated into a practical equivalence bound: if the population distance falls below a sample-size-dependent threshold, the added general column is arguably negligible even if formally required, analogous to a 'close enough' criterion for model comparisons.
- The doublet hazard implies a testable prediction for applied data: in batteries where a general factor is preferred only after residual dependence is absorbed, the general-factor loadings should fail to replicate in a new sample or show weak external validity, because they are absorbing covariance rather than a stable dimension.
- The mixed-boundary analysis leaves a concrete research target: a complete characterization of which mixed configurations admit exact K-factor rivals would let practitioners know exactly which silent clusters are fatal to the general-factor decision.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a covariance-level theory of when an additional general (bifactor) dimension is distinguishable from a set of correlated first-order factors. Proposition 1 states that if general and group loadings are proportional within every cluster, the bifactor and correlated-factor classes are covariance-equivalent; Theorem 1 states that if every cluster is non-proportional, no K-factor model with diagonal uniquenesses can reproduce the bifactor covariance. The paper defines a population distance D_K to the K-factor class, describes a mixed boundary that is characterized numerically, and proposes a two-step partially exploratory factor analysis procedure: an oblique anchored sweep to identify a stable first-order structure, followed by an anchored variational-BIC comparison of oblique versus bifactor models. Simulation studies calibrate the procedure's thresholds, identify hazards including count instability and absorbed residual dependence, and four empirical datasets illustrate delivery, non-delivery, and conditional comparisons. The manuscript is explicit that step 2 is a conditional comparison at a selected structure, not a formal estimator or test of the unrestricted population distance.
Significance. The paper's central theoretical contribution is substantial. The proof of Theorem 1 in Appendix A is internally coherent, and the key Lemma A3 algebra checks: a diagonal difference between two rank-one within-cluster terms forces either proportionality or a single-indicator spike. This gives a sharp sufficient condition for distinguishability that does not depend on estimation details. The paper also ships a reproducible OSF archive, calibrates its delivery threshold on a prespecified half of replications and checks it on the untouched half, and treats explicit non-delivery as a legitimate output. If the distinction between identifiability and distinguishability becomes standard in the bifactor literature, this paper will have done useful conceptual work. The principal weaknesses are operational: the step-2 criterion is nonstandard and not fully specified, and the paper's headline hazard of 'absorbed local dependence imitating a general factor' is not covered by Theorem 1 and is not quantified by the population distance D_5 for the adversarial population.
major comments (2)
- [Section 3.2.1, Table 5, and Section 5 (Finding 9)] The adversarial population P5+doublet is constructed to violate the diagonal-uniqueness assumption on which Theorem 1 rests: the residual doublet is nonzero off-diagonal covariance. The manuscript never reports D5(Σ) for this population. If D5 = 0, the population is actually in the K-factor class, and the growing bifactor preference in Table 5 is an artifact of the anchored comparison design rather than evidence about the population property; if D5 > 0, the bifactor preference reflects a genuine population-level distance and the error is only in the substantive interpretation. These two readings have different implications for the claimed hazard. Please compute D5 for P5+doublet with the same population-level ML procedure used in Table 11, and adjust the Finding 9 wording to state which reading holds.
- [Section 2.3 and Section 2.4 (operational variational BIC)] The operational criterion of step 2, 'variational BIC evaluated at the variational estimate under hard selection with the effective parameter count defined as the Jacobian rank of the implied covariance', is not defined precisely enough to be audited. The Jacobian of which map, evaluated at which point, and how the hard-selected zeros and boundary configurations are counted should be stated formally, or the defining equation from Chen and Jin (2026a) should be reproduced in the text. Every step-2 margin in Tables 5 and 10, and the calibration guidance in Table 6, depends on this definition, so the criterion cannot remain a verbal gloss.
minor comments (4)
- [Section 3.1] The sentence 'the choice Finding 7 revises once the two are compared end to end' is confusing because Finding 7 is introduced later in Study 3, not among the six preliminary findings listed in Section 3.1; please rephrase to say that the choice is revised by Finding 7.
- [Throughout] The text repeatedly renders 'sufficient' as 'sufficient'; this typographical error should be corrected in the final version.
- [Section 4.4, Table 10] The Holzinger 24 step-2 BIC margin is only -6.8, which is conventionally negligible; the text says it 'follows the same direction with smaller margins,' but a sensitivity note stating that this margin is weak would help readers avoid over-interpreting a single dataset.
- [Section 4.3 and Table 9] The paper notes that in simulations the L1/L2 boundary is drawn by backbone-specification agreement while in applications it is drawn by whether the delivered count is also the persisting count; this difference should be stated as a limitation in the Discussion rather than only as a parenthetical remark.
Circularity Check
No circularity: Theorem 1 is a self-contained mathematical proof, Proposition 1 is an attributed classical equivalence, and the operational thresholds are explicitly calibrated on held-out simulation halves rather than passed off as predictions.
full rationale
The central derivation, Theorem 1, is self-contained. Its proof (Appendix A) uses stated assumptions—orthogonal bifactor structure, nonzero general loadings, at least three items per cluster, at least two nonzero group loadings per cluster, and w_k != 0 in every cluster—and proceeds through Lemmas A1-A4 to show rank(C + Delta) >= K + 1 for every diagonal Delta. No fitted parameter, estimated threshold, or self-citation enters the theorem or its proof. The diagonal-uniqueness assumption is an explicit modeling assumption, not an input that is later relabeled as an output. Proposition 1 is likewise attributed to the Schmid-Leiman lineage (Yung et al., 1999) and presented as an equivalence result with a short algebraic demonstration, not as a novel prediction derived from the paper's own machinery. The simulation thresholds—20% cut, phi* = .85, RMSD .20/.10—are calibrated on a prespecified half of Study 3 replications and checked once on the untouched half, and the paper repeatedly identifies them as provisional guidance ('provisional to Study 3's populations'), not as theoretical consequences. Step 2's operational comparison is explicitly disclaimed as an estimator or formal test of the unrestricted distance D_K: the paper states it is 'evidence about whether an added general column improves on the delivered anchored structure, not an exact sample estimator or formal test of unrestricted D_K.' This prevents the empirical comparisons from reducing by construction to the population distance. Self-citations to the author's PEFA/vbpm framework supply software and selection-rule antecedents, but the load-bearing theoretical claims do not depend on those citations; the paper even openly reports a specification artifact (Finding 4) and checks specifications against each other (Finding 7). The adversarial doublet population is presented as a violation of Theorem 1's diagonal-uniqueness assumption and as a limitation, not as a consequence of the theorem, so it raises validity concerns rather than circularity. Overall, no step in the derivation chain equates a fitted input with a predicted output or imports a load-bearing conclusion solely from a self-citation.
Assumptions & free parameters
free parameters (6)
- ELBO gain selection cut =
20% primary, 10% sensitivity
- Minimum matched congruence threshold φ* =
0.85
- RMSD replacement thresholds =
aggregate ≤ 0.10, worst-column ≤ 0.20
- Posterior inclusion hard-selection cutoff =
0.5
- Spike variance v0 =
0.001
- Backbone anchor depth =
2 anchors per factor
assumptions (5)
- domain assumption Uncorrelated residuals (diagonal uniquenesses) in the factor model
- domain assumption At least 3 items per cluster, at least 2 nonzero group loadings per cluster, nonzero general loadings on every item, K ≥ 2
- domain assumption Continuous responses with covariance structure sufficient for the analysis
- domain assumption Correct anchors and cluster hypotheses in the simulation studies
- ad hoc to paper Variational Bayes hard-selection with Jacobian-rank effective parameter count is a valid operating criterion
Cite this review
Pith. "Pith review of When Is a General Factor Distinguishable? Non-Proportionality, Stable Structure, and the Bifactor Decision." pith.science (2026). https://pith.science/paper/6QA3E7X6
@misc{pith2026260810731,
author = {Pith},
title = {Pith review of: When Is a General Factor Distinguishable? Non-Proportionality, Stable Structure, and the Bifactor Decision},
year = {2026},
howpublished = {\url{https://pith.science/paper/6QA3E7X6}},
note = {Machine review of arXiv:2608.10731}
}
abstract
Whether an additional general dimension is necessary beyond correlated first-order factors is a property of the population covariance matrix, not of any estimator or design. This research establishes when that property can be decided. Where the general and group loadings are proportional within every cluster the bifactor structure is covariance-equivalent to correlated factors, so no sample size separates them (Proposition 1); where that proportionality fails in every cluster, three items per cluster and some mild regularities leave no $K$-factor model with diagonal uniquenesses able to reproduce the covariance matrix (Theorem 1); and between them lies a mixed boundary, located numerically here and turning on cluster resistance. Distinguishability is therefore graded, measured by the population distance to the $K$-factor class. Because that question is conditional on a first-order structure which is itself uncertain, a two-step procedure is developed within partially exploratory factor analysis, delivering a structure only when it reproduces across adjacent counts and treating non-delivery as legitimate. Simulation shows that a unanimous count can accompany a structure that fails to reproduce, and that absorbed local dependence can imitate a general factor, the error growing with sample size while stability indicators stay clean. Four empirical datasets illustrate the possible outcomes.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
Auerswald, M., & Moshagen, M. (2019). How to determine the number of factors to retain in exploratory factor analysis: A comparison of extraction methods under realistic conditions. Psychological Methods, 24(4), 468--491
work page 2019
-
[3]
Bonifay, W., & Cai, L. (2017). On the complexity of item response theory models. Multivariate Behavioral Research, 52(4), 465--484
work page 2017
-
[4]
Braeken, J., & Assen, M. A. L. M. van. (2017). An empirical Kaiser criterion. Psychological Methods, 22(3), 450--466
work page 2017
-
[5]
Chen, J. (2020). A partially confirmatory approach to the multidimensional item response theory with the Bayesian Lasso . Psychometrika, 85(3), 738--774
work page 2020
-
[6]
Chen, J. (2021). A generalized partially confirmatory factor analysis framework with mixed Bayesian Lasso methods. Multivariate Behavioral Research, 57(6), 879--894. https://doi.org/10.1080/00273171.2021.1925520
arXiv 2021
-
[7]
Chen, J. (2022). Partially confirmatory approach to factor analysis with Bayesian learning: A LAWBL tutorial. Structural Equation Modeling: A Multidisciplinary Journal, 29(5), 800--816. https://doi.org/10.1080/10705511.2022.2039660
arXiv 2022
-
[8]
Chen, J. (2023). Fully and partially exploratory factor analysis with bi-level Bayesian regularization. Behavior Research Methods, 55(4), 2125--2142. https://doi.org/10.3758/s13428-022-01884-7
Show all 40 references
-
[9]
Chen, J., Guo, Z., Zhang, L., & Pan, J. (2021). A partially confirmatory approach to scale development with the Bayesian Lasso . Psychological Methods, 26(2), 210--235. https://doi.org/10.1037/met0000293
2021 doi
- [10]
-
[11]
Chen, J., & Jin, Y. (2026b). Vbpm: Variational Bayes psychometric models . https://github.com/Jinsong-Chen/vbpm
2026
-
[12]
M., & Revelle, W
Condon, D. M., & Revelle, W. (2014). The International Cognitive Ability Resource : Development and initial validation of a public-domain measure. Intelligence, 43, 52--64
2014
-
[13]
Cucina, J., & Byle, K. (2017). The bifactor model fits better than the higher-order model in more than 90\ test batteries. Journal of Intelligence, 5(3), 27
2017
-
[14]
Eid, M., Geiser, C., Koch, T., & Heene, M. (2017). Anomalous results in g-factor models: Explanations and alternatives. Psychological Methods, 22(3), 541--562
2017
-
[15]
Fang, G., Guo, J., Xu, X., Ying, Z., & Zhang, S. (2021). Identifiability of bifactor models. Statistica Sinica, 31, 2309--2330
2021
-
[16]
J., & Garrido, L
Garcia-Garzon, E., Abad, F. J., & Garrido, L. E. (2019). Improving bi-factor exploratory modeling: Empirical target rotation based on loading differences. Methodology, 15(2), 45--55. https://doi.org/10.1027/1614-2241/a000163
2019 doi
-
[17]
Gignac, G. E. (2016). The higher-order model imposes a proportionality constraint: That is why the bifactor model tends to fit better. Intelligence, 55, 57--68
2016
-
[18]
J., & Swineford, F
Holzinger, K. J., & Swineford, F. (1939). A study in factor analysis: The stability of a bi-factor solution. University of Chicago, Department of Education, Supplementary Educational Monographs No. 48
1939
-
[19]
Horn, J. L. (1965). A rationale and test for the number of factors in factor analysis. Psychometrika, 30, 179--185
1965
-
[20]
I., & Bentler, P
Jennrich, R. I., & Bentler, P. M. (2011). Exploratory bi-factor analysis. Psychometrika, 76(4), 537--549
2011
-
[21]
I., & Bentler, P
Jennrich, R. I., & Bentler, P. M. (2012). Exploratory bi-factor analysis: The oblique case. Psychometrika, 77(3), 442--454
2012
-
[22]
J., Garcia-Garzon, E., & Garrido, L
Jiménez, M., Abad, F. J., Garcia-Garzon, E., & Garrido, L. E. (2023). Exploratory bi-factor analysis with multiple general factors. Multivariate Behavioral Research, 58(6), 1072--1089. https://doi.org/10.1080/00273171.2023.2189571
2023
-
[23]
Jin, Y., & Chen, J. (2025a). Regularized variational approximation for partially confirmatory factor analysis. Structural Equation Modeling: A Multidisciplinary Journal, 32(3), 437--449. https://doi.org/10.1080/10705511.2024.2432612
2025
-
[24]
Jin, Y., & Chen, J. (2025b). Regularized variational Bayesian approximations for variable selection in extended multiple-indicators multiple-causes models. Multivariate Behavioral Research, 60(5), 859--877. https://doi.org/10.1080/00273171.2025.2483253
2025
-
[25]
Jin, Y., Chen, J., Yan, Z., & Zhang, Y. (2026). Sparse residual estimation in partially confirmatory factor analysis. PsyArXiv preprint. https://doi.org/10.31234/osf.io/dehtv_v2
2026 doi
-
[26]
Kaiser, H. F. (1960). The application of electronic computers to factor analysis. Educational and Psychological Measurement, 20, 141--151
1960
-
[27]
F., Derringer, J., Markon, K
Krueger, R. F., Derringer, J., Markon, K. E., Watson, D., & Skodol, A. E. (2012). Initial construction of a maladaptive personality trait model and inventory for DSM-5 . Psychological Medicine, 42(9), 1879--1890
2012
-
[28]
Mansolf, M., & Reise, S. P. (2017). When and why the second-order and bifactor models are distinguishable. Intelligence, 61, 120--129
2017
-
[29]
L., & Johnson, W
Murray, A. L., & Johnson, W. (2013). The limitations of model fit in comparing the bi-factor versus higher-order models of human cognitive ability structure. Intelligence, 41(5), 407--422
2013
-
[30]
Peeters, C. F. W. (2012). Rotational uniqueness conditions under oblique factor correlation metric. Psychometrika, 77(2), 288--292
2012
-
[31]
Qiao, J., Chen, Y., & Ying, Z. (2025a). Exact exploratory bi-factor analysis: A constraint-based optimization approach. Psychometrika, 90(3), 998--1013. https://doi.org/10.1017/psy.2025.17
2025 doi
-
[32]
Qiao, J., Chen, Y., & Ying, Z. (2025b). Exploratory hierarchical factor analysis with an application to psychological measurement. arXiv:2505.09043
2025
-
[33]
Raykov, T., DiStefano, C., & Calvocoressi, L. (2024). A note on comparing the bifactor and second-order factor models: Is the Bayesian information criterion a routinely dependable index for model selection? Educational and Psychological Measurement, 84(2), 271--288. https://do...
2024 doi
-
[34]
Reise, S. P. (2012). The rediscovery of bifactor measurement models. Multivariate Behavioral Research, 47(5), 667--696
2012
-
[35]
Revelle, W. (2024). Psych: Procedures for psychological, psychometric, and personality research. Northwestern University
2024
-
[36]
Roskam, I., Galdiolo, S., Hansenne, M., Massoudi, K., Rossier, J., Gicquel, L., & Rolland, J.-P. (2015). The psychometric properties of the French version of the Personality Inventory for DSM-5 . PLoS ONE, 10(7), e0133413
2015
-
[37]
Schmid, J., & Leiman, J. M. (1957). The development of hierarchical factor solutions. Psychometrika, 22(1), 53--61
1957
-
[38]
Waller, N. G. (2018). Direct Schmid--Leiman transformations and rank-deficient loadings matrices. Psychometrika, 83(4), 858--870
2018
-
[39]
Yung, Y.-F., Thissen, D., & McLeod, L. D. (1999). On the relationship between the higher-order factor model and the hierarchical factor model. Psychometrika, 64(2), 113--128
1999
-
[40]
Zhang, Y., & Chen, J. (2024). Accommodating and extending various models for special effects within the generalized partially confirmatory factor analysis framework. Applied Psychological Measurement, 48(4-5), 208--229. https://doi.org/10.1177/01466216241261704 CSLReferences *...
2024 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.