REVIEW 2 major objections 7 minor 238 references
Gain rule on variational bound recovers true factor count
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-09 18:35 UTC pith:HKWL6CEU
load-bearing objection Solid, practical methodology paper. The hard/soft selection framework and gain rule are genuinely useful. Main concern is the variational gap assumption, but it's bounded. the 2 major comments →
Recovering Latent Structures after Variational Bayesian Variable Selection: Fit Assessment and Factor-Number Selection in Partially Exploratory Factor Analysis
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The scale-free gain rule with sustained-drop guard converts the over-factoring tendency of raw relative criteria into accurate dimensionality recovery by thresholding marginal criterion gains against the largest gain in the window rather than locating a global optimum. Applied to the ELBO, this rule recovers the true number of factors across a wide range of nuisance structure, sample sizes, and specification levels where raw ELBO, AIC, and BIC optima all drift upward. The mechanism is that genuine factors produce large marginal improvements while nuisance factors produce negligible ones, so a ratio threshold on gains cleanly separates the two without requiring any calibrated absolute cutoff.
What carries the argument
The gain rule computes marginal gains g(K) = c(K) - c(K-1) for a criterion c oriented so larger is better, finds the maximum gain g_max, and selects the last K whose gain exceeds delta% of g_max, with a sustained-drop guard requiring s-1 subsequent gains to also fall below the threshold. Hard selection thresholds inclusion probabilities at tau (default 0.50) into a binary loading pattern with nominal or rank-adjusted parameter counts; soft selection retains probabilities as effective weights yielding posterior-expected parameter counts analogous to p_D in DIC or p_WAIC.
Load-bearing premise
The ELBO gain rule depends on the variational gap the difference between the ELBO and the true log evidence being roughly comparable across candidate models with different numbers of factors. If this gap shrinks or grows systematically with K, the marginal ELBO gains no longer reflect genuine evidence gains, and the rule's accuracy in the simulations could be partly coincidental rather than a property of the criterion itself.
What would settle it
A simulation or empirical setting where the variational gap varies systematically with K, causing the ELBO gain path to separate at the wrong dimensionality. More directly: if the gain rule were applied to exact log-marginal-likelihood gains (removing the variational approximation) and selected a different K than the ELBO gain rule in a substantial fraction of cases, the variational gap assumption would be falsified.
If this is right
- The gain rule principle is criterion-agnostic and could be applied to any model-selection path where adding parameters yields diminishing returns, including mixture model component selection, network edge selection, or tree depth in gradient boosting.
- The hard-soft parameter-count gap (the difference between thresholded and effective parameter counts) serves as a built-in diagnostic for selection uncertainty that could be reported alongside any fit index to warn analysts when thresholding discards substantial posterior information.
- The finding that stronger specification can reduce selection accuracy because it frees added factors to capture genuine nuisance covariance suggests that partial specification has a non-monotonic relationship with recovery quality, which has implications for how much prior structure to impose in exploratory settings.
- The identification boundary that unspecified collinear clusters are resolved arbitrarily while well-separated major factors are recovered robustly provides a formal target for future work on partial identification in exploratory latent variable models.
Where Pith is reading between the lines
- If the variational gap (KL divergence between the variational approximation and the true posterior) varies systematically with the number of factors, ELBO differences across K values would not reflect genuine evidence differences, and the gain rule's accuracy could be a consequence of the gap's behavior rather than the criterion's fidelity to model evidence. The paper acknowledges this is untestab
- The gain rule's sustained-drop guard with s=2 costs one extra factor of over-factoring headroom. In settings where factor estimation is expensive (e.g., large-scale genomic or neuroimaging data), the optimal s may trade off computational cost against robustness differently than in the simulations studied.
- The finding that CFI and TLI cleanly separate under-factoring from adequate factoring while RMSEA and SRMR are uninformative at high dimensions (J=80) suggests that conventional cutoff recommendations for RMSEA and SRMR may need recalibration as the number of variables grows, independent of the variable-selection machinery proposed here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops post-selection fit assessment tools for the regularized variational approximation for partially confirmatory factor analysis (PCFA-VA), extending its use to partially exploratory factor analysis (PEFA) where both the loading structure and the number of factors are weakly specified. The core contributions are: (1) converting converged variational solutions into covariance models via hard selection (thresholding inclusion probabilities) or soft selection (retaining them as effective parameter weights); (2) deriving corresponding degrees of freedom, absolute fit indices (RMSEA, SRMR, CFI, TLI), and relative criteria (AIC, BIC, ELBO); and (3) proposing a scale-free gain rule with a sustained-drop guard for selecting the number of factors. Two simulation studies calibrate the tools, and an empirical analysis of the 100-item PID-5 demonstrates the workflow. The gain rule, applied to the ELBO path, is shown to recover true dimensionality far more robustly than raw criteria, which systematically over-factor.
Significance. The paper addresses a genuine methodological gap: Bayesian variable selection in factor analysis produces inclusion probabilities, but the field lacks principled post-selection fit statistics that respect the selection step. The hard/soft selection distinction and the connection to effective degrees of freedom (p_D, p_WAIC, generalized d.f.) is well-motivated and correctly positioned in the literature. The scale-free gain rule with a formal correctness condition (Proposition 1) is a clean, falsifiable contribution. The reproducible code and simulation archives, the exact bracketing condition, and the disjoint-backbone robustness test in the empirical example are all commendable. The finding that the ELBO gain is most robust under heavy nuisance, while raw criteria collapse, is practically important for researchers working with large personality inventories.
major comments (2)
- The ELBO comparability assumption (Eqs. 31-32 and surrounding text) is the most load-bearing assumption in the paper. ELBO differences equal log-evidence differences minus differences in the KL divergence term across candidate K values. The paper states this is 'untestable directly' and relies on simulations. Two specific concerns: (1) The simulations use data generated from linear Gaussian factor models fit with normal-theory variational inference—a well-specified regime where the variational approximation is likely good and the KL gap relatively stable. The PID-5 empirical application uses 4-category items treated as continuous, introducing misspecification that could shift the KL gap behavior across K. (2) In Study 2 under heaviest nuisance (3w+3m), ELBO gain accuracy drops from 99.3% to 88.3% (Appendix A, Table A1 note). The paper does not disentangle whether this degradation stems (
- The plug-in nature of the fit statistics (Section on Absolute and Relative Fit Indices) creates a two-sided bias that is acknowledged but not fully bounded. The paper notes that the plug-in direction is conservative (shrunken variational estimates bound the minimized discrepancy from above) while the selection direction is optimistic (hard-selected d.f. are conditional on a data-chosen pattern). The hard-soft parameter-count gap is offered as a diagnostic, not a correction. However, the paper does not provide any simulation evidence on the magnitude of the post-selection optimism in T_H (Eq. 24) or its downstream effect on RMSEA_H, CFI_H, TLI_H. Since these indices are proposed as the primary reporting tools, even a brief simulation reporting the mean and SD of the gap between the plug-in T_H and the ML-refit T_H on the same selected pattern would substantially strengthen the claim that
minor comments (7)
- The notation switches between PCFA-VA and PCFA V A (with a space) throughout the manuscript. Consistent use of PCFA-VA would improve readability.
- In the paragraph following Eq. (6), the notation qvar(·) is introduced as the variational density, but the subscript notation in bπ_jk and bλ_jk uses a hat that is not defined. Clarifying that the hat denotes posterior quantities would help.
- Table 1: the K_fit=6 row shows RMSE=0.073, which is higher than K_fit=5 (0.052) but lower than K_fit=4 (0.109). The text attributes this to over-factoring absorbing nuisance, but the non-monotonicity (K_fit=7 has RMSE=0.077, K_fit=8 has 0.085, then K_fit=9 has 0.082) could use a brief explanation.
- Figure 5A: the y-axis label 'ELBO marginal gains g(K)' could be more informative if the units (nats) were indicated, as the text references '≈700 nats' but the figure axis is unlabeled in this respect.
- The default δ=10 is described as sitting at the 'under-selection-averse end of the stable region,' but Table A2 shows that for backbone a, δ=25 moves the selection from 22 to 20. A brief discussion of why δ=10 rather than, say, δ=15 (which also selects 22 for both backbones) is the default would be helpful.
- The paper cites Jin & Chen (2025) for the PCFA-VA method, which shares authors with the present paper. This is standard and appropriate, but a brief note in the introduction explicitly stating that the present paper extends the authors' prior work would be transparent.
- Appendix Table B1: the ELBO values are negative (e.g., -313,321), which is expected for a log-likelihood bound, but a note clarifying that these are negative and that larger (less negative) is better would prevent momentary confusion.
Circularity Check
No significant circularity: the gain rule and fit indices are validated against known truth in simulations, not against the authors' own prior results.
full rationale
The paper's central contributions—the post-selection fit indices (RMSEA, CFI, TLI, AIC, BIC, ELBO), the hard/soft selection framework, and the scale-free gain rule with sustained-drop guard—are logically independent of the estimation method's correctness. The gain rule (Eq. 37) is a general elbow-detection heuristic applied to any criterion path; its performance is validated against known true dimensionality in two simulation studies (99.4% and 94.8% accuracy), not against the authors' own prior results. The fit indices are derived from standard SEM formulas (Eqs. 23-36) applied to the converged variational solution, with parameter counts defined independently (Eqs. 12-22). The self-citation to Jin & Chen (2025) for the PCFA-VA estimation machinery is not load-bearing for the present paper's claims: the post-selection assessment framework would work with any variable-selection method that produces inclusion probabilities. The ELBO (Eq. 31-32) is the algorithm-native objective whose properties follow from variational inference theory generally, not from the authors' prior work. Proposition 1 (Appendix A) provides an independent correctness condition for the gain rule based on gain separation, with a self-contained proof. No step in the derivation chain reduces to its own inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (5)
- tau (threshold) =
0.50
- delta (gain threshold) =
10
- s (sustained-drop guard) =
2
- v0, v1 (spike/slab variances) =
not specified in text
- rho (Bernoulli inclusion prob) =
0.5
axioms (4)
- domain assumption Variational gap comparability across candidate K
- domain assumption Local identification of selected loading pattern
- domain assumption Pseudo chi-square interpretation of T_H
- domain assumption Gain separation at true K
read the original abstract
In partially exploratory factor analysis (PEFA), the loading structure and factor numbers are weakly specified. The regularized variational approximation for partially confirmatory factor analysis (PCFA VA) recovers this structure via Bayesian variable selection, using spike and slab priors to assign inclusion probabilities to unspecified loadings. This research introduces a post selection assessment framework for this approach. We convert converged solutions into covariance models using either hard selection (thresholding probabilities into a sparse pattern) or soft selection (retaining them as weights for effective parameter counts). We derive the resulting degrees of freedom, absolute fit diagnostics (RMSEA, SRMR, CFI, TLI), and relative criteria (AIC, BIC, ELBO). To determine factor numbers, we propose a scale free gain rule with a sustained drop guard. Simulations show absolute indices successfully track loading recovery and flag under factoring. While raw criteria over factor, our gain rule accurately recovers true dimensionality, with the ELBO variant proving most robust. Finally, a 100 item PID 5 example demonstrates that our model fits better than a confirmatory 25 facet model and concordantly recovers major structures across disjoint specifications.
Reference graph
Works this paper leans on
-
[1]
and Derringer, Jaime and Markon, Kristian E
Krueger, Robert F. and Derringer, Jaime and Markon, Kristian E. and Watson, David and Skodol, Andrew E. , journal=. Initial construction of a maladaptive personality trait model and inventory for. 2012 , publisher=
work page 2012
-
[2]
The SAPA Personality Inventory: An empirically-derived, hierarchically-organized self-report personality assessment model , author=. PsyArXiv preprint , year=. doi:10.31234/osf.io/sc4p9 , publisher=
-
[3]
psychTools: Tools to Accompany the 'psych' Package for Psychological Research , author=. 2024 , note=
work page 2024
-
[4]
A generalized partially confirmatory factor analysis framework with mixed
Chen, Jinsong , journal=. A generalized partially confirmatory factor analysis framework with mixed. 2022 , publisher=
work page 2022
-
[5]
A partially confirmatory approach to scale development with the
Chen, Jinsong and Guo, Zhihan and Zhang, Lijin and Pan, Junhao , journal=. A partially confirmatory approach to scale development with the. 2021 , publisher=
work page 2021
-
[6]
Journal of machine Learning research , volume=
Latent dirichlet allocation , author=. Journal of machine Learning research , volume=
-
[7]
BERTopic: Neural topic modeling with a class-based TF-IDF procedure
BERTopic: Neural topic modeling with a class-based TF-IDF procedure , author=. arXiv preprint arXiv:2203.05794 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[8]
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Bert: Pre-training of deep bidirectional transformers for language understanding , author=. arXiv preprint arXiv:1810.04805 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[9]
Advances in neural information processing systems , volume=
Language models are few-shot learners , author=. Advances in neural information processing systems , volume=
-
[10]
Multimedia tools and applications , volume=
Latent Dirichlet allocation (LDA) and topic modeling: models, applications, a survey , author=. Multimedia tools and applications , volume=. 2019 , publisher=
work page 2019
-
[11]
A review of topic modeling methods , author=. Information Systems , volume=. 2020 , publisher=
work page 2020
-
[12]
ACM computing surveys (csur) , volume=
Generalizing from a few examples: A survey on few-shot learning , author=. ACM computing surveys (csur) , volume=. 2020 , publisher=
work page 2020
- [13]
-
[14]
Braeken, Johan and Van Assen, Marcel ALM , journal=. An empirical. 2017 , publisher=
work page 2017
- [15]
-
[16]
Principles and practice of structural equation modeling , author=. 2023 , publisher=
work page 2023
-
[17]
Transactions of the Association for Computational Linguistics , volume=
A primer in BERTology: What we know about how BERT works , author=. Transactions of the Association for Computational Linguistics , volume=. 2021 , publisher=
work page 2021
-
[18]
Springer google schola , volume=
Pattern recognition and machine learning , author=. Springer google schola , volume=
-
[19]
Fundamentals of artificial intelligence , pages=
Natural language processing , author=. Fundamentals of artificial intelligence , pages=. 2020 , publisher=
work page 2020
-
[20]
Natural language processing with Python: analyzing text with the natural language toolkit , author=. 2009 , publisher=
work page 2009
-
[21]
Upper Saddle River, NJ: Prentice Hall , year=
Speech and Language Processing: An introduction to speech recognition, computational linguistics and natural language processing , author=. Upper Saddle River, NJ: Prentice Hall , year=
-
[22]
Transformers for Natural Language Processing: Build innovative deep neural network architectures for NLP with Python, PyTorch, TensorFlow, BERT, RoBERTa, and more , author=. 2021 , publisher=
work page 2021
-
[23]
Text analytics with Python: a practitioner's guide to natural language processing , author=. 2019 , publisher=
work page 2019
-
[24]
LLaMA: Open and Efficient Foundation Language Models
Llama: Open and efficient foundation language models , author=. arXiv preprint arXiv:2302.13971 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[25]
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Roberta: A robustly optimized bert pretraining approach , author=. arXiv preprint arXiv:1907.11692 , year=
work page internal anchor Pith review Pith/arXiv arXiv 1907
-
[26]
Advances in neural information processing systems , volume=
Attention is all you need , author=. Advances in neural information processing systems , volume=
- [27]
-
[28]
Multivariate Behavioral Research , volume=
An overview of analytic rotation in exploratory factor analysis , author=. Multivariate Behavioral Research , volume=. 2001 , publisher=
work page 2001
-
[29]
Proceedings of the 27th international conference on computational linguistics , pages=
Contextual string embeddings for sequence labeling , author=. Proceedings of the 27th international conference on computational linguistics , pages=
-
[30]
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension , author=. arXiv preprint arXiv:1910.13461 , year=
work page internal anchor Pith review Pith/arXiv arXiv 1910
-
[31]
Utilizing BERT for Aspect-Based Sentiment Analysis via Constructing Auxiliary Sentence
Utilizing BERT for aspect-based sentiment analysis via constructing auxiliary sentence , author=. arXiv preprint arXiv:1903.09588 , year=
work page internal anchor Pith review Pith/arXiv arXiv 1903
-
[32]
Rethinking complex neural network architectures for document classification , author=. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , pages=
work page 2019
-
[33]
Keyphrase extraction as sequence labeling using contextualized embeddings , author=. Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14--17, 2020, Proceedings, Part II 42 , pages=. 2020 , organization=
work page 2020
- [34]
-
[35]
The text mining handbook: advanced approaches in analyzing unstructured data , author=. 2007 , publisher=
work page 2007
-
[36]
NLTK: The Natural Language Toolkit
Nltk: The natural language toolkit , author=. arXiv preprint cs/0205028 , year=
work page internal anchor Pith review Pith/arXiv arXiv
-
[37]
Robustness in statistics , pages=
Robustness in the strategy of scientific model building , author=. Robustness in statistics , pages=. 1979 , publisher=
work page 1979
-
[38]
Exploratory bi-factor analysis , author=. Psychometrika , volume=. 2011 , publisher=
work page 2011
-
[39]
The development of hierarchical factor solutions , author=. Psychometrika , volume=. 1957 , publisher=
work page 1957
-
[40]
Multivariate Behavioral Research , volume=
The rediscovery of bifactor measurement models , author=. Multivariate Behavioral Research , volume=. 2012 , publisher=
work page 2012
-
[41]
International Conference on Learning Representations , year=
Long Range Arena : A Benchmark for Efficient Transformers , author=. International Conference on Learning Representations , year=
-
[42]
George, Edward I and McCulloch, Robert E , journal=. Approaches for. 1997 , publisher=
work page 1997
-
[43]
Guoqiang, Lin , journal=. A
-
[44]
The annals of mathematical statistics , volume=
On information and sufficiency , author=. The annals of mathematical statistics , volume=. 1951 , publisher=
work page 1951
-
[45]
Blei, David M and Kucukelbir, Alp and McAuliffe, Jon D , journal=. Variational inference:. 2017 , publisher=
work page 2017
-
[46]
Automatic text scoring using neural networks , author=. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=. 2016 , organization=
work page 2016
-
[47]
International Journal of Advanced Computer Science and Applications , volume=
An empirical analysis of BERT embedding for automated essay scoring , author=. International Journal of Advanced Computer Science and Applications , volume=
-
[48]
Meaning modulations and stability in large language models: An analysis of BERT embeddings for psycholinguistic research , author=. PNAS Nexus , year=
-
[49]
Semantics derived automatically from language corpora contain human-like biases , author=. Science , volume=
-
[50]
Leveraging the alignment between machine learning and intersectionality: Using word embeddings to measure intersectional experiences of the nineteenth century U.S. South , author=. Poetics , volume=
-
[51]
Arshad, F. and Shaukat, A. and Latif, S. , title =. Proceedings of the 2024 International Conference on Frontiers of Information Technology , year =
work page 2024
-
[52]
Christian, H. and Suhartono, D. and Chowanda, A. and Zamli, K. Z. , title =. Journal of Big Data , volume =
- [53]
-
[54]
Zhou, X. and Yang, L. and Fan, X. and Ren, G. and Yang, Y. and Lin, H. , title =. Lecture Notes in Computer Science , year =
-
[55]
Huang, Z. and Long, Y. and Peng, K. and Tong, S. , title =. Journal of Intelligence , volume =
-
[56]
Wulff, D. U. and Mata, R. , title =. Nature Human Behaviour , year =
-
[57]
Rheault, Luke and Cochrane, Christopher , title =. Political Analysis , volume =
-
[58]
Proceedings of the National Academy of Sciences , volume =
Garg, Nikhil and Schiebinger, Londa and Jurafsky, Dan and Zou, James , title =. Proceedings of the National Academy of Sciences , volume =
-
[59]
Applied and Computational Engineering , volume =
Shu, Xiaoxue , title =. Applied and Computational Engineering , volume =
-
[60]
Faseeh, Muneeb and Jaleel, Abdul and Iqbal, Nadeem and Ghani, Abdullah and Abdusalomov, Abduvohid and Mehmood, Asad and Cho, Young , title =. Mathematics , volume =
-
[61]
Rytting, C. M. and Sorensen, T. and Argyle, L. and Busby, E. and Fulda, N. and Gubler, J. and Wingate, D. , title =. arXiv preprint , year =
-
[62]
Li, Lin and Li, Jincheng and Chen, Cheng and Gui, Feng and Yang, Hongyi and Yu, Chenyu and Wang, Zefan and others , title =. arXiv preprint , year =
- [63]
-
[64]
Petukhova, Anna and Matos-Carvalho, Joao Pedro and Fachada, Nuno , title =. arXiv preprint , year =
-
[65]
Nie, Zihan and Feng, Zhen and Li, Ming and Zhang, Cheng and Zhang, Yichen and Long, Donghong and Zhang, Rui , title =. arXiv preprint , year =
-
[66]
Tao, Chao and Shen, Tao and Gao, Shangsong and Zhang, Jun and Li, Zhiyuan and Tao, Zhanxing and Ma, Shuming , title =. arXiv preprint , year =
-
[67]
Structural Equation Modeling: A Multidisciplinary Journal , volume=
Regularized Variational Approximation for Partially Confirmatory Factor Analysis , author=. Structural Equation Modeling: A Multidisciplinary Journal , volume=. 2025 , publisher=
work page 2025
-
[68]
Exploring Psychometric Analysis of Textual Data with Large Language Models: Chances and Challenges , author=. 2024 , publisher=
work page 2024
-
[69]
Yihui Xie , publisher =. Dynamic Documents with. 2015 , edition =
work page 2015
- [70]
-
[71]
Empirical Model-Building and Response Surfaces , publisher =. 1987 , author =
work page 1987
- [72]
-
[73]
Investigating variation in replicability:
Klein, Richard A and Ratliff, Kate A and Vianello, Michelangelo and Adams Jr, Reginald B and Bahn. Investigating variation in replicability:. Social psychology , year=
-
[74]
Data analysis: A model comparison approach , author=. 2011 , publisher=
work page 2011
-
[75]
Doing Bayesian data analysis: A tutorial with R, JAGS, and Stan , author=. 2014 , publisher=
work page 2014
-
[76]
Wellcome Open Research , VOLUME =
Allen, M and Poggiali, D and Whitaker, K and Marshall, T R and Kievit, R A , TITLE =. Wellcome Open Research , VOLUME =. 2019 , NUMBER =
work page 2019
-
[77]
The probable error of a mean , author=. Biometrika , pages=. 1908 , publisher=
work page 1908
-
[78]
On student's 1908 article "the probable error of a mean"" , author=. Journal of the American Statistical Association , volume=. 2008 , publisher=
work page 1908
-
[79]
Sample Quantiles in Statistical Packages , volume =. The American Statistician , author =. 1996 , pages =. doi:10.1080/00031305.1996.10473566 , language =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.1080/00031305.1996.10473566 1996
-
[80]
Psychonomic bulletin & review , volume=
The fallacy of placing confidence in confidence intervals , author=. Psychonomic bulletin & review , volume=. 2016 , publisher=
work page 2016
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.