REVIEW 4 major objections 5 minor 20 references
HLE's text-only multiple-choice benchmark measures a single general reasoning factor: its eight subject-domain labels explain only 3.5% of item-response variance, and measurement precision drops sharply at the ability levels where frontier
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 07:35 UTC pith:RJOLMVAP
load-bearing objection First psychometric audit of HLE's MCQ subset is worth reading, but the headline unidimensionality statistic is a reliability index mislabeled as omega_h, so the central claim is weaker than it looks. the 4 major comments →
Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that on the text-only multiple-choice subset of HLE (428 items, 29 models), item responses are dominated by a single general factor: McDonald's omega_h = 0.998 (95% CI [0.998, 0.999]); domain labels explain only 3.5% of variance in the first three principal components; after removing the general factor, within-domain and between-domain residual correlations are nearly identical (Cohen's d = 0.016); and for the four well-powered domains, domain ability estimates correlate with total ability at r >= 0.81. Therefore, the eight subject-domain labels do not correspond to empirically distinct latent constructs. A second claim is that the test information function peaks at theta =
What carries the argument
The central machinery is a two-parameter logistic (2PL) item response theory model, where the probability that a model answers an item correctly depends on latent ability theta, item discrimination a_j, and item difficulty b_j. Those item parameters feed two downstream analyses: a dimensionality analysis using McDonald's omega_h (computed from factor loadings lambda_j = a_j / sqrt(1 + a_j^2)), PCA on the transposed item-by-model response matrix, residual correlations after subtracting the general-factor-implied correlation matrix, and correlations between domain-specific and total ability estimates; and a measurement-precision analysis using the test information function I(theta) = sum_j a_j
Load-bearing premise
The 29 language models scored on the 428 items are treated as a sufficient sample to estimate all 428 items' difficulty and discrimination parameters—if this small, non-random model sample cannot identify those parameters, the apparent single-factor structure and the location of the precision peak could be artifacts of which models were tested.
What would settle it
A concrete check: re-score the same 428 multiple-choice items with a much larger and more diverse model sample (e.g., more than 100 models spanning several families and ability levels), refit the 2PL model, and compute McDonald's omega_h and the domain R^2. If omega_h drops below about 0.95, or if a confirmatory bifactor model yields a second factor with nontrivial variance, or if the test-information peak moves above theta = 0, the single-factor and precision-limitation claims would be refuted.
If this is right
- HLE's multiple-choice subscores should not be reported or read as profiles of separable capability; the total score carries essentially all psychometric information.
- Score gaps between top models on HLE's multiple-choice subset are poorly supported by the instrument, because measurement information drops sharply above theta = 0.
- Engineering items, the smallest domain (~4%), provide the largest share of frontier information, while Mathematics and Humanities/Social Science—the largest domains—contribute least to separating strong models.
- As model ability rises, HLE's ability to rank models will degrade further unless new, harder, high-discrimination items are added.
Where Pith is reading between the lines
- A direct extension the paper leaves implicit: if the unidimensionality result holds on the full HLE including exact-match and image items, then HLE's overall score is the only defensible summary and per-domain leaderboard bars are systematically misleading.
- A testable extension: scoring each model multiple times per item (e.g., sampling temperature or decoding seeds) would reveal whether within-model response variability changes the item discriminations and the location of the test-information peak; the current design uses one deterministic answer per model-item pair.
- Because the model sample is small and non-random, the TIF's peak near theta = -0.35 is tied to the ability distribution of these 29 models; adding many weaker or stronger models could shift the measured precision curve even if HLE's item properties are unchanged.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates 29 LLMs on the text-only multiple-choice subset of Humanity's Last Exam (428 items) and fits a two-parameter logistic IRT model to the binary response matrix. Its central claims are (i) that HLE's eight subject-domain labels do not correspond to empirically distinct latent constructs, with McDonald's hierarchical omega = 0.998, domain labels explaining 3.5% of item-response variance, near-identical within- and between-domain residual correlations, and near-redundant domain-specific ability estimates; and (ii) that HLE's measurement precision concentrates at moderate ability levels, dropping sharply above theta = 0 where frontier models sit. The paper concludes that HLE domain subscores should not be interpreted as distinct capability profiles and that the benchmark has limited ability to differentiate among the strongest models.
Significance. The question addressed is important: HLE is widely used for capability evaluation, and the dimensionality and measurement precision of the benchmark directly affect the validity of domain-level claims. If the unidimensionality result were well supported, this would be a useful caution for leaderboard interpretation. The paper also provides a reproducible analysis pipeline and openly acknowledges several limitations. However, the central evidentiary chain is currently not sound: the primary unidimensionality statistic is mis-specified, one of the residual-correlation analyses combines incompatible correlation metrics, and the model sample is very small relative to the number of item parameters. These issues undermine the strongest claims, though the TIF analysis and the descriptive domain-difficulty results remain of interest.
major comments (4)
- [§2.6.2, Eq. (2)] The quantity labeled 'McDonald's hierarchical omega' is not the hierarchical omega statistic. Equation (2), ω = (Σλ_j)² / [(Σλ_j)² + Σ(1−λ_j²)], is the coefficient omega (total reliability) formula for a unidimensional congeneric model. Hierarchical omega requires a bifactor or nested-factor model that explicitly separates a general factor from group factors. Here the 2PL model in Eq. (1) assumes a single latent dimension, so the λ_j absorb all common variance; the ratio will approach 1 for any highly reliable one-factor composite, regardless of whether multidimensional structure exists. Thus ω_h = 0.998 is not evidence of unidimensionality; it is largely a restatement of the unidimensional model assumption. The bootstrap CIs do not address this because the model is refit under the same one-factor restriction.
- [§2.6.2, residual item correlations] The residual-correlation analysis subtracts a model-implied correlation matrix R_g = λλ^T, which lies on the tetrachoric (latent-continuous) scale, from Pearson correlations computed directly on binary item response vectors. These two quantities are on different scales: Pearson correlations on binary items are attenuated relative to tetrachoric correlations, and the degree of attenuation depends on item difficulty and marginal response rates. The paper acknowledges negative residuals but argues attenuation affects within- and between-domain pairs equally. That is not generally true: domains with different difficulty distributions (e.g., engineering items with median b = 1.65 vs. computer science items with median b = 0.10) will be attenuated differently. Consequently, the near-zero Cohen's d = 0.016 cannot be interpreted as evidence against domain-level group factors.
- [§2.5, §3.1] The analysis estimates 856 item parameters (a_j, b_j) plus 29 ability parameters from a 29×428 response matrix, after removing 85 zero-variance items and coding 0.79% missing observations as incorrect. At N = 29, marginal maximum likelihood estimates of 2PL discrimination parameters are known to be unstable and can be severely biased upward, especially for items with extreme difficulties. The paper reports no standard errors, no simulation study, and no sensitivity analysis with respect to the zero-variance-item removal or the missing-data recoding. Since the omega, residual-correlation, and TIF results all propagate these item parameters, the central quantitative claims are not supported by the evidence presented. The authors' own limitation statement concedes the small N, but the conclusion that 'IRT discrimination parameters proved stable enough' is asserted without supporting diagnos
- [§3.2, Domain-level θ correlations] The near-unity correlations between domain-specific ability estimates and total ability estimates are presented as evidence of a single factor, but these estimates come from 2PL models fit to the same 29 models, and the domain-specific and total models share both data and estimation procedure. Part of the correlation is mechanical: both theta estimates are functions of the same model outcomes. The reported p-values (p < 10⁻⁷) are also of limited meaning with N = 29. Thus the 'near-redundancy' result does not independently corroborate unidimensionality; it is another consequence of fitting the same model class to the same sample.
minor comments (5)
- [§2.6.2] The conversion λ_j = a_j / sqrt(1 + a_j²) is the normal-ogive loading formula. Equation (1) specifies a logistic 2PL model; the usual scaling constant D ≈ 1.702 should be introduced or the logistic/normal-ogive equivalence explicitly stated.
- [§2.6.2, PCA] The PCA domain R² = 0.035 is reported without a null distribution or permutation benchmark. With only 29 columns in the transposed matrix, the first three PCs may reflect noise; a permutation test would help establish whether 0.035 is smaller than expected under a no-domain-structure null.
- [§2.6.1 / references] The software is referenced as 'torch measure' but the URL shows 'torch_measure'; please standardize the name and ensure the reference is complete.
- [Figure 3] The figure is dense and hard to read with 29 model labels on the x-axis; consider rotating labels, using a table, or splitting into panels.
- [Abstract / §3.2] The abstract states 'McDonald's ω_h = 0.998' without qualification; if the corrected analysis uses a different estimator, the abstract should be updated to avoid perpetuating the mislabeling.
Circularity Check
The headline ω_h=0.998 is coefficient omega from the unidimensional 2PL model, used as if it tested dimensionality; that inference is partly self-referential, though PCA and domain-θ analyses give partial independent support.
specific steps
-
self definitional
[Section 2.6.2, Eq. (2); reported in Abstract and Section 3.2]
"Our primary evidence for unidimensionality is McDonald’s hierarchical omega (McDonald, 1999), computed analytically from the 2PL discrimination parameters. Under the normal-ogive parameterization, item j’s loading on the general factor is λj = aj/√(1+a_j²), and ωh is defined as: ωh = (Σj λj)² / [(Σj λj)² + Σj(1−λj²)] ... ωh quantifies the proportion of item variance attributable to the general factor, ranging from 0 (no general factor) to 1 (perfectly unidimensional)."
The 2PL model in Eq. (1) has exactly one latent dimension θ_i, so the 'general factor' loadings λ_j are defined by a unidimensional model. Eq. (2) is the coefficient-omega reliability formula for that one-factor composite: it measures how much variance the single fitted factor accounts for, not whether one factor suffices. The formula has no term for group factors, so by construction it cannot discriminate unidimensional from multidimensional structure. Reporting ωh=0.998 as 'primary evidence for unidimensionality' is therefore the model's unidimensionality assumption re-entering as the conclusion; the index is a function of the fitted loadings, but the dimensionality inference drawn from it is not an independent test.
full rationale
The paper does not depend on self-citation; the methodological references are standard external sources. The clearest circularity is confined to Eq. (2): the statistic labeled 'McDonald's ωh' is actually coefficient omega total reliability for a unidimensional congeneric model, so treating it as evidence of unidimensionality is self-referential. The residual-correlation analysis also reuses the same fitted λ_j to form R_g=λλ^T, but as a residual check it could in principle detect remaining group factors, so I do not count it as a separate circular step. The PCA on item response profiles is model-free, and the domain-θ correlations come from separate within-domain fits, giving the central claim some independent content. However, the paper's own 'primary evidence' is the self-referential ωh statistic, and the independent analyses are underpowered at N=29. This supports a moderate partial-circularity score of 4 rather than a higher score.
Axiom & Free-Parameter Ledger
free parameters (3)
- Item discrimination a_j (j=1..428) =
median 1.49, range 0.00-5.00 (upper cap imposed)
- Item difficulty b_j (j=1..428) =
median 0.41, range -2.06 to 5.67
- Latent model ability theta_i (i=1..29) =
e.g., Claude Opus 4.7 highest; GPT-4o-2024-11-20 near -1.8
axioms (5)
- domain assumption 2PL item response model with local independence and monotonicity is correct for HLE MCQ responses
- ad hoc to paper lambda_j = a_j / sqrt(1 + a_j^2) and Eq. 2 provide McDonald's hierarchical omega
- domain assumption N=29 models with >=95% coverage is sufficient for marginal maximum likelihood estimation of 428 two-parameter items
- ad hoc to paper Coding 0.79% missing responses as incorrect introduces negligible bias
- domain assumption The text-only multiple-choice subset represents HLE's eight subject domains
Cite this review
Pith. "Pith review of Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset." pith.science (2026). https://pith.science/paper/RJOLMVAP
@misc{pith2026260727420,
author = {Pith},
title = {Pith review of: Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset},
year = {2026},
howpublished = {\url{https://pith.science/paper/RJOLMVAP}},
note = {Machine review of arXiv:2607.27420}
}
read the original abstract
Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE ($J = 428$ items) and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find convergent evidence that HLE measures a single general reasoning factor: McDonald's $\omega_h = 0.998$, domain labels explain only 3.5\% of item response variance, within- and between-domain residual correlations are nearly identical (Cohen's $d = 0.016$), and domain-specific ability estimates are near-redundant with the total score ($r \geq 0.81$). A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above $\theta = 0$, where frontier models sit. These findings suggest that HLE's domain subscores do not warrant distinct capability interpretations and that the benchmark's ability to discriminate among the strongest models is limited.
Figures
Reference graph
Works this paper leans on
-
[1]
and Levin, John-Clark and Kazakov, Mstyslav and Feng, Fiona and Feng, Steven Y
Phan, Long and Gatti, Alice and Li, Nathaniel and Khoja, Adam and Kim, Ryan and Ren, Richard and Hausenloy, Jason and Zhang, Oliver and Mazeika, Mantas and Hendrycks, Dan and Han, Ziwen and Hu, Josephina and Zhang, Hugh and Zhang, Chen Bo Calvin and Shaaban, Mohamed and Ling, John and Shi, Sean and Choi, Michael and Agrawal, Anish and Chopra, Arnav and Na...
2026
-
[2]
arXiv preprint arXiv:2402.14992 , year =
Polo, Felipe Maia and Weber, Lucas and Choshen, Leshem and Sun, Yuekai and Xu, Gongjun and Yurochkin, Mikhail , title =. arXiv preprint arXiv:2402.14992 , year =
-
[3]
Chiang, Wei-Lin and Zheng, Lianmin and Sheng, Ying and Angelopoulos, Anastasios Nikolas and Li, Tianle and Li, Dacheng and Zhu, Banghua and Zhang, Hao and Jordan, Michael I. and Gonzalez, Joseph E. and Stoica, Ion , title =. arXiv preprint arXiv:2403.04132 , year =
-
[4]
Latent variables analysis: Applications for developmental research , editor=
Corrections to test statistics and standard errors in covariance structure analysis , author=. Latent variables analysis: Applications for developmental research , editor=. 1994 , publisher=
1994
-
[5]
, author=
Exploring the measurement invariance of psychological instruments: Applications in the substance use domain. , author=. 1997 , publisher=
1997
-
[6]
Structural equation modeling , volume=
Evaluating goodness-of-fit indexes for testing measurement invariance , author=. Structural equation modeling , volume=. 2002 , publisher=
2002
-
[7]
Structural equation modeling: a multidisciplinary journal , volume=
Sensitivity of goodness of fit indexes to lack of measurement invariance , author=. Structural equation modeling: a multidisciplinary journal , volume=. 2007 , publisher=
2007
-
[8]
2023 , eprint=
GPQA: A Graduate-Level Google-Proof Q&A Benchmark , author=. 2023 , eprint=
2023
-
[9]
arXiv preprint arXiv:2009.03300 , year =
Measuring Massive Multitask Language Understanding , author =. arXiv preprint arXiv:2009.03300 , year =. 2009.03300 , archivePrefix =
Pith/arXiv arXiv 2009
-
[10]
Transactions on Machine Learning Research , year =
Holistic Evaluation of Language Models , author =. Transactions on Machine Learning Research , year =. 2211.09110 , archivePrefix =
-
[11]
Comparing Test Sets with Item Response Theory
Vania, Clara and Htut, Phu Mon and Huang, William and Mungra, Dhara and Pang, Richard Yuanzhe and Phang, Jason and Liu, Haokun and Cho, Kyunghyun and Bowman, Samuel R. Comparing Test Sets with Item Response Theory. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural...
-
[12]
2023 , eprint=
Holistic Evaluation of Language Models , author=. 2023 , eprint=
2023
-
[13]
2022 , eprint=
On the Opportunities and Risks of Foundation Models , author=. 2022 , eprint=
2022
-
[14]
and Koyejo, Sanmi , title =
Truong, Sang T. and Koyejo, Sanmi , title =. 2026 , publisher =
2026
-
[15]
2023 , eprint=
Efficient Memory Management for Large Language Model Serving with PagedAttention , author=. 2023 , eprint=
2023
-
[16]
and Yi, Xiaoyuan and Koyejo, Sanmi and Xie, Xing and Xiao, Ziang , title =
Jiang, Han and Zhang, Susu and Zhu, Dongyao and Bai, Yuzhuo and Truong, Sang T. and Yi, Xiaoyuan and Koyejo, Sanmi and Xie, Xing and Xiao, Ziang , title =. arXiv preprint arXiv:2604.03244 , year =
-
[17]
and Curran, Patrick J
Flora, David B. and Curran, Patrick J. , title =. Psychological Methods , year =
-
[18]
, title =
McDonald, Roderick P. , title =. 1999 , doi =
1999
-
[19]
and others , title =
Truong, Sang T. and others , title =. 2026 , url =
2026
-
[20]
and Kim, Seock-Ho , title =
Baker, Frank B. and Kim, Seock-Ho , title =. 2004 , doi =
2004
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.