Pith. sign in

REVIEW 3 major objections 9 minor 42 references

LLM cognitive-task scores show a small theory-aligned prompt effect, but not stable five-dimensional ability profiles.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Across 55 models and a frozen crossed scaffold study, LLM cognitive-task scores show a dominant general factor and only a small, non-transportable grouping tendency—not stable five-dimensional profiles.

T0 review reviewed 2026-07-31 challenge →

load-bearing objection Careful negative result plus a reusable multimethod checklist: five-dimensional LLM cognitive profiles are not established on this battery, and the one positive Γ claim is thinner than the abstract suggests. the 3 major comments →

arxiv 2607.24999 v1 pith:5H6F5FEC submitted 2026-07-27 cs.CL cs.AI

CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

classification cs.CL cs.AI
keywords large language modelscognitive evaluationconstruct validityability structurepsychometric validationscaffold interventionsout-of-family predictionprocedural benchmarks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

People increasingly turn lists of cognitive-test scores from language models into per-ability profiles—working memory, control, episodic memory, theory of mind, metacognition—as if those labels name separable dimensions. This paper asks when such labels are earned. It builds CogArena, a procedurally generated 13-paradigm battery and a four-part test: do tasks show the expected behavioral signatures, do scores cluster by grouping beyond a common competence axis, do matched answer-free scaffolds raise the right groups selectively, and does any of that improve prediction for unseen model families? Across dozens of open-weight models a single broad axis explains about half the variance; the within-grouping edge is small and sensitive to scoring and family; scaffolds show only a weak battery-level diagonal tendency; and the pre-set confirmation rule, including transport, fails—including under alternate wording. The practical upshot is a stricter workflow before cognitive labels are attached to model scores, and a boundary result: the five groupings remain organizing labels, not validated transportable dimensions.

Core claim

Theory-aligned prompting produces a small in-battery matched-grouping tendency, but the present evidence does not establish stable five-dimensional cognitive profiles. Nearly all paradigm correlations are positive and one common axis explains roughly half the variance; the within-grouping covariance advantage is small, scoring-sensitive, and uncertain across families; no scaffold-specific contrast survives multiplicity correction; and neither observational grouping scores nor intervention selectivity improve held-out-family prediction. The frozen all-nine confirmation criterion fails, as does a post-hoc alternate-wording replication.

What carries the argument

CogArena’s multimethod validation workflow: the same five theory-motivated groupings are tested jointly through within-paradigm behavioral signatures, between-model covariance (convergent/discriminant structure), a fully crossed matched-scaffold intervention against a length-matched neutral placebo, and prediction to held-out model families—before dimensional cognitive labels are attached.

Load-bearing premise

That five theory groupings, each built from only two or three text-adapted tasks (some weakly signed or only partly faithful to the original human procedure), are a fair enough stand-in for the abilities being judged.

What would settle it

Rerun the frozen fully crossed scaffold study so that the matched-grouping advantage stays positive under family-by-item uncertainty, scaffold-to-group mapping and family consistency clear their gates, and selective terms actually improve leave-one-family-out prediction—passing all nine pre-set confirmation gates rather than failing on transport and cell-minimum rules.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Per-ability LLM “cognitive profiles” should not be treated as validated latent traits on the strength of task coverage or positive correlations alone.
  • Reporting should prefer paradigms with replicated behavioral signatures over unvalidated grouping means.
  • Claims of separable cognitive structure need family-aware covariance, matched intervention selectivity, and out-of-family prediction—not only in-battery accuracy gains.
  • Theory-aligned scaffolds can show a small diagonal tendency without proving transportable dimensions.
  • Future batteries can reuse the same four-level workflow before attaching dimensional labels to model scores.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Benchmark leaders that publish spider charts of named abilities without signature, selectivity, and transport checks are making a stronger scientific claim than their designs support.
  • If broad competence dominates text batteries, multi-ability “cognitive radar” plots may mostly restate overall capability under different task names.
  • Thickening each grouping (more paradigms, harder items, better-adapted modalities) is a direct next experiment that could flip the boundary without changing the workflow.
  • The same confirmation stack could be applied to other fashionable LLM taxonomies—personality, values, agency—before those labels are treated as stable dimensions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 9 minor

Summary. The paper introduces CogArena, a 13-paradigm, procedurally generated benchmark adapting established cognitive-science tasks into five theory-motivated groupings (working memory, cognitive control, episodic memory, theory of mind, metacognition), evaluated on 55 open-weight LLMs with a 12-model fully crossed scaffold-intervention study. Its central contribution is a multimethod validation protocol — within-paradigm behavioral signatures, between-model covariance, crossed matched-scaffold interventions against a length-matched placebo, and leave-one-family-out prediction — gated by a frozen nine-criterion confirmation rule. The results are largely negative/boundary: PC1 explains ~49.8% of variance, the within-grouping correlation advantage is small (δ≈.081), family-clustered intervals include zero, construct-native rescoring reverses the contrast, no scaffold-specific contrast survives BH correction, transport to held-out families fails, and the all-nine rule fails under both the frozen and an alternate scaffold wording. The one positive clause is a small matched-scaffold diagonal tendency (Γ≈.0199) relative to placebo. The manuscript is unusually transparent about its own sensitivity analyses, limitations, and post-hoc audits, and releases code, generators, frozen specifications, and run manifests.

Significance. If the results stand, this is a useful corrective to a growing practice of reporting per-ability LLM profiles without testing separability. The paper ships several things the field needs more of: procedurally generated items with a direct contamination probe, deterministic LLM-judge-free scoring with scorer-specification sensitivities, a construct-native rescoring battery (d′, interference differences, Brier, type-2 d′) with split-half reliabilities, simulation calibration of the separability test's power and type-I rates, an exact 5! mapping test, family-clustered and crossed family×item bootstrap inference, and an outcome-frozen nine-gate decision rule whose failure is reported plainly. The boundary conclusion is credible and the workflow is reusable even where confirmation fails. The thin operationalization (2–3 paradigms per grouping, one modality) limits generality but is explicitly scoped in §6 rather than hidden. The one overreach risk is the small positive Γ clause, detailed in the major comments.

major comments (3)
  1. [§5.3 and Appendix S1.11 (post-hoc audit, item 3); Abstract] The headline Γ=.0199 [CI .0041,.0360] is computed under intention-to-treat scoring (invalid = 0), while the placebo controls only prompt presence and approximate length (§6, 'Scope of the Intervention Evidence') and the five scaffolds explicitly 'specify how to organize a response' (§4). The authors' own audit shows this matters: restricted to protocol-valid, nonempty, parseable responses, Γ falls to .0072 [CI −.0067,.0246], and .0062 of the .0199 comes from target-only-evaluable pairs — cells where the scaffolded response was parseable and the placebo response was not. This is exactly the signature of a differential format-compliance channel rather than theory-aligned facilitation. The main text and abstract should (a) report the evaluability-restricted estimate alongside the headline Γ, (b) soften the causal reading of the abstract clause, and (c) discuss the needed control: a format-m
  2. [Table S9 (frozen all-required confirmation rule, gates 7–8)] Two of the three failed gates (empty-response exclusion, 855 pairs; operation-span parse exclusion, 262 pairs) are scored FAIL only because the exclusions leave cells below the frozen minimum and are therefore 'unestimable' — not because the ≥.5Γ preservation criterion was evaluated and failed. Given that M1 shows asymmetric evaluability carries a substantial share of Γ, these two exclusions are precisely the informative ones. The paper should report, even if post-hoc, the Γ estimate and preservation ratio over the cells that remain estimable under each exclusion, so readers can tell whether the gates would have failed on the effect-size criterion rather than on cell-count bookkeeping. As written, the all-nine failure is partly a consequence of the frozen minimum-cell rule rather than of the selectivity evidence itself, and the manuscript does not separate these.
  3. [§5.2 ('Where Grouping Structure Strengthens') and Appendix S1.5 (Table S4)] The joint exclusion of text Stroop, Go/No-Go, and CVLT yields accuracy δ=.147 (p2=.021), and the difficulty-tier analysis gives δ up to .169 with merged-family intervals excluding zero at every tier. These are the strongest positive separation results in the paper and receive prominent main-text placement, yet the manuscript itself notes that no family-clustered interval was computed for the joint deletion — i.e., the analysis is not held to the family-aware inferential standard the paper applies everywhere else (and under which the primary δ does not survive). Either compute and report the family-clustered interval for the joint-deletion analysis, or move these restricted-view analyses to the appendix as clearly-labeled sensitivities; the current framing invites readers to weight them more heavily than the paper's own protocol licenses.
minor comments (9)
  1. [Figure 4 caption] The caption references 'grouping abbreviations from Table 1,' but Table 1 contains no grouping abbreviations; presumably Table S10/S13 or Figure 1 is meant.
  2. [Table S14] The caption says 'Bold marks a column maximum,' but no boldface appears in the table as printed. Also, Qwen2.5-7B false belief is listed as 100% here while §S1.3 cites the same comparison as 100% vs. Mistral 68% — consistent, but the Stroop text mean (92%) should be reconciled with the §5.1 aggregate congruent/incongruent figures (94.2%/89.4%).
  3. [Table 1] The column header 'p2' (two-sided p) is never defined in the caption or text; §S1.5 reports one-sided p-values (.037 for the full matrix) against the main text's two-sided .057 without flagging the convention change. Please define p2 and state the sidedness of each reported p.
  4. [Table 2] The column 'Families+' (5/6, 4/6) is undefined; readers must infer it counts positive family-level Γ estimates. Define in the caption and cross-reference the exact sign-flip test (p=.063) that gives the inferential version.
  5. [Abstract / §5.2] Please state in the abstract that 77 of 78 correlations are positive; 'nearly all' undersells a striking descriptive result.
  6. [§3.1 (Adaptation Distance)] The adaptation-distance ratings are acknowledged as author judgments; a short documented rubric (what would move a paradigm from Low to Medium) or a second rater would strengthen the audit, since the ratings gate which paradigms enter the battery.
  7. [Appendix S1.11 (Freeze and reporting amendment) / code release] Raw response text is withheld from the repository (described as 'anonymous,' though the submission is not anonymized). The evaluability decomposition in S1.11 — central to interpreting Γ — cannot be independently verified without response-level parseability labels. Please release the stored responses or, at minimum, the per-record evaluability flags with the SHA-256 manifests.
  8. [References] Several 2025–2026 citations appear to be preprints or workshop papers (e.g., He et al. 2026; Javadov et al. 2026; Contreras 2026; Bugaud 2026); please update to published versions where they exist.
  9. [§5.3 (ICC diagnostic) and S1.12] The ICC=.979 replay-stability diagnostic is a useful control, but it is computed on adjacent greedy-decoding administrations of identical items; please note explicitly that it bounds serving/replay variance only and says nothing about seed sensitivity of the procedurally generated item sets, which is the more relevant variance source for the transport analyses.

Circularity Check

0 steps flagged

No significant circularity: boundary claims are tested against external criteria (permutations, family bootstrap, held-out transport, frozen gates), not defined into existence from fitted inputs.

full rationale

CogArena’s load-bearing chain is empirical and largely negative. The five groupings are imported from external cognitive taxonomies (CHC, Miyake EF, false-belief and metacognition traditions), not estimated from the same model-by-paradigm matrix and then re-declared validated. Separability is assessed with label-permutation tests, family-clustered intervals, construct-native rescoring, and leave-one-family-out RMSE/ΔLL; matched-scaffold selectivity Γ is defined as placebo-adjusted matched-minus-nonmatched gain, with an exact 5! mapping test and a pre-frozen nine-gate confirmation rule that the authors report as failing. The small positive in-battery Γ and the dominant PC1/positive manifold are measured quantities, not parameters fitted then renamed as predictions. Self-citation (e.g., WMF-AM) is peripheral background, not a uniqueness theorem or ansatz that forces the boundary conclusion. Concerns about thin operationalization (2–3 paradigms per grouping) or format-compliance channels in Γ are scope/validity issues, not circular reductions of outputs to inputs by construction.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 3 invented entities

The load-bearing content is empirical measurement design, not a formal derivation. Claims rest on standard psychometrics plus domain choices about which human paradigms and five groupings operationalize ‘cognitive ability structure’ in text LLMs, and on statistical decision rules frozen for the intervention study.

free parameters (4)
  • Frozen confirmation gate thresholds (nine-gate rule) = all-nine must pass; several numeric floors in Table S9
    Pass/fail cutoffs such as crossed CI lower bound >0, ≥4/6 families positive, mapping p≤.05, protocol-invalid ≤1%, exclusion preserving ≥0.5Γ and ≥3 items/cell are author-chosen decision parameters that define ‘confirmation’.
  • Primary within-grouping separation estimand and 0.15 profiling threshold = 0.15 increment / observed primary δ≈0.081
    Interpretation of δ and simulation power arms uses a pre-specified 0.15 within-grouping correlation increment as a substantive profiling threshold; not fitted to confirm the taxonomy but chosen as a bar.
  • Scaffold wording and length-matched placebo content = Γ_frozen=.0199; Γ_alternate=.0134
    Answer-free scaffold texts and placebo length matching are design choices that quantitatively determine Γ; alternate wording changes Γ from .0199 to .0134.
  • Adaptation-distance ratings (Low/Medium) and paradigm inclusion set = 13 paradigms; Low/Medium only
    Author judgments exclude High-distance paradigms and retain Medium ones (Stroop, Flanker, Go/No-Go, wagering), shaping which constructs enter the taxonomy test.
axioms (6)
  • standard math Standard multivariate statistics apply to the model-by-paradigm accuracy matrix (Pearson correlations, PCA, label permutation, family bootstrap, BH correction).
    Used throughout §5.2–5.3 and Appendices S1.5–S1.7 without unusual estimators beyond disclosed sensitivities.
  • domain assumption Human cognitive taxonomies (CHC-inspired WM/control/episodic groupings; false-belief and metacognitive-monitoring traditions; Miyake unity-diversity EF) are appropriate theory sources for labeling five LLM score groups.
    Stated in §2–3; the paper tests rather than assumes empirical separability, but still assumes these labels are the right candidate dimensions.
  • domain assumption Text-adapted paradigms with behavioral-signature checks can still support profile-level dimensional inference even when some signatures are weak or Medium-distance.
    §3.1 and §5.1 retain mixed-signature paradigms in the battery while cautioning interpretation; profile tests pool them.
  • domain assumption Open-weight checkpoints served at default quantization with greedy decoding yield scores informative about family-general cognitive structure claims.
    §4 and Limitations: no closed/frontier models; quantization mixed with full-precision external benchmarks in exploratory correlations.
  • ad hoc to paper Selectivity defined as placebo-adjusted matched minus nonmatched gain, averaged equally across five scaffolds (Γ), is the right interventional estimand for dimensional labels.
    §4 Intervention-Validity Study and Appendix S1.11 define S_j and Γ; confirmation requires this diagonal plus transport.
  • ad hoc to paper Outcome-frozen protocol without public preregistration is sufficient to treat the nine-gate decision as confirmatory for the intervention claim.
    Authors explicitly note non-preregistration while freezing panel, items, prompts, and gates before formal outcome inspection (§4, S1.11).
invented entities (3)
  • CogArena multimethod confirmation framework (signatures + covariance + crossed scaffolds + held-out-family transport) independent evidence
    purpose: Decide when cognitive-task scores warrant dimensional labels rather than only organizing labels.
    Core methodological contribution; criteria are stipulated by the authors then applied to yield a fail decision on five-dimensional profiles.
  • Five answer-free theory-targeted scaffolds (ledger, rule rehearsal, source binding, belief-state ledger, metacognitive forecast) no independent evidence
    purpose: Provide matched interventions for discriminant validity without supplying item answers.
    Prompt instruments invented for this study; selectivity is measured against placebo and full cross.
  • Construct-native rescoring battery (interference differences, d′, false-memory contrast, Brier, type-2 d′) independent evidence
    purpose: Separate shared response-format variance from intended construct variance in separability tests.
    Alternative endpoints defined in S1.6; they change δ sign and PC1 share, supporting scoring-sensitivity of structure.

reviewed 2026-07-31 · how reviews work

0 comments
Cite this review

Pith. "Pith review of CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models." pith.science (2026). https://pith.science/paper/5H6F5FEC

@misc{pith2026260724999,
  author       = {Pith},
  title        = {Pith review of: CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5H6F5FEC}},
  note         = {Machine review of arXiv:2607.24999}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.

Figures

Figures reproduced from arXiv: 2607.24999 by Dengzhe Hou, Fangzhou Lin, Kazunori D Yamada, Lingyu Jiang.

Figure 1
Figure 1. Figure 1: CogArena overview. Two of 13 paradigms illustrate how established cognitive procedures become procedurally [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Paradigm-level construct diagnostics. Bars show corrected mean accuracy, except that DRM shows false-recognition [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Pearson correlations among corrected paradigm [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Intervention-validity evidence. (A) Target-minus [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

42 extracted references · 5 linked inside Pith

  1. [1]

    M.; and Frith, U

    Baron-Cohen, S.; Leslie, A. M.; and Frith, U. 1985. Does the autistic child have a ``theory of mind''? Cognition, 21(1): 37--46

  2. [2]

    M.; Kearns, R

    Bean, A. M.; Kearns, R. O.; Romanou, A.; et al. 2025. Measuring what Matters: Construct Validity in Large Language Model Benchmarks. In Advances in Neural Information Processing Systems 38 (NeurIPS), Datasets and Benchmarks Track

  3. [3]

    Binz, M.; Akata, E.; Bethge, M.; Br \"a ndle, F.; Callaway, F.; Coda-Forno, J.; et al. 2025. A foundation model to predict and capture human cognition. Nature, 644(8078): 1002--1009

  4. [4]

    Binz, M.; and Schulz, E. 2023. Using cognitive psychology to understand GPT-3 . Proceedings of the National Academy of Sciences, 120(6): e2218523120

  5. [5]

    Bugaud, Z. 2026. A Cognitive Battery for Foundation Models: Theory-Grounded Benchmarks for Attention, Learning, Metacognition, Executive Function, and Social Cognition. In ICML 2026 Workshop on Combining Theory and Benchmarks

  6. [6]

    Burnell, R.; Hao, H.; Conway, A. R. A.; and Hern \'a ndez-Orallo, J. 2023. Revealing the structure of language model capabilities. arXiv preprint arXiv:2306.10062

  7. [7]

    X.; and Schulz, E

    Coda-Forno, J.; Binz, M.; Wang, J. X.; and Schulz, E. 2024. CogBench : A large language model walks into a psychology lab. In Proceedings of the 41st International Conference on Machine Learning (ICML)

  8. [8]

    Contreras, J. M. 2026. An LLM -Native Psychometric Instrument Does Not Predict LLM Behavior: Evidence Across 25 Models. arXiv preprint arXiv:2606.09843

  9. [9]

    I.; Hu, B.; Le, K

    de Langis, K.; Park, J. I.; Hu, B.; Le, K. C.; Schramm, A.; Mensink, M. C.; Elfenbein, A.; and Kang, D. 2026. Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs . In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 5971--5986

  10. [10]

    C.; Kramer, J

    Delis, D. C.; Kramer, J. H.; Kaplan, E.; and Ober, B. A. 2000. California Verbal Learning Test--Second Edition (CVLT-II) : Adult Version Manual

  11. [11]

    A.; and Eriksen, C

    Eriksen, B. A.; and Eriksen, C. W. 1974. Effects of noise letters upon the identification of a target letter in a nonsearch task. Perception & Psychophysics, 16(1): 143--149

  12. [12]

    Fischhoff, B.; Slovic, P.; and Lichtenstein, S. 1977. Knowing with certainty: The appropriateness of extreme confidence. Journal of Experimental Psychology: Human Perception and Performance, 3(4): 552--564

  13. [13]

    G.; Ardi, F

    Haznitrama, F. G.; Ardi, F. R.; and Oh, A. 2026. A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities. arXiv preprint arXiv:2603.02540

  14. [14]

    He, J.; Dai, S.; Qiao, X.; Li, J.; Yan, Y.; and Hu, X. 2026. Beyond Direct Gains: Matched Controls for Evaluating Concept Scaffolds. In ICML 2026 AI4Math Workshop

  15. [15]

    Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021. Measuring massive multitask language understanding. In International Conference on Learning Representations (ICLR)

  16. [16]

    Hou, D.; Jiang, L.; Li, D.; Li, Z.; Lin, F.; and Yamada, K. D. 2026. WMF-AM : Probing LLM Working Memory via Depth-Parameterized Cumulative State Tracking. arXiv preprint arXiv:2603.27343

  17. [17]

    Ili \'c , D.; and Gignac, G. E. 2024. Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement? Intelligence, 106: 101858

  18. [18]

    Javadov, A.; Aitkazinov, S.; Hoesli, T.; von Wangenheim, F.; Schuller, B.; and Ollier, J. 2026. NeuReasoner : Theory-Grounded Mapping of Reasoning Elicitation Boundaries. arXiv preprint arXiv:2606.29971

  19. [19]

    K.; Hashtroudi, S.; and Lindsay, D

    Johnson, M. K.; Hashtroudi, S.; and Lindsay, D. S. 1993. Source monitoring. Psychological Bulletin, 114(1): 3--28

  20. [20]

    R.; Trott, S.; and Bergen, B

    Jones, C. R.; Trott, S.; and Bergen, B. 2024. Comparing Humans and Large Language Models on an Experimental Protocol Inventory for Theory of Mind Evaluation ( EPITOME ). Transactions of the Association for Computational Linguistics, 12: 803--819

  21. [21]

    Jung, J.; Lutz, M.; Sen, I.; and Strohmaier, M. 2026. Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 8143--8173. Association for Computational Linguistics

  22. [22]

    M.; and Schulz, E

    Kipnis, A.; Voudouris, K.; Schulze Buschoff, L. M.; and Schulz, E. 2025. metabench : A Sparse Benchmark of Reasoning and Knowledge in Large Language Models. In International Conference on Learning Representations (ICLR)

  23. [23]

    Lichtenstein, S.; and Fischhoff, B. 1977. Do those who know more also know more about how much they know? Organizational Behavior and Human Performance, 20(2): 159--183

  24. [24]

    MacLeod, C. M. 1991. Half a century of research on the S troop effect: An integrative review. Psychological Bulletin, 109(2): 163--203

  25. [25]

    McGrew, K. S. 2009. CHC theory and the human cognitive abilities project: Standing on the shoulders of the giants of psychometric intelligence research. Intelligence, 37(1): 1--10

  26. [26]

    Mirzadeh, I.; Alizadeh, K.; Shahrokhi, H.; Tuzel, O.; Bengio, S.; and Farajtabar, M. 2025. GSM -Symbolic: Understanding the limitations of mathematical reasoning in large language models. In International Conference on Learning Representations

  27. [27]

    P.; Emerson, M

    Miyake, A.; Friedman, N. P.; Emerson, M. J.; Witzki, A. H.; Howerter, A.; and Wager, T. D. 2000. The unity and diversity of executive functions and their contributions to complex ``frontal lobe'' tasks: A latent variable analysis. Cognitive Psychology, 41(1): 49--100

  28. [28]

    Moment \`e , F.; Suglia, A.; Giulianelli, M.; Ferrari, A.; Koller, A.; Lemon, O.; Schlangen, D.; Fern \'a ndez, R.; and Bernardi, R. 2025. Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests. In Findings of the Association for Computational Linguistics: EMNLP 2025, 20051--20072. Association for Computational Linguistics

  29. [29]

    T.; Garc \'i a-Madruga, J

    Pelegrina, S.; Lechuga, M. T.; Garc \'i a-Madruga, J. A.; Elos \'u a, M. R.; Macizo, P.; Carreiras, M.; Fuentes, L. J.; and Bajo, M. T. 2015. Normative Data on the N -Back Task for Children and Young Adolescents. Frontiers in Psychology, 6: 1544

  30. [30]

    Persaud, N.; McLeod, P.; and Cowey, A. 2007. Post-decision wagering objectively measures awareness. Nature Neuroscience, 10(2): 257--261

  31. [31]

    S.; Broadway, J

    Redick, T. S.; Broadway, J. M.; Meier, M. E.; Kuriakose, P. S.; Unsworth, N.; Kane, M. J.; and Engle, R. W. 2012. Measuring working memory capacity with automated complex span tasks. European Journal of Psychological Assessment, 28(3): 164--171

  32. [32]

    L.; and McDermott, K

    Roediger, H. L.; and McDermott, K. B. 1995. Creating false memories: Remembering words not presented in lists. Journal of Experimental Psychology: Learning, Memory, and Cognition, 21(4): 803--814

  33. [33]

    Serapio-Garc \' a, G.; Safdari, M.; Crepy, C.; Sun, L.; Fitz, S.; Romero, P.; Abdulhai, M.; Faust, A.; and Matari \'c , M. 2025. A psychometric framework for evaluating and shaping personality traits in large language models. Nature Machine Intelligence, 7(12): 1954--1968

  34. [34]

    Strachan, J. W. A.; Albergo, D.; Borghini, G.; Pansardi, O.; Scaliti, E.; Gupta, S.; Saxena, K.; Rufo, A.; Panzeri, S.; Manzi, G.; Graziano, M. S. A.; and Becchio, C. 2024. Testing theory of mind in large language models and humans. Nature Human Behaviour, 8: 1285--1295

  35. [35]

    Stroop, J. R. 1935. Studies of interference in serial verbal reactions. Journal of Experimental Psychology, 18(6): 643--662

  36. [36]

    D.; and Jones, C

    Trott, S.; Rivi \`e re, P. D.; and Jones, C. R. 2026. Do Different Theory of Mind Tasks for LLMs Measure the Same Thing? In ACL 2026 Workshop on Evaluating Evaluations (EvalEval)

  37. [37]

    Van der Elst, W.; Van Boxtel, M. P. J.; Van Breukelen, G. J. P.; and Jolles, J. 2006. The S troop Color-Word Test: Influence of age, sex, and education; and normative data for a large sample across the adult age range. Assessment, 13(1): 62--79

  38. [38]

    L.; and Langenecker, S

    Votruba, K. L.; and Langenecker, S. A. 2013. Factor structure, construct validity, and age- and education-based normative data for the P arametric Go/No-Go T est. Journal of Clinical and Experimental Neuropsychology, 35(2): 132--146

  39. [39]

    Wechsler, D. 2008. WAIS-IV Administration and Scoring Manual

  40. [40]

    M.; Cross, D.; and Watson, J

    Wellman, H. M.; Cross, D.; and Watson, J. 2001. Meta-analysis of theory-of-mind development: The truth about false belief. Child Development, 72(3): 655--684

  41. [41]

    Yang, Y.; Miao, C.; Li, W.; and Wu, Y. 2026. ActTraitBench : Quantifying the Knowledge--Decision Gap in Large Language Models via Human-Grounded Behavioral Validation. arXiv preprint arXiv:2605.29791

  42. [42]

    M.; et al

    Zhou, L.; Pacchiardi, L.; Mart \' nez-Plumed, F.; Collins, K. M.; et al. 2026. General scales unlock AI evaluation with explanatory and predictive power. Nature, 652: 58--67

This paper was first reviewed by grok-4.5 on July 31, 2026.