Pith. sign in

REVIEW 4 major objections 5 minor 4 cited by

This paper argues that the jagged intelligence of AI systems stems from a missing training signal—cognitive dark matter—and that collecting process-level data from human minds and brains can supply that signal.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 02:32 UTC pith:UZWF4JYM

load-bearing objection A wide-framing hypothesis paper that names a real pattern and a data agenda; the empirical support is thin, but the idea is worth taking seriously. the 4 major comments →

arxiv 2603.03414 v2 pith:UZWF4JYM submitted 2026-03-03 q-bio.NC

Cognitive Dark Matter: Measuring What AI Misses

classification q-bio.NC
keywords cognitive dark matterjagged intelligencemetacognitioncognitive flexibilityneural-behavioral dataprocess tracingAI benchmarkscognitive process training
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper proposes that the uneven, 'jagged' intelligence of modern AI—adept at some tasks, surprisingly poor at others—is not a fixed ceiling but a measurement gap. The missing ingredient is cognitive dark matter: brain functions such as metacognition, cognitive flexibility, social reasoning, and emotional intelligence that shape behavior yet are hard to infer from behavior alone. The paper surveys current AI benchmarks and large-scale neuroscience datasets and finds both are skewed toward capabilities AI has already mastered, while these hard-to-measure functions are largely absent. To close the gap, it proposes collecting three kinds of data—latent variables from cognitive models, process-tracing data such as eye-tracking and think-aloud protocols, and paired neural-behavioral data—so models can be trained on cognitive process rather than behavioral outcome alone. If this is right, the path to more general, less jagged AI runs through measuring the mind, with better understanding of human cognition as a dual payoff.

Core claim

On the paper's own terms, the central claim is that the jagged intelligence landscape of AI systems arises from a missing training signal, not from lack of scale or raw capability. CDM-loaded functions are largely unmeasured in current AI benchmarks and in large-scale neuroscience datasets, so models never receive the signal needed to acquire them. The paper argues that making hidden cognitive processes visible—through latent variables, process-tracing data, and paired neural-behavioral data—would let AI train on process rather than outcome, producing models that generalize more smoothly and fail in human-legible ways. The authors present the measurement gap as both an opportunity and a warn

What carries the argument

The central object is cognitive dark matter (CDM), defined as brain functions that meaningfully shape behavior yet are hard to infer from behavior alone. The argument runs through a three-tier taxonomy used in the paper (L1: mastered capabilities like vision and language; L2: partial progress; L3: rarely explored CDM-loaded functions such as cognitive flexibility, social reasoning, and metacognition). The proposed mechanism is a trio of data types: latent variables sampled from large-scale cognitive models, process-tracing data (eye-tracking, mouse-tracking, think-aloud), and paired neural-behavioral data. These are meant to supply the missing training signal by exposing the causal cognitive

Load-bearing premise

The load-bearing premise is that hidden cognitive processes—metacognitive states, attention, neural activity—when added to training data are causally efficacious signals that transfer human-like capabilities to models; if they are only correlates of behavior, training on them may not close the jaggedness gap.

What would settle it

A controlled comparison: train matched models on the same CDM-heavy tasks with and without process-tracing or neural supervision, then test on novel L3 tasks such as rule-switch flexibility and confidence calibration. If process-supervised models show no improvement over behavior-only baselines at comparable data scale, the thesis that the missing signal is latent process data is falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Benchmarking practice would shift: frontier model evaluation suites should include L3-heavy tests, because current suites systematically under-test the functions where jaggedness hides.
  • The bottleneck for less jagged AI becomes data collection, not architecture: large, intensive datasets for metacognition, flexibility, social reasoning, and related functions become a prerequisite.
  • Training on process data should make failures more human-legible—detectable and interpretable—because models would learn the cognitive processes that generate human outputs.
  • Process supervision alone is insufficient for most tasks because correct intermediate steps are unknown; the paper's three data types are meant to supply those latent processes at scale.
  • A dual benefit follows: even if AI transfer falls short, the datasets would be foundational for understanding human cognition.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • An implicit prediction is that returns from process-level data should show up specifically on L3 capabilities; a scaling-law-style curve for eye-tracking or neural data against flexibility and calibration metrics would test this.
  • The chess-endgame case suggests a concrete diagnostic: measuring whether a model abandons a failing strategy after feedback could serve as a quantitative metacognition score before and after CDM-style training.
  • The thesis connects jaggedness to safety: chronic miscalibration, hallucination, and perseveration may share a single missing metacognitive signal, so better measurement could improve reliability as well as capability.
  • If the latent signals turn out to be epiphenomenal, the paper's own dual-benefit framing still justifies the dataset program, but the AI-transfer claim would need a different mechanism.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes that the jagged capability landscape of modern AI systems is caused by a missing training signal, termed 'cognitive dark matter' (CDM): cognitive functions that shape human behavior but are hard to infer from observable behavior alone (metacognition, cognitive flexibility, lifelong learning, abductive reasoning, social/emotional reasoning). It argues that current AI benchmarks and large-scale neuroscience datasets both under-represent these L3 functions, and it proposes a research program based on three data types—latent variables from cognitive models, process-tracing data, and paired neural-behavioral data—to train AI on cognitive process rather than outcome. Three small empirical analyses are presented: a chess-app example, a survey of benchmark tiers in model release documents, and a survey of neuroimaging/literature coverage. The central claim is that better measurement of CDM will produce more general, less jagged AI, with the same data also advancing cognitive neuroscience.

Significance. If the transfer premise held, the paper would provide an actionable and fairly original research program that connects AI benchmarking, cognitive modeling, and large-scale neuroscience. The dual-benefit framing is attractive, and the authors are explicit that scaling laws for neural-data training are currently unknown. The paper also makes a falsifiable prediction: collecting the proposed data types should improve performance on CDM-loaded tasks beyond what outcome-only data of the same scale would provide. However, the empirical support is exploratory: the three analyses are small, partially manual, and not accompanied by inter-rater reliability or confidence intervals, and the central causal claim rests on an untested transfer assumption. The paper's strength is as a synthesis and proposal, not as a demonstration.

major comments (4)
  1. [CDM, Jagged Intelligence, and human–AI interactions; 'What to collect next'] The load-bearing claim is that process-tracing data, neural recordings, and cognitive-model latents can serve as training signals that instill metacognition, flexibility, and social reasoning into AI. No pilot, derivation, or mechanism is provided. The cited brain-tuning studies (refs 28–34) concern perceptual/semantic robustness, not L3 functions; process supervision (ref 18) rewards correct intermediate steps, which is not the same as latent cognitive process. The paper itself admits the absence of scaling laws. This transfer premise needs at least a proof-of-concept or a concrete falsifiable test, e.g., comparing CDM-enriched versus outcome-only supervision on the same L3 benchmarks with matched compute.
  2. [Methods: Chess endgame evaluation] The chess example is presented as evidence of missing metacognition, but it is based on N=3 models × 10 runs, with manual scoring and no inter-rater reliability, confidence intervals, or quantified failure rates. The interactive Claude Code sessions are described qualitatively as 'failed to elicit it.' As an illustrative anecdote this is acceptable, but as evidence for the core claim it is under-powered. Report the exact counts, scoring rubric, and inter-rater agreement, or explicitly label it as anecdotal.
  3. [Methods: AI benchmark analysis and neuroimaging surveys; Figures 2–3] The categorization of benchmarks and neuroimaging studies into the Liu et al. L1/L2/L3 taxonomy is subjective and mostly manual, with LLM-assisted tagging for some datasets. The 50 GPT-5.2 runs for benchmark classification are reported as 'stability' but no agreement statistic is given. Without inter-rater reliability or a validation of the LLM labels, the quantitative 'gap' shown in Figures 2 and 3 is suggestive but not established. The qualitative conclusion may survive, but the numerical presentation overstates precision.
  4. [CDM, Jagged Intelligence, and human–AI interactions] The paper asserts that under-measurement *causes* jagged performance, but it does not rule out alternative explanations: L3 tasks may be harder because they require more data in general, different architectures, or better world models, independent of whether the missing signal is 'cognitive process.' The constant-hazard-rate analysis of Ord (ref 9) suggests subtask success rates p and effective step count n; improving p could occur through many mechanisms, not specifically CDM-enriched data. The central hypothesis should be stated with a concrete discriminating experiment, otherwise it risks being unfalsifiable.
minor comments (5)
  1. [Abstract and body text] Typographical errors: 'isjagged', 'musthave', and missing spaces before citations in a few places. Figure 3 labels 'Claude 4.5' while the text and Methods say 'Claude Opus 4.5'.
  2. [Figure 2 caption] The figure caption does not define what counts as 'intensive neuroimaging datasets' or report the number of papers per source, making the y-axis scale hard to interpret. Please clarify the counts and sources.
  3. [Methods: AI benchmark analysis] The Methods say 'GPT-5.2 with high reasoning effort' was used to assign cognitive functions, but no details are given on the prompt, temperature, or how the 50 runs were aggregated. If the analysis is to be reproducible, these details are needed.
  4. [What to collect next] The '~500 hours' milestone for large-scale datasets is stated without a citation or quantitative justification. It is a reasonable heuristic but should be framed as such.
  5. [References] Reference [43] is an incomplete author list with 'K. Allen ... et al.'; [54] is a test instrument rather than a peer-reviewed article. Please verify the reference metadata.

Circularity Check

1 steps flagged

Minor definitional circularity in the measurement-gap evidence; the central proposal remains an independent hypothesis.

specific steps
  1. self definitional [Section 'Why is performance jagged?'; 'Charting the Data Gap'; 'Charting the Measurement Gap']
    "We call the latter cognitive dark matter (CDM): capabilities that materially shape intelligent behavior yet are under-measured and under-represented in the data and benchmarks that drive modern AI ... L3 covers functions aligned with cognitive dark matter ... the evaluation suites largely focus on L2 cognitive functions with little L3 function characterization."

    CDM is defined as being under-measured and under-represented in data and benchmarks. The paper then uses Liu et al.'s L3 tier, which is defined as 'rarely explored,' as the operational category for CDM. The finding that CDM-loaded functions are largely unmeasured therefore partly restates the definitions rather than being an independent discovery. The accompanying surveys and benchmark categorizations add specificity, but the headline 'measurement gap' is in part guaranteed by the taxonomy.

full rationale

This is a position paper with no fitted parameters, no equations, and no quantitative prediction that reduces to its inputs. The central claim that jagged intelligence arises from a missing training signal is an explanatory hypothesis, not a derivation. The empirical analyses use external data and an external taxonomy, and the supporting self-citations are background references, not load-bearing forced choices. The only noticeable circularity is definitional: CDM is defined as under-measured, and the L3 category used to demonstrate under-measurement is itself defined as 'rarely explored,' so part of the evidence for a measurement gap is built into the terminology. This weakens but does not collapse the paper's argument, which still offers an independently plausible research program.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 1 invented entities

No numbers are fitted; the free-parameter ledger is empty. The main weight-bearing choices are conceptual: the L1/L2/L3 taxonomy, the causal role of under-measurement, and the transferability of latent/process/neural signals. The only invented entity is CDM itself, which is an explanatory label, not a fitted constant.

axioms (3)
  • domain assumption The Liu et al. [23] three-tier taxonomy (L1/L2/L3) maps cognitive functions onto AI capability levels and is suitable for classifying benchmarks and datasets.
    The two main quantitative analyses (Figures 2 and 3) label benchmarks and papers according to this taxonomy; its validity is assumed rather than tested. AI-tier labels are conceptual, not measured.
  • ad hoc to paper Under-measurement in training data is a cause of under-performance in AI systems.
    This is the central causal premise: 'progress concentrates in capabilities where data and metrics are rich, and stalls where they are thin.' It is plausible and consistent with scaling experience, but it is an assumption; no causal test or counterfactual is provided.
  • ad hoc to paper Hidden mental processes are present in latent variables, process-tracing signals (eye-tracking, think-aloud), and neural recordings, and are learnable by AI from such data.
    The proposed training signal presupposes that these measurements carry usable information that is not already in behavioral outcome data. In the 'What to collect next' section, this is asserted as the 'loop' but not demonstrated for L3 functions.
invented entities (1)
  • Cognitive dark matter (CDM) no independent evidence
    purpose: Explanatory construct for the missing training signal; also the proposed target of measurement.
    It is defined by inference from jagged performance, not by a direct measurement or a falsifiable handle of its own; the paper explicitly analogizes it to cosmological dark matter, whose existence is inferred. The falsifiable consequences (improved AI via data collection) require future experiments.

pith-pipeline@v1.3.0-alltime-deepseek · 14193 in / 10802 out tokens · 79116 ms · 2026-08-03T02:32:43.313806+00:00 · methodology

0 comments
read the original abstract

We propose that the jagged intelligence landscape of modern AI systems arises from a missing training signal that we call ``cognitive dark matter'' (CDM): brain functions that meaningfully shape behavior yet are hard to infer from behavior alone. We identify key CDM domains---metacognition, cognitive flexibility, lifelong learning, abductive reasoning, social and common-sense reasoning, and emotional intelligence---and present evidence CDM-loaded functions are largely unmeasured in current AI benchmarks, and that large-scale neuroscience training datasets which could be used to instill these capabilities do not yet exist. We then outline a research program centered on three complementary data types designed to surface CDM for model training: (i) latent variables from large-scale cognitive models, (ii) process-tracing data such as eye-tracking and think-aloud protocols, and (iii) paired neural--behavioral data. These data will enable AI training on cognitive process rather than behavioral outcome alone, producing models with more general, less jagged intelligence. As a dual benefit, the same data will advance our understanding of human intelligence itself.

Figures

Figures reproduced from arXiv: 2603.03414 by Patrick J. Mineault, Sean Escola, Thomas L. Griffiths.

Figure 1
Figure 1. Figure 1: White to play. Notice that the position is [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Distribution across AI tiers of the cognitive capabilities focused on in intensive neuroimaging datasets, and [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Distribution across AI tiers of the benchmarks used by leading AI labs. We assigned AI tier labels to each [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners

    cs.AI 2026-05 unverdicted novelty 6.0

    Frontier LRMs match human game-learning behavior and predict fMRI signals an order of magnitude better than RL or Bayesian agents because of their in-context game-state representations.

  2. Process Matters more than Output for Distinguishing Humans from Machines

    cs.AI 2026-05 unverdicted novelty 6.0

    Process-level features from 30 cognitive tasks distinguish humans from frontier AI agents more effectively than task performance or output matching, achieving mean classifier AUC of 0.88, with fine-tuning experiments ...

  3. Process Matters more than Output for Distinguishing Humans from Machines

    cs.AI 2026-05 unverdicted novelty 6.0

    A new battery of 30 cognitive tasks demonstrates that process-level behavioral features distinguish humans from frontier AI agents better than performance metrics (mean AUC 0.88), with process-specific fine-tuning imp...

  4. Hypothesis generation and updating in large language models

    cs.LG 2026-05 unverdicted novelty 6.0

    LLMs exhibit Bayesian-like hypothesis updating with strong-sampling bias and an evaluation-generation gap but generalize poorly outside observed data.

Reference graph

Works this paper leans on

85 extracted references · 2 linked inside Pith · cited by 3 Pith papers

  1. [1]

    Im- ageNet classification with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Im- ageNet classification with deep convolutional neural networks,” inAdvances in Neural Information Pro- cessing Systems, F. Pereira, C. J. Burges, L. Bottou, and K. Q. Weinberger, Eds., vol. 25. Curran Asso- ciates, Inc., 2012

  2. [2]

    A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,

    D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Has- sabis, “A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,” Science, vol. 362, no. 6419, pp. 1140–1144, Dec. 2018

  3. [3]

    Accurate structure prediction of biomolecular interactions with AlphaFold 3,

    J. Abramson, J. Adler, J. Dunger, R. Evans, T. Green, A. Pritzel, O. Ronneberger, L. Willmore, A. J. Bal- lard, J. Bambrick, S. W. Bodenstein, D. A. Evans, 7 C.-C. Hung, M. O’Neill, D. Reiman, K. Tunyasuvu- nakool, Z. Wu, A. ˇZemgulyt˙ e, E. Arvaniti, C. Beattie, O. Bertolli, A. Bridgland, A. Cherepanov, M. Con- greve, A. I. Cowen-Rivers, A. Cowie, M. Fig...

  4. [4]

    Language models are Few-Shot learners,

    T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amod...

  5. [5]

    Dynabench: Rethinking benchmarking in NLP,

    D. Kiela, M. Bartolo, Y. Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia, Z. Ma, T. Thrush, S. Riedel, Z. Waseem, P. Stene- torp, R. Jia, M. Bansal, C. Potts, and A. Williams, “Dynabench: Rethinking benchmarking in NLP,” Apr. 2021

  6. [6]

    Navi- gating the jagged technological frontier: Field exper- imental evidence of the effects of AI on knowledge worker productivity and quality,

    F. Dell’Acqua, E. McFowland, E. R. Mollick, H. Lifshitz-Assaf, K. Kellogg, S. Rajendran, L. Krayer, F. Candelon, and K. R. Lakhani, “Navi- gating the jagged technological frontier: Field exper- imental evidence of the effects of AI on knowledge worker productivity and quality,”SSRN Electron. J., Sep. 2023

  7. [7]

    Measuring AI abil- ity to complete long tasks,

    T. Kwa, B. West, J. Becker, A. Deng, K. Gar- cia, M. Hasin, S. Jawhar, M. Kinniment, N. Rush, S. Von Arx, R. Bloom, T. Broadley, H. Du, B. Goodrich, N. Jurkovic, L. H. Miles, S. Nix, T. Lin, N. Parikh, D. Rein, L. J. K. Sato, H. Wijk, D. M. Ziegler, E. Barnes, and L. Chan, “Measuring AI abil- ity to complete long tasks,” Mar. 2025

  8. [8]

    Measuring AI ability to complete long tasks,

    METR, “Measuring AI ability to complete long tasks,” https://metr.org/blog/2025-03-19-measuring-ai- ability-to-complete-long-tasks/, Mar. 2025, accessed: 2026-1-12

  9. [9]

    Is there a half-life for the success rates of AI agents?

    T. Ord, “Is there a half-life for the success rates of AI agents?” May 2025

  10. [10]

    Toolformer: Language models can teach themselves to use tools,

    T. Schick, J. Dwivedi-Yu, R. Dessi, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” 2023. [Online]. Available: https://arxiv.org/abs/2302.04761

  11. [11]

    Faith and fate: Limits of trans- formers on compositionality,

    N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jiang, B. Y. Lin, P. West, C. Bhagavatula, R. L. Bras, J. D. Hwang, S. Sanyal, S. Welleck, X. Ren, A. Ettinger, Z. Har- chaoui, and Y. Choi, “Faith and fate: Limits of trans- formers on compositionality,” May 2023

  12. [12]

    SSL, JEPA, world models and the future of AI,

    Y. LeCun, “SSL, JEPA, world models and the future of AI,” Seminar, Sep. 2025

  13. [13]

    The breakdown of vigilance during prolonged visual search,

    N. H. Mackworth, “The breakdown of vigilance during prolonged visual search,”Q. J. Exp. Psychol., vol. 1, no. 1, pp. 6–21, Apr. 1948

  14. [14]

    Why are some problems hard? evidence from tower of hanoi,

    K. Kotovsky, J. R. Hayes, and H. A. Simon, “Why are some problems hard? evidence from tower of hanoi,”Cogn. Psychol., vol. 17, no. 2, pp. 248–294, Apr. 1985

  15. [15]

    Intuition in insight and noninsight problem solving,

    J. Metcalfe and D. Wiebe, “Intuition in insight and noninsight problem solving,”Mem. Cognit., vol. 15, no. 3, pp. 238–246, May 1987

  16. [16]

    Does incubation enhance problem solving? a meta-analytic review,

    U. N. Sio and T. C. Ormerod, “Does incubation enhance problem solving? a meta-analytic review,” Psychol. Bull., vol. 135, no. 1, pp. 94–120, Jan. 2009

  17. [17]

    Metacognition and confidence: A review and synthesis,

    S. M. Fleming, “Metacognition and confidence: A review and synthesis,”Annu. Rev. Psychol., vol. 75, no. 1, pp. 241–268, Jan. 2024

  18. [18]

    Improve mathematical reasoning in language models by automated process supervision,

    L. Luo, Y. Liu, R. Liu, S. Phatale, H. Lara, Y. Li, L. Shu, Y. Zhu, L. Meng, J. Sun, and A. Rastogi, “Improve mathematical reasoning in language models by automated process supervision,” Jun. 2024

  19. [19]

    Newell and H

    A. Newell and H. A. Simon,Human Problem Solving. Englewood Cliffs, N.J.: Prentice-Hall, 1972

  20. [20]

    T. L. Griffiths, N. Chater, and J. B. Tenenbaum, Bayesian Models of Cognition. Cambridge, MA: The MIT Press, Dec. 2024

  21. [21]

    Using large-scale exper- iments and machine learning to discover theories of human decision-making,

    J. C. Peterson, D. D. Bourgin, M. Agrawal, D. Re- ichman, and T. L. Griffiths, “Using large-scale exper- iments and machine learning to discover theories of human decision-making,”Science, vol. 372, no. 6547, pp. 1209–1214, Jun. 2021

  22. [22]

    Beyond playing 20 questions with nature: Integrative experiment design in the social and behavioral sciences,

    A. Almaatouq, T. L. Griffiths, J. W. Suchow, M. E. Whiting, J. Evans, and D. J. Watts, “Beyond playing 20 questions with nature: Integrative experiment design in the social and behavioral sciences,”Behav. Brain Sci., vol. 47, p. e33, Dec. 2022

  23. [23]

    Advances and challenges in foun- dation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems,

    B. Liu, X. Li, J. Zhang, J. Wang, T. He, S. Hong, H. Liu, S. Zhang, K. Song, K. Zhu, Y. Cheng, S. Wang, X. Wang, Y. Luo, H. Jin, P. Zhang, O. Liu, J. Chen, H. Zhang, Z. Yu, H. Shi, B. Li, D. Wu, F. Teng, X. Jia, J. Xu, J. Xiang, Y. Lin, T. Liu, T. Liu, Y. Su, H. Sun, G. Berseth, J. Nie, I. Fos- ter, L. Ward, Q. Wu, Y. Gu, M. Zhuge, X. Liang, 8 X. Tang, H....

  24. [24]

    Eye movement monitoring as a process tracing methodology in deci- sion making research,

    M. G. Glaholt and E. M. Reingold, “Eye movement monitoring as a process tracing methodology in deci- sion making research,”J. Neurosci. Psychol. Econ., vol. 4, no. 2, pp. 125–146, May 2011

  25. [25]

    Mouse tracking as a window into decision making,

    M. Maldonado, E. Dunbar, and E. Chemla, “Mouse tracking as a window into decision making,”Behav. Res. Methods, vol. 51, no. 3, pp. 1085–1101, Jun. 2019

  26. [26]

    Aligning machine and human visual representations across abstraction levels,

    L. Muttenthaler, K. Greff, F. Born, B. Spitzer, S. Ko- rnblith, M. C. Mozer, K.-R. M¨ uller, T. Unterthiner, and A. K. Lampinen, “Aligning machine and human visual representations across abstraction levels,”Na- ture, vol. 647, no. 8089, pp. 349–355, Nov. 2025

  27. [27]

    Schulte-Mecklenbeck, A

    M. Schulte-Mecklenbeck, A. Kuehberger, and J. G. Johnson,A handbook of process tracing methods: 2nd edition, 2nd ed., ser. The Society for Judgment and Decision Making Series, M. Schulte-Mecklenbeck, A. Kuehberger, and J. G. Johnson, Eds. London, England: Routledge, Jun. 2019

  28. [28]

    Improving semantic understanding in speech language models via brain-tuning,

    O. Moussa, D. Klakow, and M. Toneva, “Improving semantic understanding in speech language models via brain-tuning,” Oct. 2024

  29. [29]

    BrainWavLM: Fine-tuning speech rep- resentations with brain responses to language,

    N. Vattikonda, A. R. Vaidya, R. J. Antonello, and A. G. Huth, “BrainWavLM: Fine-tuning speech rep- resentations with brain responses to language,” Feb. 2025

  30. [30]

    The one where they brain-tune for social cognition: Multi- modal brain-tuning on friends,

    N. Policzer, C. Braunstein, and M. Toneva, “The one where they brain-tune for social cognition: Multi- modal brain-tuning on friends,” Nov. 2025

  31. [31]

    Learning from brains how to regularize machines,

    Z. Li, W. Brendel, E. Walker, E. Cobos, T. Muham- mad, J. Reimer, M. Bethge, F. Sinz, Z. Pitkow, and A. Tolias, “Learning from brains how to regularize machines,” inAdvances in Neural Information Pro- cessing Systems, vol. 32. Curran Associates, Inc., 2019

  32. [32]

    Simulating a primary vi- sual cortex at the front of CNNs improves robustness to image perturbations,

    J. Dapello, T. Marques, M. Schrimpf, F. Geiger, D. Cox, and J. J. DiCarlo, “Simulating a primary vi- sual cortex at the front of CNNs improves robustness to image perturbations,”Advances in Neural Infor- mation Processing Systems, vol. 33, pp. 13 073–13 087, 2020

  33. [33]

    Towards robust vision by multi-task learning on mon- key visual cortex,

    S. Safarani, A. Nix, K. Willeke, S. A. Cadena, K. Restivo, G. Denfield, A. S. Tolias, and F. H. Sinz, “Towards robust vision by multi-task learning on mon- key visual cortex,” Jul. 2021

  34. [34]

    Aligning model and macaque inferior temporal cortex representations improves model-to-human behavioral alignment and adversarial robustness,

    J. Dapello, K. Kar, M. Schrimpf, R. B. Geary, M. Ferguson, D. D. Cox, and J. J. DiCarlo, “Aligning model and macaque inferior temporal cortex representations improves model-to-human behavioral alignment and adversarial robustness,” inThe Eleventh International Conference on Learning Representations, 2023. [Online]. Available: https://openreview.net/forum?...

  35. [35]

    Using human brain activity to guide machine learning,

    R. C. Fong, W. J. Scheirer, and D. D. Cox, “Using human brain activity to guide machine learning,”Sci. Rep., vol. 8, no. 1, p. 5397, Mar. 2018

  36. [36]

    NeuroQuery, comprehensive meta-analysis of human brain mapping,

    J. Dock` es, R. A. Poldrack, R. Primet, H. G¨ oz¨ ukan, T. Yarkoni, F. Suchanek, B. Thirion, and G. Varo- quaux, “NeuroQuery, comprehensive meta-analysis of human brain mapping,”Elife, vol. 9, Mar. 2020

  37. [37]

    Gemini 3 pro: Model evalua- tion – approach, methodology and results,

    Google DeepMind, “Gemini 3 pro: Model evalua- tion – approach, methodology and results,” Google DeepMind, Tech. Rep., 2025

  38. [38]

    System card: Claude opus 4.5,

    Anthropic, “System card: Claude opus 4.5,” Tech. Rep., 2025

  39. [39]

    Introducing GPT-5.2,

    OpenAI, “Introducing GPT-5.2,” https://openai.c om/index/introducing-gpt-5-2/, 2025, accessed: 2026-2-2

  40. [40]

    GPT-4 technical report,

    OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ah- mad, I. Akkaya, F. L. Aleman, D. Almeida, J. Al- tenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. Boyd, A.-L. Brakman, G. Brockman, T. Brooks, M. Bru...

  41. [41]

    ARC-AGI-2: A new challenge for fron- tier AI reasoning systems,

    F. Chollet, M. Knoop, G. Kamradt, B. Landers, and H. Pinkard, “ARC-AGI-2: A new challenge for fron- tier AI reasoning systems,” May 2025

  42. [42]

    τ 2-bench: Evaluating conversational agents in a dual- control environment,

    V. Barres, H. Dong, S. Ray, X. Si, and K. Narasimhan, “τ 2-bench: Evaluating conversational agents in a dual- control environment,” Jun. 2025

  43. [43]

    Using games to understand the mind,

    K. Allen, F. Br¨ andle, M. Botvinick, J. E. Fan, S. J. Gershman, A. Gopnik, T. L. Griffiths, J. K. Hartshorne, T. U. Hauser, M. K. Hoet al., “Using games to understand the mind,”Nature Human Be- haviour, vol. 8, no. 6, pp. 1035–1043, 2024

  44. [44]

    Principles of intensive human neuroimaging,

    E. R. Kupers, T. Knapen, E. P. Merriam, and K. N. Kay, “Principles of intensive human neuroimaging,” Trends Neurosci., Oct. 2024

  45. [45]

    Large-scale neural recordings with single neuron resolution using neuropixels probes in human cortex,

    A. C. Paulk, Y. Kfir, A. R. Khanna, M. L. Mus- troph, E. M. Trautmann, D. J. Soper, S. D. Stavisky, M. Welkenhuysen, B. Dutta, K. V. Shenoy, L. R. Hochberg, R. M. Richardson, Z. M. Williams, and 9 S. S. Cash, “Large-scale neural recordings with single neuron resolution using neuropixels probes in human cortex,”Nat. Neurosci., vol. 25, no. 2, pp. 252–263, ...

  46. [46]

    AJILE12: Long- term naturalistic human intracranial neural record- ings and pose,

    S. M. Peterson, S. H. Singh, B. Dichter, M. Scheid, R. P. N. Rao, and B. W. Brunton, “AJILE12: Long- term naturalistic human intracranial neural record- ings and pose,”Sci Data, vol. 9, no. 1, p. 184, Apr. 2022

  47. [47]

    Resource-rational anal- ysis: Understanding human cognition as the optimal use of limited computational resources,

    F. Lieder and T. L. Griffiths, “Resource-rational anal- ysis: Understanding human cognition as the optimal use of limited computational resources,”Behav. Brain Sci., vol. 43, no. e1, p. e1, Feb. 2019

  48. [48]

    Self-evaluation of decision-making: A general bayesian framework for metacognitive computation,

    S. M. Fleming and N. D. Daw, “Self-evaluation of decision-making: A general bayesian framework for metacognitive computation,”Psychol. Rev., vol. 124, no. 1, pp. 91–114, Jan. 2017

  49. [49]

    Meta-reasoning: Monitoring and control of thinking and reasoning,

    R. Ackerman and V. A. Thompson, “Meta-reasoning: Monitoring and control of thinking and reasoning,” Trends Cogn. Sci., vol. 21, no. 8, pp. 607–617, Aug. 2017

  50. [50]

    Strategy selection as rational metareasoning,

    F. Lieder and T. L. Griffiths, “Strategy selection as rational metareasoning,”Psychol. Rev., vol. 124, no. 6, pp. 762–794, Nov. 2017

  51. [51]

    The trouble with overconfidence,

    D. A. Moore and P. J. Healy, “The trouble with overconfidence,”Psychol. Rev., vol. 115, no. 2, pp. 502–517, Apr. 2008

  52. [52]

    Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs,

    M. Xiong, Z. Hu, X. Lu, Y. Li, J. Fu, J. He, and B. Hooi, “Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs,” Jun. 2023

  53. [53]

    Why language models hallucinate,

    A. T. Kalai, O. Nachum, S. S. Vempala, and E. Zhang, “Why language models hallucinate,” Sep. 2025

  54. [54]

    Stroop color and word test,

    C. Golden, S. M. Freshwater, and Z. Golden, “Stroop color and word test,” Jul. 2012, title of the publication associated with this dataset: PsycTESTS Dataset

  55. [55]

    Prefrontal cortex lesions disrupt the contextual control of response conflict,

    J. E. Haddon and S. Killcross, “Prefrontal cortex lesions disrupt the contextual control of response conflict,”J. Neurosci., vol. 26, no. 11, pp. 2933–2940, Mar. 2006

  56. [56]

    Stroop-like effects for monkeys and humans: processing speed or strength of association?

    D. A. Washburn, “Stroop-like effects for monkeys and humans: processing speed or strength of association?” Psychol. Sci., vol. 5, no. 6, pp. 375–379, Nov. 1994

  57. [57]

    A definition of AGI,

    D. Hendrycks, D. Song, C. Szegedy, H. Lee, Y. Gal, E. Brynjolfsson, S. Li, A. Zou, L. Levine, B. Han, J. Fu, Z. Liu, J. Shin, K. Lee, M. Mazeika, L. Phan, G. Ingebretsen, A. Khoja, C. Xie, O. Salaudeen, M. Hein, K. Zhao, A. Pan, D. Duvenaud, B. Li, S. Omohundro, G. Alfour, M. Tegmark, K. McGrew, G. Marcus, J. Tallinn, E. Schmidt, and Y. Bengio, “A definit...

  58. [58]

    Human-level concept learning through probabilistic program induction,

    B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum, “Human-level concept learning through probabilistic program induction,”Science, vol. 350, no. 6266, pp. 1332–1338, Dec. 2015

  59. [59]

    The reversal curse: LLMs trained on “a is b

    L. Berglund, M. Tong, M. Kaufmann, M. Balesni, A. C. Stickland, T. Korbak, and O. Evans, “The reversal curse: LLMs trained on “a is b” fail to learn “b is a”,” Sep. 2023

  60. [60]

    Chain- of-thought prompting elicits reasoning in large lan- guage models,

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, “Chain- of-thought prompting elicits reasoning in large lan- guage models,” Jan. 2022

  61. [61]

    Bayes in the age of intelligent machines,

    T. L. Griffiths, J.-Q. Zhu, E. Grant, and R. Thomas McCoy, “Bayes in the age of intelligent machines,”Current Directions in Psychological Sci- ence, vol. 33, no. 5, pp. 283–291, 2024

  62. [62]

    An explanation of in-context learning as implicit bayesian inference,

    S. M. Xie, A. Raghunathan, P. Liang, and T. Ma, “An explanation of in-context learning as implicit bayesian inference,” Nov. 2021

  63. [63]

    Walton,Abductive Reasoning

    D. Walton,Abductive Reasoning. Tuscaloosa, AL: University of Alabama Press, May 2014

  64. [64]

    Intuitive physics,

    M. McCloskey, “Intuitive physics,”Sci. Am., vol. 248, no. 4, pp. 122–131, 1983

  65. [65]

    Theory- of-mind deficits and causal attributions,

    P. Kinderman, R. Dunbar, and R. P. Bentall, “Theory- of-mind deficits and causal attributions,”Br. J. Psy- chol., vol. 89, no. 2, pp. 191–204, May 1998

  66. [66]

    Evaluating the world model implicit in a generative model,

    K. Vafa, J. Y. Chen, A. Rambachan, J. Kleinberg, and S. Mullainathan, “Evaluating the world model implicit in a generative model,” Jun. 2024

  67. [67]

    Measuring emo- tional intelligence with the Mayer-Salovery-Caruso emotional intelligence test (MSCEIT),

    M. A. Brackett and P. Salovey, “Measuring emo- tional intelligence with the Mayer-Salovery-Caruso emotional intelligence test (MSCEIT),”Psicothema, vol. 18 Suppl, pp. 34–41, 2006

  68. [68]

    Emotion and decision making,

    J. S. Lerner, Y. Li, P. Valdesolo, and K. S. Kassam, “Emotion and decision making,”Annu. Rev. Psychol., vol. 66, no. 1, pp. 799–823, Jan. 2015

  69. [69]

    The role of affect in decision making,

    G. Loewenstein and J. S. Lerner, “The role of affect in decision making,” inHandbook of Affective Sciences, R. J. Davidson, K. R. Scherer, and H. H. Goldsmith, Eds. Oxford, England: Oxford University Press, 2003

  70. [70]

    The theory of constructed emotion: an active inference account of interoception and cate- gorization,

    L. F. Barrett, “The theory of constructed emotion: an active inference account of interoception and cate- gorization,”Soc. Cogn. Affect. Neurosci., p. nsw154, Oct. 2016. 10

  71. [71]

    R. W. Picard,Affective Computing, ser. The MIT Press. London, England: MIT Press, Jul. 2000

  72. [72]

    The social neuroscience of empathy,

    T. Singer and C. Lamm, “The social neuroscience of empathy,”Ann. N. Y. Acad. Sci., vol. 1156, no. 1, pp. 81–96, Mar. 2009

  73. [73]

    Crawford,Atlas of AI: Power, politics, and the planetary costs of artificial intelligence

    K. Crawford,Atlas of AI: Power, politics, and the planetary costs of artificial intelligence. Yale Uni- versity Press, Apr. 2021

  74. [74]

    On the morality of artificial agents,

    L. Floridi and J. W. Sanders, “On the morality of artificial agents,”Minds and Machines, 2004

  75. [75]

    A teen was suicidal. ChatGPT was the friend he confided in,

    K. Hill, “A teen was suicidal. ChatGPT was the friend he confided in,”The New York Times, Aug. 2025

  76. [76]

    How a chatbot encouraged a man who wanted to kill the queen,

    T. Singleton, T. Gerken, and L. McMahon, “How a chatbot encouraged a man who wanted to kill the queen,”BBC News, Oct. 2023

  77. [77]

    A massive 7T fMRI dataset to bridge cognitive neu- roscience and artificial intelligence,

    E. J. Allen, G. St-Yves, Y. Wu, J. L. Breedlove, J. S. Prince, L. T. Dowdle, M. Nau, B. Caron, F. Pestilli, I. Charest, J. B. Hutchinson, T. Naselaris, and K. Kay, “A massive 7T fMRI dataset to bridge cognitive neu- roscience and artificial intelligence,”Nat. Neurosci., vol. 25, no. 1, pp. 116–126, Jan. 2022

  78. [78]

    A large-scale standardized physiological survey reveals functional organization of the mouse visual cortex,

    S. E. J. de Vries, J. A. Lecoq, M. A. Buice, P. A. Groblewski, G. K. Ocker, M. Oliver, D. Feng, N. Cain, P. Ledochowitsch, and D. Millman, “A large-scale standardized physiological survey reveals functional organization of the mouse visual cortex,”Nat. Neu- rosci., vol. 23, no. 1, pp. 138–151, 2020

  79. [79]

    Brain-wide representations of prior infor- mation in mouse decision-making,

    C. Findling, F. Hubert, International Brain Labora- tory, L. Acerbi, B. Benson, J. Benson, D. Birman, N. Bonacchi, E. K. Buchanan, S. Bruijns, M. Caran- dini, J. A. Catarino, G. A. Chapuis, A. K. Churchland, Y. Dan, F. Davatolhagh, E. E. J. DeWitt, T. A. Engel, M. Fabbri, M. A. Faulkner, I. R. Fiete, L. Freitas- Silva, B. Gercek, K. D. Harris, M. H¨ ausse...

  80. [80]

    Time spent think- ing in online chess reflects the value of computation,

    E. M. Russek, D. Acosta-Kane, B. van Opheusden, M. G. Mattar, and T. L. Griffiths, “Time spent think- ing in online chess reflects the value of computation,” Cogn. Sci., vol. 49, no. 10, p. e70119, Oct. 2025

Showing first 80 references.