REVIEW 4 major objections 5 minor 4 cited by
This paper argues that the jagged intelligence of AI systems stems from a missing training signal—cognitive dark matter—and that collecting process-level data from human minds and brains can supply that signal.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 02:32 UTC pith:UZWF4JYM
load-bearing objection A wide-framing hypothesis paper that names a real pattern and a data agenda; the empirical support is thin, but the idea is worth taking seriously. the 4 major comments →
Cognitive Dark Matter: Measuring What AI Misses
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central claim is that the jagged intelligence landscape of AI systems arises from a missing training signal, not from lack of scale or raw capability. CDM-loaded functions are largely unmeasured in current AI benchmarks and in large-scale neuroscience datasets, so models never receive the signal needed to acquire them. The paper argues that making hidden cognitive processes visible—through latent variables, process-tracing data, and paired neural-behavioral data—would let AI train on process rather than outcome, producing models that generalize more smoothly and fail in human-legible ways. The authors present the measurement gap as both an opportunity and a warn
What carries the argument
The central object is cognitive dark matter (CDM), defined as brain functions that meaningfully shape behavior yet are hard to infer from behavior alone. The argument runs through a three-tier taxonomy used in the paper (L1: mastered capabilities like vision and language; L2: partial progress; L3: rarely explored CDM-loaded functions such as cognitive flexibility, social reasoning, and metacognition). The proposed mechanism is a trio of data types: latent variables sampled from large-scale cognitive models, process-tracing data (eye-tracking, mouse-tracking, think-aloud), and paired neural-behavioral data. These are meant to supply the missing training signal by exposing the causal cognitive
Load-bearing premise
The load-bearing premise is that hidden cognitive processes—metacognitive states, attention, neural activity—when added to training data are causally efficacious signals that transfer human-like capabilities to models; if they are only correlates of behavior, training on them may not close the jaggedness gap.
What would settle it
A controlled comparison: train matched models on the same CDM-heavy tasks with and without process-tracing or neural supervision, then test on novel L3 tasks such as rule-switch flexibility and confidence calibration. If process-supervised models show no improvement over behavior-only baselines at comparable data scale, the thesis that the missing signal is latent process data is falsified.
If this is right
- Benchmarking practice would shift: frontier model evaluation suites should include L3-heavy tests, because current suites systematically under-test the functions where jaggedness hides.
- The bottleneck for less jagged AI becomes data collection, not architecture: large, intensive datasets for metacognition, flexibility, social reasoning, and related functions become a prerequisite.
- Training on process data should make failures more human-legible—detectable and interpretable—because models would learn the cognitive processes that generate human outputs.
- Process supervision alone is insufficient for most tasks because correct intermediate steps are unknown; the paper's three data types are meant to supply those latent processes at scale.
- A dual benefit follows: even if AI transfer falls short, the datasets would be foundational for understanding human cognition.
Where Pith is reading between the lines
- An implicit prediction is that returns from process-level data should show up specifically on L3 capabilities; a scaling-law-style curve for eye-tracking or neural data against flexibility and calibration metrics would test this.
- The chess-endgame case suggests a concrete diagnostic: measuring whether a model abandons a failing strategy after feedback could serve as a quantitative metacognition score before and after CDM-style training.
- The thesis connects jaggedness to safety: chronic miscalibration, hallucination, and perseveration may share a single missing metacognitive signal, so better measurement could improve reliability as well as capability.
- If the latent signals turn out to be epiphenomenal, the paper's own dual-benefit framing still justifies the dataset program, but the AI-transfer claim would need a different mechanism.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes that the jagged capability landscape of modern AI systems is caused by a missing training signal, termed 'cognitive dark matter' (CDM): cognitive functions that shape human behavior but are hard to infer from observable behavior alone (metacognition, cognitive flexibility, lifelong learning, abductive reasoning, social/emotional reasoning). It argues that current AI benchmarks and large-scale neuroscience datasets both under-represent these L3 functions, and it proposes a research program based on three data types—latent variables from cognitive models, process-tracing data, and paired neural-behavioral data—to train AI on cognitive process rather than outcome. Three small empirical analyses are presented: a chess-app example, a survey of benchmark tiers in model release documents, and a survey of neuroimaging/literature coverage. The central claim is that better measurement of CDM will produce more general, less jagged AI, with the same data also advancing cognitive neuroscience.
Significance. If the transfer premise held, the paper would provide an actionable and fairly original research program that connects AI benchmarking, cognitive modeling, and large-scale neuroscience. The dual-benefit framing is attractive, and the authors are explicit that scaling laws for neural-data training are currently unknown. The paper also makes a falsifiable prediction: collecting the proposed data types should improve performance on CDM-loaded tasks beyond what outcome-only data of the same scale would provide. However, the empirical support is exploratory: the three analyses are small, partially manual, and not accompanied by inter-rater reliability or confidence intervals, and the central causal claim rests on an untested transfer assumption. The paper's strength is as a synthesis and proposal, not as a demonstration.
major comments (4)
- [CDM, Jagged Intelligence, and human–AI interactions; 'What to collect next'] The load-bearing claim is that process-tracing data, neural recordings, and cognitive-model latents can serve as training signals that instill metacognition, flexibility, and social reasoning into AI. No pilot, derivation, or mechanism is provided. The cited brain-tuning studies (refs 28–34) concern perceptual/semantic robustness, not L3 functions; process supervision (ref 18) rewards correct intermediate steps, which is not the same as latent cognitive process. The paper itself admits the absence of scaling laws. This transfer premise needs at least a proof-of-concept or a concrete falsifiable test, e.g., comparing CDM-enriched versus outcome-only supervision on the same L3 benchmarks with matched compute.
- [Methods: Chess endgame evaluation] The chess example is presented as evidence of missing metacognition, but it is based on N=3 models × 10 runs, with manual scoring and no inter-rater reliability, confidence intervals, or quantified failure rates. The interactive Claude Code sessions are described qualitatively as 'failed to elicit it.' As an illustrative anecdote this is acceptable, but as evidence for the core claim it is under-powered. Report the exact counts, scoring rubric, and inter-rater agreement, or explicitly label it as anecdotal.
- [Methods: AI benchmark analysis and neuroimaging surveys; Figures 2–3] The categorization of benchmarks and neuroimaging studies into the Liu et al. L1/L2/L3 taxonomy is subjective and mostly manual, with LLM-assisted tagging for some datasets. The 50 GPT-5.2 runs for benchmark classification are reported as 'stability' but no agreement statistic is given. Without inter-rater reliability or a validation of the LLM labels, the quantitative 'gap' shown in Figures 2 and 3 is suggestive but not established. The qualitative conclusion may survive, but the numerical presentation overstates precision.
- [CDM, Jagged Intelligence, and human–AI interactions] The paper asserts that under-measurement *causes* jagged performance, but it does not rule out alternative explanations: L3 tasks may be harder because they require more data in general, different architectures, or better world models, independent of whether the missing signal is 'cognitive process.' The constant-hazard-rate analysis of Ord (ref 9) suggests subtask success rates p and effective step count n; improving p could occur through many mechanisms, not specifically CDM-enriched data. The central hypothesis should be stated with a concrete discriminating experiment, otherwise it risks being unfalsifiable.
minor comments (5)
- [Abstract and body text] Typographical errors: 'isjagged', 'musthave', and missing spaces before citations in a few places. Figure 3 labels 'Claude 4.5' while the text and Methods say 'Claude Opus 4.5'.
- [Figure 2 caption] The figure caption does not define what counts as 'intensive neuroimaging datasets' or report the number of papers per source, making the y-axis scale hard to interpret. Please clarify the counts and sources.
- [Methods: AI benchmark analysis] The Methods say 'GPT-5.2 with high reasoning effort' was used to assign cognitive functions, but no details are given on the prompt, temperature, or how the 50 runs were aggregated. If the analysis is to be reproducible, these details are needed.
- [What to collect next] The '~500 hours' milestone for large-scale datasets is stated without a citation or quantitative justification. It is a reasonable heuristic but should be framed as such.
- [References] Reference [43] is an incomplete author list with 'K. Allen ... et al.'; [54] is a test instrument rather than a peer-reviewed article. Please verify the reference metadata.
Circularity Check
Minor definitional circularity in the measurement-gap evidence; the central proposal remains an independent hypothesis.
specific steps
-
self definitional
[Section 'Why is performance jagged?'; 'Charting the Data Gap'; 'Charting the Measurement Gap']
"We call the latter cognitive dark matter (CDM): capabilities that materially shape intelligent behavior yet are under-measured and under-represented in the data and benchmarks that drive modern AI ... L3 covers functions aligned with cognitive dark matter ... the evaluation suites largely focus on L2 cognitive functions with little L3 function characterization."
CDM is defined as being under-measured and under-represented in data and benchmarks. The paper then uses Liu et al.'s L3 tier, which is defined as 'rarely explored,' as the operational category for CDM. The finding that CDM-loaded functions are largely unmeasured therefore partly restates the definitions rather than being an independent discovery. The accompanying surveys and benchmark categorizations add specificity, but the headline 'measurement gap' is in part guaranteed by the taxonomy.
full rationale
This is a position paper with no fitted parameters, no equations, and no quantitative prediction that reduces to its inputs. The central claim that jagged intelligence arises from a missing training signal is an explanatory hypothesis, not a derivation. The empirical analyses use external data and an external taxonomy, and the supporting self-citations are background references, not load-bearing forced choices. The only noticeable circularity is definitional: CDM is defined as under-measured, and the L3 category used to demonstrate under-measurement is itself defined as 'rarely explored,' so part of the evidence for a measurement gap is built into the terminology. This weakens but does not collapse the paper's argument, which still offers an independently plausible research program.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption The Liu et al. [23] three-tier taxonomy (L1/L2/L3) maps cognitive functions onto AI capability levels and is suitable for classifying benchmarks and datasets.
- ad hoc to paper Under-measurement in training data is a cause of under-performance in AI systems.
- ad hoc to paper Hidden mental processes are present in latent variables, process-tracing signals (eye-tracking, think-aloud), and neural recordings, and are learnable by AI from such data.
invented entities (1)
-
Cognitive dark matter (CDM)
no independent evidence
read the original abstract
We propose that the jagged intelligence landscape of modern AI systems arises from a missing training signal that we call ``cognitive dark matter'' (CDM): brain functions that meaningfully shape behavior yet are hard to infer from behavior alone. We identify key CDM domains---metacognition, cognitive flexibility, lifelong learning, abductive reasoning, social and common-sense reasoning, and emotional intelligence---and present evidence CDM-loaded functions are largely unmeasured in current AI benchmarks, and that large-scale neuroscience training datasets which could be used to instill these capabilities do not yet exist. We then outline a research program centered on three complementary data types designed to surface CDM for model training: (i) latent variables from large-scale cognitive models, (ii) process-tracing data such as eye-tracking and think-aloud protocols, and (iii) paired neural--behavioral data. These data will enable AI training on cognitive process rather than behavioral outcome alone, producing models with more general, less jagged intelligence. As a dual benefit, the same data will advance our understanding of human intelligence itself.
Figures
Forward citations
Cited by 4 Pith papers
-
Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners
Frontier LRMs match human game-learning behavior and predict fMRI signals an order of magnitude better than RL or Bayesian agents because of their in-context game-state representations.
-
Process Matters more than Output for Distinguishing Humans from Machines
Process-level features from 30 cognitive tasks distinguish humans from frontier AI agents more effectively than task performance or output matching, achieving mean classifier AUC of 0.88, with fine-tuning experiments ...
-
Process Matters more than Output for Distinguishing Humans from Machines
A new battery of 30 cognitive tasks demonstrates that process-level behavioral features distinguish humans from frontier AI agents better than performance metrics (mean AUC 0.88), with process-specific fine-tuning imp...
-
Hypothesis generation and updating in large language models
LLMs exhibit Bayesian-like hypothesis updating with strong-sampling bias and an evaluation-generation gap but generalize poorly outside observed data.
Reference graph
Works this paper leans on
-
[1]
Im- ageNet classification with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Im- ageNet classification with deep convolutional neural networks,” inAdvances in Neural Information Pro- cessing Systems, F. Pereira, C. J. Burges, L. Bottou, and K. Q. Weinberger, Eds., vol. 25. Curran Asso- ciates, Inc., 2012
2012
-
[2]
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Has- sabis, “A general reinforcement learning algorithm that masters chess, shogi, and go through self-play,” Science, vol. 362, no. 6419, pp. 1140–1144, Dec. 2018
2018
-
[3]
Accurate structure prediction of biomolecular interactions with AlphaFold 3,
J. Abramson, J. Adler, J. Dunger, R. Evans, T. Green, A. Pritzel, O. Ronneberger, L. Willmore, A. J. Bal- lard, J. Bambrick, S. W. Bodenstein, D. A. Evans, 7 C.-C. Hung, M. O’Neill, D. Reiman, K. Tunyasuvu- nakool, Z. Wu, A. ˇZemgulyt˙ e, E. Arvaniti, C. Beattie, O. Bertolli, A. Bridgland, A. Cherepanov, M. Con- greve, A. I. Cowen-Rivers, A. Cowie, M. Fig...
2024
-
[4]
Language models are Few-Shot learners,
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amod...
2020
-
[5]
Dynabench: Rethinking benchmarking in NLP,
D. Kiela, M. Bartolo, Y. Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia, Z. Ma, T. Thrush, S. Riedel, Z. Waseem, P. Stene- torp, R. Jia, M. Bansal, C. Potts, and A. Williams, “Dynabench: Rethinking benchmarking in NLP,” Apr. 2021
2021
-
[6]
Navi- gating the jagged technological frontier: Field exper- imental evidence of the effects of AI on knowledge worker productivity and quality,
F. Dell’Acqua, E. McFowland, E. R. Mollick, H. Lifshitz-Assaf, K. Kellogg, S. Rajendran, L. Krayer, F. Candelon, and K. R. Lakhani, “Navi- gating the jagged technological frontier: Field exper- imental evidence of the effects of AI on knowledge worker productivity and quality,”SSRN Electron. J., Sep. 2023
2023
-
[7]
Measuring AI abil- ity to complete long tasks,
T. Kwa, B. West, J. Becker, A. Deng, K. Gar- cia, M. Hasin, S. Jawhar, M. Kinniment, N. Rush, S. Von Arx, R. Bloom, T. Broadley, H. Du, B. Goodrich, N. Jurkovic, L. H. Miles, S. Nix, T. Lin, N. Parikh, D. Rein, L. J. K. Sato, H. Wijk, D. M. Ziegler, E. Barnes, and L. Chan, “Measuring AI abil- ity to complete long tasks,” Mar. 2025
2025
-
[8]
Measuring AI ability to complete long tasks,
METR, “Measuring AI ability to complete long tasks,” https://metr.org/blog/2025-03-19-measuring-ai- ability-to-complete-long-tasks/, Mar. 2025, accessed: 2026-1-12
2025
-
[9]
Is there a half-life for the success rates of AI agents?
T. Ord, “Is there a half-life for the success rates of AI agents?” May 2025
2025
-
[10]
Toolformer: Language models can teach themselves to use tools,
T. Schick, J. Dwivedi-Yu, R. Dessi, R. Raileanu, M. Lomeli, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” 2023. [Online]. Available: https://arxiv.org/abs/2302.04761
Pith/arXiv arXiv 2023
-
[11]
Faith and fate: Limits of trans- formers on compositionality,
N. Dziri, X. Lu, M. Sclar, X. L. Li, L. Jiang, B. Y. Lin, P. West, C. Bhagavatula, R. L. Bras, J. D. Hwang, S. Sanyal, S. Welleck, X. Ren, A. Ettinger, Z. Har- chaoui, and Y. Choi, “Faith and fate: Limits of trans- formers on compositionality,” May 2023
2023
-
[12]
SSL, JEPA, world models and the future of AI,
Y. LeCun, “SSL, JEPA, world models and the future of AI,” Seminar, Sep. 2025
2025
-
[13]
The breakdown of vigilance during prolonged visual search,
N. H. Mackworth, “The breakdown of vigilance during prolonged visual search,”Q. J. Exp. Psychol., vol. 1, no. 1, pp. 6–21, Apr. 1948
1948
-
[14]
Why are some problems hard? evidence from tower of hanoi,
K. Kotovsky, J. R. Hayes, and H. A. Simon, “Why are some problems hard? evidence from tower of hanoi,”Cogn. Psychol., vol. 17, no. 2, pp. 248–294, Apr. 1985
1985
-
[15]
Intuition in insight and noninsight problem solving,
J. Metcalfe and D. Wiebe, “Intuition in insight and noninsight problem solving,”Mem. Cognit., vol. 15, no. 3, pp. 238–246, May 1987
1987
-
[16]
Does incubation enhance problem solving? a meta-analytic review,
U. N. Sio and T. C. Ormerod, “Does incubation enhance problem solving? a meta-analytic review,” Psychol. Bull., vol. 135, no. 1, pp. 94–120, Jan. 2009
2009
-
[17]
Metacognition and confidence: A review and synthesis,
S. M. Fleming, “Metacognition and confidence: A review and synthesis,”Annu. Rev. Psychol., vol. 75, no. 1, pp. 241–268, Jan. 2024
2024
-
[18]
Improve mathematical reasoning in language models by automated process supervision,
L. Luo, Y. Liu, R. Liu, S. Phatale, H. Lara, Y. Li, L. Shu, Y. Zhu, L. Meng, J. Sun, and A. Rastogi, “Improve mathematical reasoning in language models by automated process supervision,” Jun. 2024
2024
-
[19]
Newell and H
A. Newell and H. A. Simon,Human Problem Solving. Englewood Cliffs, N.J.: Prentice-Hall, 1972
1972
-
[20]
T. L. Griffiths, N. Chater, and J. B. Tenenbaum, Bayesian Models of Cognition. Cambridge, MA: The MIT Press, Dec. 2024
2024
-
[21]
Using large-scale exper- iments and machine learning to discover theories of human decision-making,
J. C. Peterson, D. D. Bourgin, M. Agrawal, D. Re- ichman, and T. L. Griffiths, “Using large-scale exper- iments and machine learning to discover theories of human decision-making,”Science, vol. 372, no. 6547, pp. 1209–1214, Jun. 2021
2021
-
[22]
Beyond playing 20 questions with nature: Integrative experiment design in the social and behavioral sciences,
A. Almaatouq, T. L. Griffiths, J. W. Suchow, M. E. Whiting, J. Evans, and D. J. Watts, “Beyond playing 20 questions with nature: Integrative experiment design in the social and behavioral sciences,”Behav. Brain Sci., vol. 47, p. e33, Dec. 2022
2022
-
[23]
Advances and challenges in foun- dation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems,
B. Liu, X. Li, J. Zhang, J. Wang, T. He, S. Hong, H. Liu, S. Zhang, K. Song, K. Zhu, Y. Cheng, S. Wang, X. Wang, Y. Luo, H. Jin, P. Zhang, O. Liu, J. Chen, H. Zhang, Z. Yu, H. Shi, B. Li, D. Wu, F. Teng, X. Jia, J. Xu, J. Xiang, Y. Lin, T. Liu, T. Liu, Y. Su, H. Sun, G. Berseth, J. Nie, I. Fos- ter, L. Ward, Q. Wu, Y. Gu, M. Zhuge, X. Liang, 8 X. Tang, H....
2025
-
[24]
Eye movement monitoring as a process tracing methodology in deci- sion making research,
M. G. Glaholt and E. M. Reingold, “Eye movement monitoring as a process tracing methodology in deci- sion making research,”J. Neurosci. Psychol. Econ., vol. 4, no. 2, pp. 125–146, May 2011
2011
-
[25]
Mouse tracking as a window into decision making,
M. Maldonado, E. Dunbar, and E. Chemla, “Mouse tracking as a window into decision making,”Behav. Res. Methods, vol. 51, no. 3, pp. 1085–1101, Jun. 2019
2019
-
[26]
Aligning machine and human visual representations across abstraction levels,
L. Muttenthaler, K. Greff, F. Born, B. Spitzer, S. Ko- rnblith, M. C. Mozer, K.-R. M¨ uller, T. Unterthiner, and A. K. Lampinen, “Aligning machine and human visual representations across abstraction levels,”Na- ture, vol. 647, no. 8089, pp. 349–355, Nov. 2025
2025
-
[27]
Schulte-Mecklenbeck, A
M. Schulte-Mecklenbeck, A. Kuehberger, and J. G. Johnson,A handbook of process tracing methods: 2nd edition, 2nd ed., ser. The Society for Judgment and Decision Making Series, M. Schulte-Mecklenbeck, A. Kuehberger, and J. G. Johnson, Eds. London, England: Routledge, Jun. 2019
2019
-
[28]
Improving semantic understanding in speech language models via brain-tuning,
O. Moussa, D. Klakow, and M. Toneva, “Improving semantic understanding in speech language models via brain-tuning,” Oct. 2024
2024
-
[29]
BrainWavLM: Fine-tuning speech rep- resentations with brain responses to language,
N. Vattikonda, A. R. Vaidya, R. J. Antonello, and A. G. Huth, “BrainWavLM: Fine-tuning speech rep- resentations with brain responses to language,” Feb. 2025
2025
-
[30]
The one where they brain-tune for social cognition: Multi- modal brain-tuning on friends,
N. Policzer, C. Braunstein, and M. Toneva, “The one where they brain-tune for social cognition: Multi- modal brain-tuning on friends,” Nov. 2025
2025
-
[31]
Learning from brains how to regularize machines,
Z. Li, W. Brendel, E. Walker, E. Cobos, T. Muham- mad, J. Reimer, M. Bethge, F. Sinz, Z. Pitkow, and A. Tolias, “Learning from brains how to regularize machines,” inAdvances in Neural Information Pro- cessing Systems, vol. 32. Curran Associates, Inc., 2019
2019
-
[32]
Simulating a primary vi- sual cortex at the front of CNNs improves robustness to image perturbations,
J. Dapello, T. Marques, M. Schrimpf, F. Geiger, D. Cox, and J. J. DiCarlo, “Simulating a primary vi- sual cortex at the front of CNNs improves robustness to image perturbations,”Advances in Neural Infor- mation Processing Systems, vol. 33, pp. 13 073–13 087, 2020
2020
-
[33]
Towards robust vision by multi-task learning on mon- key visual cortex,
S. Safarani, A. Nix, K. Willeke, S. A. Cadena, K. Restivo, G. Denfield, A. S. Tolias, and F. H. Sinz, “Towards robust vision by multi-task learning on mon- key visual cortex,” Jul. 2021
2021
-
[34]
Aligning model and macaque inferior temporal cortex representations improves model-to-human behavioral alignment and adversarial robustness,
J. Dapello, K. Kar, M. Schrimpf, R. B. Geary, M. Ferguson, D. D. Cox, and J. J. DiCarlo, “Aligning model and macaque inferior temporal cortex representations improves model-to-human behavioral alignment and adversarial robustness,” inThe Eleventh International Conference on Learning Representations, 2023. [Online]. Available: https://openreview.net/forum?...
2023
-
[35]
Using human brain activity to guide machine learning,
R. C. Fong, W. J. Scheirer, and D. D. Cox, “Using human brain activity to guide machine learning,”Sci. Rep., vol. 8, no. 1, p. 5397, Mar. 2018
2018
-
[36]
NeuroQuery, comprehensive meta-analysis of human brain mapping,
J. Dock` es, R. A. Poldrack, R. Primet, H. G¨ oz¨ ukan, T. Yarkoni, F. Suchanek, B. Thirion, and G. Varo- quaux, “NeuroQuery, comprehensive meta-analysis of human brain mapping,”Elife, vol. 9, Mar. 2020
2020
-
[37]
Gemini 3 pro: Model evalua- tion – approach, methodology and results,
Google DeepMind, “Gemini 3 pro: Model evalua- tion – approach, methodology and results,” Google DeepMind, Tech. Rep., 2025
2025
-
[38]
System card: Claude opus 4.5,
Anthropic, “System card: Claude opus 4.5,” Tech. Rep., 2025
2025
-
[39]
Introducing GPT-5.2,
OpenAI, “Introducing GPT-5.2,” https://openai.c om/index/introducing-gpt-5-2/, 2025, accessed: 2026-2-2
2025
-
[40]
GPT-4 technical report,
OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ah- mad, I. Akkaya, F. L. Aleman, D. Almeida, J. Al- tenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. Boyd, A.-L. Brakman, G. Brockman, T. Brooks, M. Bru...
2023
-
[41]
ARC-AGI-2: A new challenge for fron- tier AI reasoning systems,
F. Chollet, M. Knoop, G. Kamradt, B. Landers, and H. Pinkard, “ARC-AGI-2: A new challenge for fron- tier AI reasoning systems,” May 2025
2025
-
[42]
τ 2-bench: Evaluating conversational agents in a dual- control environment,
V. Barres, H. Dong, S. Ray, X. Si, and K. Narasimhan, “τ 2-bench: Evaluating conversational agents in a dual- control environment,” Jun. 2025
2025
-
[43]
Using games to understand the mind,
K. Allen, F. Br¨ andle, M. Botvinick, J. E. Fan, S. J. Gershman, A. Gopnik, T. L. Griffiths, J. K. Hartshorne, T. U. Hauser, M. K. Hoet al., “Using games to understand the mind,”Nature Human Be- haviour, vol. 8, no. 6, pp. 1035–1043, 2024
2024
-
[44]
Principles of intensive human neuroimaging,
E. R. Kupers, T. Knapen, E. P. Merriam, and K. N. Kay, “Principles of intensive human neuroimaging,” Trends Neurosci., Oct. 2024
2024
-
[45]
Large-scale neural recordings with single neuron resolution using neuropixels probes in human cortex,
A. C. Paulk, Y. Kfir, A. R. Khanna, M. L. Mus- troph, E. M. Trautmann, D. J. Soper, S. D. Stavisky, M. Welkenhuysen, B. Dutta, K. V. Shenoy, L. R. Hochberg, R. M. Richardson, Z. M. Williams, and 9 S. S. Cash, “Large-scale neural recordings with single neuron resolution using neuropixels probes in human cortex,”Nat. Neurosci., vol. 25, no. 2, pp. 252–263, ...
2022
-
[46]
AJILE12: Long- term naturalistic human intracranial neural record- ings and pose,
S. M. Peterson, S. H. Singh, B. Dichter, M. Scheid, R. P. N. Rao, and B. W. Brunton, “AJILE12: Long- term naturalistic human intracranial neural record- ings and pose,”Sci Data, vol. 9, no. 1, p. 184, Apr. 2022
2022
-
[47]
Resource-rational anal- ysis: Understanding human cognition as the optimal use of limited computational resources,
F. Lieder and T. L. Griffiths, “Resource-rational anal- ysis: Understanding human cognition as the optimal use of limited computational resources,”Behav. Brain Sci., vol. 43, no. e1, p. e1, Feb. 2019
2019
-
[48]
Self-evaluation of decision-making: A general bayesian framework for metacognitive computation,
S. M. Fleming and N. D. Daw, “Self-evaluation of decision-making: A general bayesian framework for metacognitive computation,”Psychol. Rev., vol. 124, no. 1, pp. 91–114, Jan. 2017
2017
-
[49]
Meta-reasoning: Monitoring and control of thinking and reasoning,
R. Ackerman and V. A. Thompson, “Meta-reasoning: Monitoring and control of thinking and reasoning,” Trends Cogn. Sci., vol. 21, no. 8, pp. 607–617, Aug. 2017
2017
-
[50]
Strategy selection as rational metareasoning,
F. Lieder and T. L. Griffiths, “Strategy selection as rational metareasoning,”Psychol. Rev., vol. 124, no. 6, pp. 762–794, Nov. 2017
2017
-
[51]
The trouble with overconfidence,
D. A. Moore and P. J. Healy, “The trouble with overconfidence,”Psychol. Rev., vol. 115, no. 2, pp. 502–517, Apr. 2008
2008
-
[52]
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs,
M. Xiong, Z. Hu, X. Lu, Y. Li, J. Fu, J. He, and B. Hooi, “Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs,” Jun. 2023
2023
-
[53]
Why language models hallucinate,
A. T. Kalai, O. Nachum, S. S. Vempala, and E. Zhang, “Why language models hallucinate,” Sep. 2025
2025
-
[54]
Stroop color and word test,
C. Golden, S. M. Freshwater, and Z. Golden, “Stroop color and word test,” Jul. 2012, title of the publication associated with this dataset: PsycTESTS Dataset
2012
-
[55]
Prefrontal cortex lesions disrupt the contextual control of response conflict,
J. E. Haddon and S. Killcross, “Prefrontal cortex lesions disrupt the contextual control of response conflict,”J. Neurosci., vol. 26, no. 11, pp. 2933–2940, Mar. 2006
2006
-
[56]
Stroop-like effects for monkeys and humans: processing speed or strength of association?
D. A. Washburn, “Stroop-like effects for monkeys and humans: processing speed or strength of association?” Psychol. Sci., vol. 5, no. 6, pp. 375–379, Nov. 1994
1994
-
[57]
A definition of AGI,
D. Hendrycks, D. Song, C. Szegedy, H. Lee, Y. Gal, E. Brynjolfsson, S. Li, A. Zou, L. Levine, B. Han, J. Fu, Z. Liu, J. Shin, K. Lee, M. Mazeika, L. Phan, G. Ingebretsen, A. Khoja, C. Xie, O. Salaudeen, M. Hein, K. Zhao, A. Pan, D. Duvenaud, B. Li, S. Omohundro, G. Alfour, M. Tegmark, K. McGrew, G. Marcus, J. Tallinn, E. Schmidt, and Y. Bengio, “A definit...
2025
-
[58]
Human-level concept learning through probabilistic program induction,
B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum, “Human-level concept learning through probabilistic program induction,”Science, vol. 350, no. 6266, pp. 1332–1338, Dec. 2015
2015
-
[59]
The reversal curse: LLMs trained on “a is b
L. Berglund, M. Tong, M. Kaufmann, M. Balesni, A. C. Stickland, T. Korbak, and O. Evans, “The reversal curse: LLMs trained on “a is b” fail to learn “b is a”,” Sep. 2023
2023
-
[60]
Chain- of-thought prompting elicits reasoning in large lan- guage models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou, “Chain- of-thought prompting elicits reasoning in large lan- guage models,” Jan. 2022
2022
-
[61]
Bayes in the age of intelligent machines,
T. L. Griffiths, J.-Q. Zhu, E. Grant, and R. Thomas McCoy, “Bayes in the age of intelligent machines,”Current Directions in Psychological Sci- ence, vol. 33, no. 5, pp. 283–291, 2024
2024
-
[62]
An explanation of in-context learning as implicit bayesian inference,
S. M. Xie, A. Raghunathan, P. Liang, and T. Ma, “An explanation of in-context learning as implicit bayesian inference,” Nov. 2021
2021
-
[63]
Walton,Abductive Reasoning
D. Walton,Abductive Reasoning. Tuscaloosa, AL: University of Alabama Press, May 2014
2014
-
[64]
Intuitive physics,
M. McCloskey, “Intuitive physics,”Sci. Am., vol. 248, no. 4, pp. 122–131, 1983
1983
-
[65]
Theory- of-mind deficits and causal attributions,
P. Kinderman, R. Dunbar, and R. P. Bentall, “Theory- of-mind deficits and causal attributions,”Br. J. Psy- chol., vol. 89, no. 2, pp. 191–204, May 1998
1998
-
[66]
Evaluating the world model implicit in a generative model,
K. Vafa, J. Y. Chen, A. Rambachan, J. Kleinberg, and S. Mullainathan, “Evaluating the world model implicit in a generative model,” Jun. 2024
2024
-
[67]
Measuring emo- tional intelligence with the Mayer-Salovery-Caruso emotional intelligence test (MSCEIT),
M. A. Brackett and P. Salovey, “Measuring emo- tional intelligence with the Mayer-Salovery-Caruso emotional intelligence test (MSCEIT),”Psicothema, vol. 18 Suppl, pp. 34–41, 2006
2006
-
[68]
Emotion and decision making,
J. S. Lerner, Y. Li, P. Valdesolo, and K. S. Kassam, “Emotion and decision making,”Annu. Rev. Psychol., vol. 66, no. 1, pp. 799–823, Jan. 2015
2015
-
[69]
The role of affect in decision making,
G. Loewenstein and J. S. Lerner, “The role of affect in decision making,” inHandbook of Affective Sciences, R. J. Davidson, K. R. Scherer, and H. H. Goldsmith, Eds. Oxford, England: Oxford University Press, 2003
2003
-
[70]
The theory of constructed emotion: an active inference account of interoception and cate- gorization,
L. F. Barrett, “The theory of constructed emotion: an active inference account of interoception and cate- gorization,”Soc. Cogn. Affect. Neurosci., p. nsw154, Oct. 2016. 10
2016
-
[71]
R. W. Picard,Affective Computing, ser. The MIT Press. London, England: MIT Press, Jul. 2000
2000
-
[72]
The social neuroscience of empathy,
T. Singer and C. Lamm, “The social neuroscience of empathy,”Ann. N. Y. Acad. Sci., vol. 1156, no. 1, pp. 81–96, Mar. 2009
2009
-
[73]
Crawford,Atlas of AI: Power, politics, and the planetary costs of artificial intelligence
K. Crawford,Atlas of AI: Power, politics, and the planetary costs of artificial intelligence. Yale Uni- versity Press, Apr. 2021
2021
-
[74]
On the morality of artificial agents,
L. Floridi and J. W. Sanders, “On the morality of artificial agents,”Minds and Machines, 2004
2004
-
[75]
A teen was suicidal. ChatGPT was the friend he confided in,
K. Hill, “A teen was suicidal. ChatGPT was the friend he confided in,”The New York Times, Aug. 2025
2025
-
[76]
How a chatbot encouraged a man who wanted to kill the queen,
T. Singleton, T. Gerken, and L. McMahon, “How a chatbot encouraged a man who wanted to kill the queen,”BBC News, Oct. 2023
2023
-
[77]
A massive 7T fMRI dataset to bridge cognitive neu- roscience and artificial intelligence,
E. J. Allen, G. St-Yves, Y. Wu, J. L. Breedlove, J. S. Prince, L. T. Dowdle, M. Nau, B. Caron, F. Pestilli, I. Charest, J. B. Hutchinson, T. Naselaris, and K. Kay, “A massive 7T fMRI dataset to bridge cognitive neu- roscience and artificial intelligence,”Nat. Neurosci., vol. 25, no. 1, pp. 116–126, Jan. 2022
2022
-
[78]
A large-scale standardized physiological survey reveals functional organization of the mouse visual cortex,
S. E. J. de Vries, J. A. Lecoq, M. A. Buice, P. A. Groblewski, G. K. Ocker, M. Oliver, D. Feng, N. Cain, P. Ledochowitsch, and D. Millman, “A large-scale standardized physiological survey reveals functional organization of the mouse visual cortex,”Nat. Neu- rosci., vol. 23, no. 1, pp. 138–151, 2020
2020
-
[79]
Brain-wide representations of prior infor- mation in mouse decision-making,
C. Findling, F. Hubert, International Brain Labora- tory, L. Acerbi, B. Benson, J. Benson, D. Birman, N. Bonacchi, E. K. Buchanan, S. Bruijns, M. Caran- dini, J. A. Catarino, G. A. Chapuis, A. K. Churchland, Y. Dan, F. Davatolhagh, E. E. J. DeWitt, T. A. Engel, M. Fabbri, M. A. Faulkner, I. R. Fiete, L. Freitas- Silva, B. Gercek, K. D. Harris, M. H¨ ausse...
2025
-
[80]
Time spent think- ing in online chess reflects the value of computation,
E. M. Russek, D. Acosta-Kane, B. van Opheusden, M. G. Mattar, and T. L. Griffiths, “Time spent think- ing in online chess reflects the value of computation,” Cogn. Sci., vol. 49, no. 10, p. e70119, Oct. 2025
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.