Pith. sign in

REVIEW 2 major objections 4 minor 112 references

On the use of foundation models in cognitive science

T0 review · 2 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper argues that behavioral fit alone cannot justify treating foundation models as explanatory cognitive models; alignment becomes meaningful only within explicit theory, diagnostic tasks, and contrastive evaluation.

desk verdict A clear, honest synthesis of the inferential steps between foundation-model outputs and human behavior; the framework is useful even though the strongest necessity claims stay underdetermined. read the letter →

arxiv 2608.07812 v1 pith:6ZHX24EP submitted 2026-08-07 cs.CL

classification cs.CL
keywords foundationmodelscognitivemodelinglinkinghypothesesbehavioralalignmentmodelevaluationsciencelargelanguagedevelopmental
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Foundation models can reproduce many behavioral signatures of human cognition, but the paper argues that such behavioral fit is not itself evidence that a model explains how humans think. The authors propose a four-stage inferential framework—task adaptation, linking hypothesis specification, behavioral evaluation, and contrastive model comparison—as the minimal structure under which alignment becomes scientifically meaningful. The central message is that alignment is a property of a model evaluated under explicit theoretical commitments and diagnostic contrasts, not a property of the model alone. A sympathetic reader would take away that most current demonstrations of model–human similarity are proof-of-possibility results, not explanatory claims, until embedded in this structure.

What carries the argument

The central object is the four-stage inferential framework, split into an inner alignment loop (Stages 1–3: task adaptation, linking hypothesis, goodness-of-fit evaluation) and an outer contrastive loop (Stage 4: cross-model and manipulation comparison). The linking hypothesis is the load-bearing pivot of the framework: it defines what counts as evidence of alignment by mapping model outputs to human behavioral measures, and the paper argues that a strong fit under one linking hypothesis can vanish under another. The running example is the translation of progressive matrix analogies into symbolic digit matrices [107], which the paper uses to show how each stage changes the interpretation of an alignment result.

What would settle it

Find two foundation models with architecturally distinct mechanisms that both pass all four stages on the same human dataset, and show that no manipulation of task, linking hypothesis, or training regime separates their predictions; such a result would show that the contrastive stage cannot identify which computational features are necessary, undermining the framework's explanatory criterion.

Watch

Extended reading notes

Core claim

The paper's central claim is that behavioral alignment justifies treating a foundation model as an explanatory cognitive model only when the evaluation embeds the model in a theory-diagnostic design: adapt the task so model and humans perform functionally equivalent problems, specify a linking hypothesis that maps model outputs to human measures, evaluate fit on theoretically diagnostic contrasts rather than aggregate scores, and compare across candidate models or manipulations to identify which computational features are necessary for the fit. Under this view, a model that simply gets high accuracy or matches average human judgments remains a behavioral proxy. The paper supports the claim by showing how each of four common linking hypotheses—similarity in representational space, surprisal, prompting, and process-trace analysis—carries different theoretical commitments, and by identifying four challenges that constrain alignment claims: theoretical underdetermination, mechanistic opacity, training and developmental mismatch, and population-level variability.

Load-bearing premise

The framework assumes that adding explicit theory, diagnostic contrasts, and contrastive model comparison can overcome underdetermination and mechanistic opacity—if the same behavior can arise from very different internal computations, even a model that passes all four stages may not reveal the mechanisms of human cognition.

Editorial extensions

If this is right

  • Single-model demonstrations, however strong, establish at most that a cognitive signature is reproducible, not that the model's mechanisms explain it.
  • Evaluation of alignment without diagnostic contrasts—comparing conditions that discriminate between theories—cannot adjudicate between cognitive accounts.
  • Developmental alignment claims require more than matching learning curves with children; they need causal manipulations of training regime or data ordering.
  • Because linking hypotheses are not neutral, results should be triangulated across multiple linking assumptions before drawing explanatory conclusions.
  • The framework gives concrete shape to reporting standards for model–cognition studies: state the theory, the adaptation, the linking hypothesis, the contrast, and the comparison set.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the framework is right, a large share of published model–behavior benchmarks should be reinterpreted as capability or similarity studies rather than cognitive models, unless they already include diagnostic contrasts and model comparison.
  • A testable extension: reanalyzing landmark alignment results under alternative linking hypotheses should change the apparent fit, and the direction of change would reveal which theoretical commitment is doing the work.
  • The framework implies that the field's next bottleneck is not larger models but better theory: designing tasks whose contrasts discriminate between computational mechanisms, and reporting null or negative contrasts as informative.
  • The four-stage structure could generalize to other 'black box' scientific models, wherever mere behavioral fit risks being mistaken for mechanistic insight.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper is a perspective article that asks under what conditions behavioral alignment between foundation models (FMs) and human performance can justify treating FMs as explanatory models of cognition. It proposes a four-stage inferential framework: (1) adapting human tasks to model-compatible formats, (2) specifying linking hypotheses that map model outputs to human measures, (3) evaluating behavioral correspondence with attention to diagnostic contrasts, and (4) comparing across candidate models or manipulations. The authors argue that behavioral fit alone is insufficient and that meaningful alignment requires explicit theoretical commitments, theory-diagnostic tasks, and contrastive evaluation. They discuss four linking hypotheses (similarity, surprisal, prompting, process-trace) and four challenges (theoretical underdetermination, mechanistic opacity, training/developmental mismatch, and population variability), then distill the framework into five research guidelines, using the digit-matrix analogical reasoning task as a running example.

Significance. If the framework is adopted, it would provide a common vocabulary and set of standards for a rapidly growing literature that evaluates FMs against human and developmental data. The paper's main contributions are conceptual: it separates task adaptation from linking hypotheses and evaluation, emphasizes diagnostic contrasts over aggregate fit, and insists that alignment is a relation between model, task, linking hypothesis, and theoretical claim rather than a property of the model alone. The treatment is careful and self-consciously hedged: the authors acknowledge multiple realizability, unfaithful chain-of-thought traces, and the correlational status of developmental correspondences. The paper also offers concrete, actionable guidelines and a running example that makes the abstract stages easy to follow. No new empirical validation is provided, but for a perspective article this is appropriate; the value lies in organizing and constraining future practice.

major comments (2)
  1. [Section 2, Stage 4; Guideline 4] Stage 4 is described as identifying computational features that are 'necessary' to reproduce a behavioral signature, and Guideline 4 repeats this language. Finite contrastive comparisons over a few architectures, scales, ablations, or training regimes can establish at most that a feature is necessary within that particular model family and manipulation set; because of multiple realizability, which the paper itself acknowledges in Section 4.2, they cannot establish that the feature is necessary for any computational account of the behavior, let alone that it corresponds to a human mechanism. I recommend rewording these passages to say 'necessary within the class of models and manipulations under consideration' and adding a sentence that unconditional mechanistic necessity would require additional theoretical constraints beyond contrastive evaluation. This is a local but important precision issue, because the abstract and conclusion also use 'necessary' when describing what the framework can illuminate.
  2. [Section 3, intro and Section 3.2] The paper reviews four linking hypotheses and explains their commitments, but it does not give the reader guidance for choosing among them beyond saying that appropriateness is 'empirical and task-dependent' in the surprisal section. Since the framework makes the linking hypothesis the central theoretical commitment, I would welcome a brief selection principle, for example, prefer the most proximal mapping that preserves the theoretical construct of interest, and triangulate across at least two linking hypotheses when possible. This would make the framework more actionable.
minor comments (4)
  1. [References] The reference list contains several formatting artifacts from LaTeX source, such as 'Y u', 'V arma', 'F orty-third', and 'ET AL .' in headings; these should be cleaned before publication.
  2. [Section 3.3] The term 'proximal' linking hypothesis is used without a definition; a brief gloss (e.g., mapping inputs and outputs directly rather than through internal representations) would help readers who are not familiar with the distal/proximal distinction.
  3. [Section 5, Guideline 2] The guideline notes that adaptation artifacts can lower observed alignment, but it is equally possible for surface cues to inflate alignment; adding this symmetric warning would make the point more complete.
  4. [Section 4.4] The discussion of persona-based prompting cites conflicting findings, but does not specify which findings conflict or how a reader should interpret them; one or two concrete examples would strengthen the caution.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper argues normatively for a four-stage framework using underdetermination and mechanistic opacity as premises, and its self-citations are illustrative rather than load-bearing.

full rationale

The paper is a perspective piece that proposes a methodological framework rather than deriving empirical predictions from fitted parameters or formal equations. Its central claim—that behavioral alignment is scientifically meaningful only when embedded in explicit theoretical commitments, theory-diagnostic tasks, and contrastive evaluation—is argued from conceptual premises: theoretical underdetermination (Section 4.1), mechanistic opacity and multiple realizability (Section 4.2), training/developmental mismatch (Section 4.3), and variability (Section 4.4). The framework is explicitly introduced as a systematization of prior proposals (“Our framework systematizes such proposals” and “we distill and expand guidance from prior commentaries”), which is a legitimate synthesis rather than a disguised derivation. The paper contains many self-citations (e.g., [36], [44], [45], [56], [88], [89], [102], [105], [28], [13], [15]), but they are used as examples of existing FM alignment studies or as background for domain-specific claims, not as premises that force the framework’s conclusions. No equation is equated to another by construction, no fitted parameter is relabeled as a prediction, and no uniqueness theorem or ansatz is imported from the authors’ prior work to forbid alternatives. The skeptical concern that Stage 4’s contrastive comparisons cannot fully establish necessity of computational features under multiple realizability is a substantive validity/cogency limitation of the framework, and the paper itself acknowledges it (“two systems might produce similar outputs while relying on different internal computations”); but this is not a circularity, because the framework’s claim does not reduce to its own definition or to a self-citation chain. The framework is a proposal about how to interpret alignment evidence, and its force depends on the strength of the methodological argument, not on an input-output equivalence. Therefore no circular step is identifiable, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities are introduced. The framework rests on three domain assumptions about the possibility of valid task adaptation, linking hypotheses, and contrastive mechanistic inference.

assumptions (3)
  • domain assumption Model outputs can be meaningfully mapped to human behavioral measures through a linking hypothesis.
    The framework's Stages 2 and 3 presuppose that some mapping between model probabilities, representations, or traces and human measures can be valid. The paper itself notes that this validity is task-dependent and empirical.
  • domain assumption Task adaptations can preserve the diagnostic structure of the original human experiment.
    Stage 1 guidance assumes that translating a task into a model-compatible format can retain the relational dependencies and distractor structure that make the original task informative. The paper acknowledges formatting artifacts can break this.
  • domain assumption Contrastive comparisons and mechanistic interpretability can in principle identify which computational features are necessary for a behavioral signature.
    Stage 4 relies on this to move from descriptive fit to explanation. Section 4.2 acknowledges multiple realizability and mechanistic opacity but treats them as surmountable through ablations and interpretability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the use of foundation models in cognitive science." pith.science (2026). https://pith.science/paper/6ZHX24EP

@misc{pith2026260807812,
  author       = {Pith},
  title        = {Pith review of: On the use of foundation models in cognitive science},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6ZHX24EP}},
  note         = {Machine review of arXiv:2608.07812}
}
read the original abstract

A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive domains, as well as whether aspects of model training track children's cognitive development. However, using FMs as candidate cognitive models poses significant methodological and conceptual challenges. A key question underlies this effort: under what conditions does behavioral alignment justify treating FMs as explanatory models of cognition? In this paper, we articulate a four-stage inferential framework for evaluating FMs as cognitive and developmental models: adapting human experimental tasks to model-compatible formats, specifying linking hypotheses that map model outputs to human measures, evaluating behavioral correspondence, and comparing across candidate models or manipulations. We clarify the role of linking hypotheses in mapping model outputs to human behavioral measures, identify challenges that constrain alignment claims, and propose principles for theory-driven and comparative evaluation. Throughout, we argue that behavioral fit alone is insufficient. Alignment becomes scientifically meaningful only when embedded within explicit theoretical commitments, theory-diagnostic tasks, and systematic contrastive evaluation across candidate models.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

112 extracted references · 64 canonical work pages

  1. [1]

    Using large language models to simulate multiple humans and replicate human subject studies

    Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai. Using large language models to simulate multiple humans and replicate human subject studies. In International Conference on Machine Learning , pages 337–371. PMLR, 2023

  2. [2]

    Large language models for mathematical reasoning: Progresses and challenges

    Janice Ahn, Rishu V erma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. Large language models for mathematical reasoning: Progresses and challenges. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages 225–237, 2024

  3. [3]

    How can the human mind occur in the physical universe? Oxford University Press, 2009

    John R Anderson. How can the human mind occur in the physical universe? Oxford University Press, 2009

  4. [4]

    Claude [large language model]

    Anthropic. Claude [large language model]. https://www.anthropic.com, 2026

  5. [5]

    Transformer networks of human conceptual knowledge

    Sudeep Bhatia and Russell Richie. Transformer networks of human conceptual knowledge. Psychological Review, 2022

  6. [6]

    Pythia: A suite for analyzing large language models across training and scaling

    Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle OBrien, Eric Hallahan, Moham- mad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning , pages 2397–2430. PMLR, 2023

  7. [7]

    Using cognitive psychology to understand gpt-3

    Marcel Binz and Eric Schulz. Using cognitive psychology to understand gpt-3. Proceedings of the National Academy of Sciences, 120(6):e2218523120, 2023

  8. [8]

    A foundation model to predict and capture human cognition

    Marcel Binz, Elif Akata, Matthias Bethge, Franziska Brändle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K Eckstein, Noémi Éltet ˝o, et al. A foundation model to predict and capture human cognition. Nature, pages 1–8, 2025. On the use of Foundation models in cognitive science 9

Show all 112 references
  1. [9]

    Using bayesian regression to test hypotheses about re- lationships between parameters and covariates in cognitive models

    Udo Boehm, Helen Steingroever, and Eric-Jan Wagenmakers. Using bayesian regression to test hypotheses about re- lationships between parameters and covariates in cognitive models. Behavior research methods , 50(3):1248–1269, 2018

  2. [10]

    Antagonistic ai

    Alice Cai, Ian Arawjo, and Elena L Glassman. Antagonistic ai. arXiv preprint arXiv:2402.07350, 2024

  3. [11]

    What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test

    Patricia A Carpenter, Marcel A Just, and Peter Shell. What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test. Psychological review, 97(3):404, 1990

  4. [12]

    Word acquisition in neural language models

    Tyler A Chang and Benjamin K Bergen. Word acquisition in neural language models. Transactions of the Association for Computational Linguistics, 10:1–16, 2022

  5. [13]

    Hu, Jing Liu, Jaap Jumelet, Tal Linzen, Aaron Mueller, Candace Ross, Raj Sanjay Shah, Alex Warstadt, Ethan Gotlieb Wilcox, and Adina Williams

    Lucas Charpentier, Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul, Michael Y . Hu, Jing Liu, Jaap Jumelet, Tal Linzen, Aaron Mueller, Candace Ross, Raj Sanjay Shah, Alex Warstadt, Ethan Gotlieb Wilcox, and Adina Williams. Findings of the third BabyLM challenge: Accelerating ...

  6. [14]

    Think deep, not just long: Measuring llm reasoning effort via deep-thinking tokens

    Wei-Lin Chen, Liqian Peng, Tian Tan, Chao Zhao, Jianhang Chen, Ziqian Lin, Alec Go, and Y u Meng. Think deep, not just long: Measuring llm reasoning effort via deep-thinking tokens. In F orty-third International Conference on Machine Learning, 2026

  7. [15]

    Babylm turns 4: Call for papers for the 2026 babylm workshop

    Leshem Choshen, Ryan Cotterell, Mustafa Omer Gul, Jaap Jumelet, Tal Linzen, Aaron Mueller, Suchir Salhan, Raj San- jay Shah, Alex Warstadt, and Ethan Gotlieb Wilcox. Babylm turns 4: Call for papers for the 2026 babylm workshop. arXiv preprint arXiv:2602.20092, 2026

  8. [16]

    Deep neural networks as scientific models

    Radoslaw M Cichy and Daniel Kaiser. Deep neural networks as scientific models. Trends in cognitive sciences , 23(4): 305–317, 2019

  9. [17]

    Cogbench: a large language model walks into a psychology lab

    Julian Coda-Forno, Marcel Binz, Jane X Wang, and Eric Schulz. Cogbench: a large language model walks into a psychology lab. arXiv preprint arXiv:2402.18225, 2024

  10. [18]

    Decoding answers before chain-of-thought: Evidence from pre-cot probes and activation steering

    Kyle Cox, Darius Kianersi, and Adrià Garriga-Alonso. Decoding answers before chain-of-thought: Evidence from pre-cot probes and activation steering. In Mechanistic Interpretability Workshop at ICML 2026 , 2026

  11. [19]

    Ai surrogates and illusions of generalizability in cognitive science

    MJ Crockett and Lisa Messeri. Ai surrogates and illusions of generalizability in cognitive science. Trends in Cognitive Sciences, 2025

  12. [20]

    The two disciplines of scientific psychology

    Lee Joseph Cronbach. The two disciplines of scientific psychology. American Psychologist, 12:671–684, 1957. URL https://api.semanticscholar.org/CorpusID:144287695

  13. [21]

    The limitations of large language models for understanding human language and cognition

    Christine Cuskley, Rebecca Woods, and Molly Flaherty. The limitations of large language models for understanding human language and cognition. Open Mind, 8:1058–1083, 2024

  14. [22]

    The cost of thinking is similar between large reasoning models and humans

    Andrea Gregor de V arda, Ferdinando Pio DElia, Hope Kean, Andrew Lampinen, and Evelina Fedorenko. The cost of thinking is similar between large reasoning models and humans. Proceedings of the National Academy of Sciences , 122 (47):e2520077122, 2025

  15. [23]

    Systematic testing of three language models reveals low language accuracy, absence of response stability, and a yes-response bias

    Vittoria Dentella, Fritz Günther, and Evelina Leivada. Systematic testing of three language models reveals low language accuracy, absence of response stability, and a yes-response bias. Proceedings of the National Academy of Sciences , 120 (51):e2309583120, 2023

  16. [24]

    Hlb: Benchmarking llms’ humanlikeness in language use

    Xufeng Duan, Bei Xiao, Xuemei Tang, and Zhenguang G Cai. Hlb: Benchmarking llms’ humanlikeness in language use. arXiv preprint arXiv:2409.15890, 2024

  17. [25]

    Linnea Evanson, Y air Lakretz, and Jean-Rémi King. Language acquisition: do children and language models follow similar learning stages? In Findings of the Association for Computational Linguistics: ACL 2023 , pages 12205–12218, 2023

  18. [26]

    Costa-jussà

    Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R. Costa-jussà. A Primer on the Inner Workings of Transformer-based Language Models, October 2024. URL http://arxiv.org/abs/2405.00208

  19. [27]

    A distributional perspective on word learning in neural language mod- els

    Filippo Ficarra, Ryan Cotterell, and Alex Warstadt. A distributional perspective on word learning in neural language mod- els. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technolo...

  20. [28]

    Bridging the data gap between children and large language models

    Michael C Frank. Bridging the data gap between children and large language models. Trends in Cognitive Sciences , 2023. 10 Shah ET AL

  21. [29]

    Cognitive modeling using artificial intelligence

    Michael C Frank and Noah D Goodman. Cognitive modeling using artificial intelligence. Annual Review of Psychology, 77, 2025

  22. [30]

    Individual differences in artificial neural networks capture individual differences in human behavior

    Herrick Fung, N Apurva Ratan Murty, and Dobromir Rahnev. Individual differences in artificial neural networks capture individual differences in human behavior. bioRxiv, pages 2026–02, 2026

  23. [31]

    Relational reasoning and generalization using nonsymbolic neural networks

    Atticus Geiger, Alexandra Carstensen, Michael C Frank, and Christopher Potts. Relational reasoning and generalization using nonsymbolic neural networks. Psychological Review, 130(2):308, 2023

  24. [32]

    What have we learned about artificial intelligence from studying the brain? Biological cybernetics, 118(1):1–5, 2024

    Samuel J Gershman. What have we learned about artificial intelligence from studying the brain? Biological cybernetics, 118(1):1–5, 2024

  25. [33]

    Gemini 3.1 pro

    Google. Gemini 3.1 pro. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/ ,

  26. [34]

    Olmo: Accelerating the science of language models

    Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. Olmo: Accelerating the science of language models. In Proceedings of the 62nd Annual Meeting of the Association for Computatio...

  27. [35]

    On logical inference over brains, behaviour, and artificial neural networks

    Olivia Guest and Andrea E Martin. On logical inference over brains, behaviour, and artificial neural networks. Computational Brain & Behavior , 6(2):213–227, 2023

  28. [36]

    Understanding graphical perception in data visualization through vision-language models

    Grace Guo, Jenna Kang, Raj Sanjay Shah, Hanspeter Pfister, and Sashank V arma. Understanding graphical perception in data visualization through vision-language models. In NeurIPS 2024 Workshop on Behavioral Machine Learning , 2024

  29. [37]

    Individual differences in non-verbal number acuity correlate with maths achievement

    Justin Halberda, Michèle MM Mazzocco, and Lisa Feigenson. Individual differences in non-verbal number acuity correlate with maths achievement. Nature, 455(7213):665–668, 2008

  30. [38]

    A probabilistic Earley parser as a psycholinguistic model

    John Hale. A probabilistic Earley parser as a psycholinguistic model. In Second Meeting of the North American Chapter of the Association for Computational Linguistics , 2001. URL https://aclanthology.org/N01-1021

  31. [39]

    Artificial neural network language models align neurally and behaviorally with humans even after a developmentally realistic amount of training

    Eghbal A Hosseini, Martin Schrimpf, Yian Zhang, Samuel Bowman, Noga Zaslavsky, and Evelina Fedorenko. Artificial neural network language models align neurally and behaviorally with humans even after a developmentally realistic amount of training. BioRxiv, pages 2022–10, 2022

  32. [40]

    Auxiliary task demands mask the capabilities of smaller language models

    Jennifer Hu and Michael Frank. Auxiliary task demands mask the capabilities of smaller language models. In First Conference on Language Modeling, 2024

  33. [41]

    Language models align with human judg- ments on key grammatical constructions

    Jennifer Hu, Kyle Mahowald, Gary Lupyan, Anna Ivanova, and Roger Levy. Language models align with human judg- ments on key grammatical constructions. Proceedings of the National Academy of Sciences , 121(36):e2400917121, 2024

  34. [42]

    In-context analogical reasoning with pre-trained language models

    Xiaoyang Hu, Shane Storks, Richard L Lewis, and Joyce Chai. In-context analogical reasoning with pre-trained language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), pages 1953–1969, 2023

  35. [43]

    thinking traces in large reasoning models: Cognitive cost or performative scaffolding? Proceedings of the National Academy of Sciences , 123(17):e2604554123, 2026

    Y ueqing Hu. thinking traces in large reasoning models: Cognitive cost or performative scaffolding? Proceedings of the National Academy of Sciences , 123(17):e2604554123, 2026

  36. [44]

    The representational geometry of number

    Zhimin Hu, Lanhao Niu, and Sashank V arma. The representational geometry of number. arXiv preprint arXiv:2602.06843, 2026

  37. [45]

    Are more tokens rational? inference-time scaling in language models as adaptive resource rationality

    Zhimin Hu, Riya Roshan, and Sashank V arma. Are more tokens rational? inference-time scaling in language models as adaptive resource rationality. arXiv preprint arXiv:2602.10329, 2026

  38. [46]

    How to evaluate the cognitive abilities of llms

    Anna A Ivanova. How to evaluate the cognitive abilities of llms. Nature Human Behaviour, pages 1–4, 2025

  39. [47]

    How many instructions can llms follow at once? In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025

    Daniel Jaroslawicz, Brendan Whiting, Parth Shah, and Karime Maamari. How many instructions can llms follow at once? In NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling, 2025

  40. [48]

    Llms can find mathematical reasoning mis- takes by pedagogical chain-of-thought

    Zhuoxuan Jiang, Haoyuan Peng, Shanshan Feng, Fan Li, and Dongsheng Li. Llms can find mathematical reasoning mis- takes by pedagogical chain-of-thought. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, pages 3439–3447, 2024

  41. [49]

    The organization of thinking: What functional brain imaging reveals about the neuroarchitecture of complex cognition

    Marcel Adam Just and Sashank V arma. The organization of thinking: What functional brain imaging reveals about the neuroarchitecture of complex cognition. Cognitive, Affective, & Behavioral Neuroscience, 7(3):153–191, 2007

  42. [50]

    Interpretability of artificial neural network models in artificial intelligence versus neuroscience

    Kohitij Kar, Simon Kornblith, and Evelina Fedorenko. Interpretability of artificial neural network models in artificial intelligence versus neuroscience. Nature Machine Intelligence, 4(12):1065–1067, 2022

  43. [51]

    Psychometric predictive power of large language models

    Tatsuki Kuribayashi, Y ohei Oseki, and Timothy Baldwin. Psychometric predictive power of large language models. In Findings of the Association for Computational Linguistics: NAACL 2024 , pages 1983–2005, 2024. On the use of Foundation models in cognitive science 11

  44. [52]

    Large language models are human-like internally

    Tatsuki Kuribayashi, Y ohei Oseki, Souhaib Ben Taieb, Kentaro Inui, and Timothy Baldwin. Large language models are human-like internally. Transactions of the Association for Computational Linguistics , 13:1743–1766, 2025

  45. [53]

    Language models, like humans, show content effects on reasoning tasks

    Andrew K Lampinen, Ishita Dasgupta, Stephanie CY Chan, Hannah R Sheahan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill. Language models, like humans, show content effects on reasoning tasks. PNAS nexus, 3(7):pgae233, 2024

  46. [54]

    Representation biases: will we achieve complete understanding by analyzing representations? arXiv preprint arXiv:2507.22216, 2025

    Andrew Kyle Lampinen, Stephanie CY Chan, Y uxuan Li, and Katherine Hermann. Representation biases: will we achieve complete understanding by analyzing representations? arXiv preprint arXiv:2507.22216, 2025

  47. [55]

    Expectation-based syntactic comprehension

    Roger Levy. Expectation-based syntactic comprehension. Cognition, 106(3):1126–1177, 2008. ISSN 0010-0277. . URL https://www.sciencedirect.com/science/article/pii/S0010027707001436

  48. [56]

    Incremental comprehension of garden-path sentences by large language models: Semantic interpretation, syntactic re-analysis, and attention

    Andrew Li, Xianle Feng, Siddhant Narang, Austin Peng, Tianle Cai, Raj Sanjay Shah, and Sashank V arma. Incremental comprehension of garden-path sentences by large language models: Semantic interpretation, syntactic re-analysis, and attention. In Proceedings of the Annual Meeti...

  49. [57]

    Cogmath: Assessing llms’ authentic mathematical ability from a human cognitive perspective

    Jiayu Liu, Zhenya Huang, Wei Dai, Cheng Cheng, Jinze Wu, Jing Sha, Song Li, Qi Liu, Shijin Wang, and Enhong Chen. Cogmath: Assessing llms’ authentic mathematical ability from a human cognitive perspective. In F orty-second International Conference on Machine Learning , 2025

  50. [58]

    Llm360: Towards fully transparent open-source llms

    Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Y uqi Wang, Suqi Sun, Omkar Pangarkar, et al. Llm360: Towards fully transparent open-source llms. In First Conference on Language Modeling, 2023

  51. [59]

    Response times: Their role in inferring elementary mental organization

    R Duncan Luce. Response times: Their role in inferring elementary mental organization . Oxford University Press, 1991

  52. [60]

    Rethinking thinking tokens: Llms as improvement operators

    Lovish Madaan, Aniket Didolkar, Suchin Gururangan, John Quan, Ruan Silva, Ruslan Salakhutdinov, Manzil Zaheer, Sanjeev Arora, and Anirudh Goyal. Rethinking thinking tokens: Llms as improvement operators. arXiv preprint arXiv:2510.01123, 2025

  53. [61]

    Dissociating language and thought in large language models

    Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. Dissociating language and thought in large language models. Trends in Cognitive Sciences, 2024

  54. [62]

    Vision: A computational investigation into the human representation and processing of visual information

    David Marr. Vision: A computational investigation into the human representation and processing of visual information . MIT press, 2010

  55. [63]

    Cognitive assessment of language models

    Daniel McDuff, David Munday, Xin Liu, and Isaac Galatzer-Levy. Cognitive assessment of language models. In ICML 2024 Workshop on LLMs and Cognition , 2024

  56. [64]

    How can deep neural networks inform theory in psychological science? 2023

    Sam Whitman McGrath, Jacob Russin, Ellie Pavlick, and Roman Feiman. How can deep neural networks inform theory in psychological science? 2023

  57. [65]

    N-gram-like language models predict reading time best

    James A Michaelov and Roger P Levy. N-gram-like language models predict reading time best. arXiv preprint arXiv:2603.09872, 2026

  58. [66]

    Large language models are able to downplay their cognitive abilities to fit the persona they simulate

    Jiˇrí Miliˇcka, Anna Marklová, Klára V anSlambrouck, Eva Pospíšilová, Jana Šimsová, Samuel Harvan, and Ondˇrej Drobil. Large language models are able to downplay their cognitive abilities to fit the persona they simulate. Plos one , 19(3): e0298522, 2024

  59. [67]

    Interventionist methods for interpreting deep neural networks

    Raphaël Millière and Cameron Buckner. Interventionist methods for interpreting deep neural networks. In Neurocogni- tive F oundations of Mind. Routledge, 1st edition, 2025

  60. [68]

    Do language models learn typicality judgments from text? In Proceedings of the Annual Meeting of the Cognitive Science Society , volume 43, 2021

    Kanishka Misra, Allyson Ettinger, and Julia Rayz. Do language models learn typicality judgments from text? In Proceedings of the Annual Meeting of the Cognitive Science Society , volume 43, 2021

  61. [69]

    Moyer and Thomas K

    Robert S. Moyer and Thomas K. Landauer. Time required for judgements of numerical inequality. Nature, 215(5109): 1519–1520, 1967

  62. [70]

    To model human linguistic prediction, make llms less superhuman

    Byung-Doh Oh and Tal Linzen. To model human linguistic prediction, make llms less superhuman. arXiv preprint arXiv:2510.05141, 2025

  63. [71]

    Transformer-based language model surprisal predicts human reading times best with about two billion training tokens

    Byung-Doh Oh and William Schuler. Transformer-based language model surprisal predicts human reading times best with about two billion training tokens. In Findings of the association for computational linguistics: EMNLP 2023 , pages 1915–1921, 2023

  64. [72]

    Desmond Ong. Gpt-ology, computational models, silicon sampling: How should we think about llms in cognitive science? In Proceedings of the Annual Meeting of the Cognitive Science Society , volume 46, 2024

  65. [73]

    OpenAI. Gpt-5.6. https://openai.com/index/gpt-5-6/, 2026. Accessed: 2026-07-16

  66. [74]

    So- cial simulacra: Creating populated prototypes for social computing systems

    Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. So- cial simulacra: Creating populated prototypes for social computing systems. In Proceedings of the 35th Annual ACM 12 Shah ET AL . Symposium on User Interface Softwar...

  67. [75]

    Temporal aspects of digit and letter inequality judgments

    John M Parkman. Temporal aspects of digit and letter inequality judgments. Journal of experimental psychology, 91(2): 191, 1971

  68. [76]

    Mapping language models to grounded conceptual spaces

    Roma Patel and Ellie Pavlick. Mapping language models to grounded conceptual spaces. In International conference on learning representations, 2021

  69. [77]

    Modern language models refute chomskys approach to language

    Steven Piantadosi. Modern language models refute chomskys approach to language. Lingbuzz Preprint, lingbuzz, 7180, 2023

  70. [78]

    Why concepts are (probably) vectors

    Steven T Piantadosi, Dyana CY Muller, Joshua S Rule, Karthikeya Kaushik, Mark Gorenstein, Elena R Leib, and Emily Sanford. Why concepts are (probably) vectors. Trends in Cognitive Sciences, 2024

  71. [79]

    Conjectures and refutations: The growth of scientific knowledge

    Karl Popper. Conjectures and refutations: The growth of scientific knowledge . routledge, 2014

  72. [80]

    A controlled reevaluation of coreference resolution models

    Ian Porada, Xiyuan Zou, and Jackie Chi Kit Cheung. A controlled reevaluation of coreference resolution models. In Pro- ceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 256–263, 2024

  73. [81]

    Frank, and Gary Lupyan

    Eva Portelance, Y uguang Duan, Michael C. Frank, and Gary Lupyan. Predicting age of acquisition for children’s early vocabulary in five languages using language model surprisal. Cognitive science , 47 9:e13334, 2023. URL https://api. semanticscholar.org/CorpusID:261696384

  74. [82]

    Psychological predicates

    Hilary Putnam. Psychological predicates. In W. H. Capitan and D. D. Merrill, editors, Art, Mind, and Religion , pages 37–48. University of Pittsburgh Press, 1967

  75. [83]

    Does spatial cognition emerge in frontier models? In The Thirteenth International Conference on Learning Representations , 2024

    Santhosh Kumar Ramakrishnan, Erik Wijmans, Philipp Kraehenbuehl, and Vladlen Koltun. Does spatial cognition emerge in frontier models? In The Thirteenth International Conference on Learning Representations , 2024

  76. [84]

    Similarity judgment within and across categories: A comprehensive model comparison

    Russell Richie and Sudeep Bhatia. Similarity judgment within and across categories: A comprehensive model comparison. Cognitive science, 45(8):e13030, 2021

  77. [85]

    Perturbation: A simple and efficient adversarial tracer for representation learning in language models

    Joshua Rozner and Cory Shain. Perturbation: A simple and efficient adversarial tracer for representation learning in language models. arXiv preprint arXiv:2603.23821, 2026

  78. [86]

    Pdp models and general issues in cognitive science

    David E Rumelhart and James L McClelland. Pdp models and general issues in cognitive science. In Parallel distributed processing: Explorations in the microstructure of cognition, V ol. 1: F oundations, pages 110–146. 1986

  79. [87]

    In-context impersonation reveals large language models’ strengths and biases

    Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. In-context impersonation reveals large language models’ strengths and biases. Advances in Neural Information Processing Systems , 36, 2024

  80. [88]

    Development of cognitive intelligence in pre-trained language models

    Raj Shah, Khushi Bhardwaj, and Sashank V arma. Development of cognitive intelligence in pre-trained language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages 9632–9657, 2024

  81. [89]

    Numeric magnitude comparison effects in large language models, 2023

    Raj Sanjay Shah, Vijay Marupudi, Reba Koenen, Khushi Bhardwaj, and Sashank V arma. Numeric magnitude comparison effects in large language models, 2023

  82. [90]

    Word frequency and predictability dissociate in naturalistic reading

    Cory Shain. Word frequency and predictability dissociate in naturalistic reading. Open Mind, 8:177–201, 2024

  83. [91]

    Open problems in mechanistic interpretability

    Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeffrey Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Isaac Bloom, et al. Open problems in mechanistic interpretability. Transactions on Machine Learning Research, 2025

  84. [92]

    The bitter lesson

    Rich Sutton. The bitter lesson. 2019

  85. [93]

    Devbench: A multimodal developmental benchmark for language learning

    Alvin W Tan, Sunny Y u, Bria Long, Wanjing Anya, Tonya Murray, Rebecca D Silverman, Jason D Y eatman, and Michael C Frank. Devbench: A multimodal developmental benchmark for language learning. Advances in Neural Information Processing Systems, 37:77445–77467, 2024

  86. [94]

    Numerosity discrimination in deep neural networks: Initial competence, developmental refinement and experience statistics

    Alberto Testolin, Will Y Zou, and James L McClelland. Numerosity discrimination in deep neural networks: Initial competence, developmental refinement and experience statistics. Developmental science, 23(5):e12940, 2020

  87. [95]

    Two tales of persona in llms: A survey of role-playing and personalization

    Y u-Min Tseng, Y u-Chao Huang, Teng-Y un Hsiao, Wei-Lin Chen, Chao-Wei Huang, Y u Meng, and Y un-Nung Chen. Two tales of persona in llms: A survey of role-playing and personalization. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 16612–16631, 2024

  88. [96]

    Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting

    Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems , 36:74952– 74965, 2023

  89. [97]

    Large language models fail on trivial alterations to theory-of-mind tasks

    Tomer Ullman. Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399, 2023. On the use of Foundation models in cognitive science 13

  90. [98]

    Individual differences as a crucible in theory construction

    Benton J Underwood. Individual differences as a crucible in theory construction. American Psychologist , 30(2):128, 1975

  91. [99]

    Reclaiming ai as a theoretical tool for cognitive science

    Iris V an Rooij, Olivia Guest, Federico Adolfi, Ronald De Haan, Antonina Kolokolova, and Patricia Rich. Reclaiming ai as a theoretical tool for cognitive science. Computational Brain & Behavior , 7(4):616–636, 2024

  92. [100]

    Correlations without causa- tion do not support claims of human–llm reasoning alignment

    Ivan I V ankov, Federico Adolfi, Rachel F Heaton, Guillermo Puebla, and Jeffrey S Bowers. Correlations without causa- tion do not support claims of human–llm reasoning alignment. Proceedings of the National Academy of Sciences , 123 (12):e2536362123, 2026

  93. [101]

    Capturing human cognitive styles with language: Towards an experimental evaluation paradigm

    V asudha V aradarajan, Syeda Mahwish, Xiaoran Liu, Julia Buffolino, Christian C Luhmann, Ryan Boyd, and H Andrew Schwartz. Capturing human cognitive styles with language: Towards an experimental evaluation paradigm. In Proceed- ings of the 2025 Conference of the Nations of the...

  94. [102]

    How well do deep learning models capture human concepts? the case of the typicality effect

    Siddhartha V emuri, Raj Sanjay Shah, and Sashank V arma. How well do deep learning models capture human concepts? the case of the typicality effect. In Proceedings of the Annual Meeting of the Cognitive Science Society , volume 46, 2024

  95. [103]

    Is a picture worth a thousand words? delving into spatial reasoning for vision language models

    Jiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet, Xin Wang, Sharon Li, and Neel Joshi. Is a picture worth a thousand words? delving into spatial reasoning for vision language models. Advances in Neural Information Processing Systems, 37:75392–75421, 2024

  96. [104]

    Coglm: Tracking cognitive development of large language models

    Xinglin Wang, Peiwen Y uan, Shaoxiong Feng, Yiwei Li, Boyuan Pan, Heda Wang, Y ao Hu, and Kan Li. Coglm: Tracking cognitive development of large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational L...

  97. [105]

    What artificial neural networks can tell us about human language ac- quisition

    Alex Warstadt and Samuel R Bowman. What artificial neural networks can tell us about human language ac- quisition. In Shalom Lappin and Jean-Philippe Bernardy, editors, Algebraic Structures in Natural Language , pages 17–60. CRC Press, 2022. URL https://www.taylorfrancis.com/ch...

  98. [106]

    Blimp: The benchmark of linguistic minimal pairs for english

    Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. Blimp: The benchmark of linguistic minimal pairs for english. Transactions of the Association for Computational Linguistics, 8:377–392, 2020

  99. [107]

    Emergent analogical reasoning in large language models

    Taylor Webb, Keith J Holyoak, and Hongjing Lu. Emergent analogical reasoning in large language models. Nature Human Behaviour, 7(9):1526–1541, 2023

  100. [108]

    Evidence from counterfactual tasks supports emergent analogical reasoning in large language models

    Taylor W Webb, Keith J Holyoak, and Hongjing Lu. Evidence from counterfactual tasks supports emergent analogical reasoning in large language models. PNAS nexus, 4(5):pgaf135, 2025

  101. [109]

    On the need to improve the way individual differences in cognitive function are measured with reaction time tasks

    Corey N White and Kiah N Kitchen. On the need to improve the way individual differences in cognitive function are measured with reaction time tasks. Current directions in psychological science , 31(3):223–230, 2022

  102. [110]

    In- context learning strategies emerge rationally

    Daniel Wurgaft, Ekdeep Singh Lubana, Core Francisco Park, Hidenori Tanaka, Gautam Reddy, and Noah Goodman. In- context learning strategies emerge rationally. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025

  103. [111]

    Choosing prediction over explanation in psychology: Lessons from machine learning

    Tal Y arkoni and Jacob Westfall. Choosing prediction over explanation in psychology: Lessons from machine learning. Perspectives on Psychological Science, 12(6):1100–1122, 2017

  104. [2026]

    Accessed: 2026-05-09

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.