REVIEW 3 major objections 4 minor 3 cited by
This review argues that large language models can address cognitive science's chronic problems—fragmented literatures, vague theories, construct redundancy, narrow models, and missing context—when deployed as tools that complement rather th
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 06:51 UTC pith:LFD4JQZT
load-bearing objection A balanced, well-scoped review that frames LLMs as tools for cognitive science; the 'can help' claim is programmatic and honestly hedged, but the evidence base is thin in places. the 3 major comments →
Addressing Longstanding Challenges in Cognitive Science with Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is programmatic: large language models can function as instruments for a more integrative, cumulative cognitive science. Five chronic challenges are identified—fragmentation, insufficient formalization, conceptual and measurement confusion, lack of generalizability, and neglect of context—and for each the paper assembles evidence that LLMs already provide useful traction. LLM embeddings of article titles and abstracts produce research maps that expose thematic clusters, temporal development, and cross-field bridges; LLM-assisted pipelines translate verbal theories into executable code and even generate new computational models that can match or exceed established on
What carries the argument
The load-bearing mechanism is the semantic embedding: texts, questionnaire items, construct labels, and task descriptions are converted into numerical vectors in a shared high-dimensional space where proximity encodes similarity in meaning. That single representation carries most of the argument—it is what lets research maps reveal conceptual links, lets measurement taxonomies flag jingle–jangle fallacies, and lets a foundation model such as Centaur (a transformer trained on behavioral datasets to predict choices across tasks) treat disparate experimental paradigms as tokens in one language. Around this core, the paper assembles two further mechanisms: LLM generation pipelines that convert v
Load-bearing premise
The entire proposal depends on the assumption that the semantic structure LLMs learn from text faithfully tracks the conceptual and behavioral structure of the mind—an assumption the paper itself flags as an open validation question.
What would settle it
Take a preregistered set of established personality scales and ask whether LLM embeddings of their items recover the scales' known factor structure out-of-sample; if the embeddings systematically place items from different constructs closer than items from the same construct, the measurement-taxonomy claim is falsified.
If this is right
- Research maps built from LLM embeddings can expose conceptual and methodological links between subfields that citation-based tools miss, giving researchers a practical route out of disciplinary silos.
- LLM-assisted formalization can turn vague verbal theories into executable models with testable predictions; at least one reported pipeline produced models that matched or outperformed established computational models across decision making, learning, planning, and working memory.
- Embedding-based measurement taxonomies can detect redundant constructs and relabel or eliminate them; the paper reports a demonstration that condensed a hypothesized personality construct set by roughly 75 percent.
- Multitask foundation models such as Centaur offer a platform for prediction across many cognitive tasks and generalization to unseen task structures, a direct alternative to one-model-per-phenomenon research.
- LLM analysis of naturalistic data can bring ecological, cultural, and individual variation into cognitive modeling—for instance, by extracting decision attributes and trade-offs from over 100,000 real-world choice dilemmas.
Where Pith is reading between the lines
- If embedding-based mapping matures, the binding constraint shifts from discovering redundant constructs to governing how they get consolidated; the same tools that expose overlap could be used to run transparent, versioned conceptual-revision processes rather than one-off pruning exercises.
- The evidence that text-trained LLMs predict behavior on nonverbal tasks suggests a testable conjecture the paper does not state: a large share of human task behavior is driven by linguistically accessible task features. Comparing LLM predictions on verbal versus purely visuospatial task variants would bound that share.
- The paper's open-infrastructure argument implies a concrete discipline-level experiment: teams using open-weight models should, in aggregate, produce more independently replicable measurement taxonomies and predictions than teams using closed APIs, because contamination and version drift can be audited.
- The paper's contrasting dystopian and utopian futures imply that the real bottleneck is institutional, not technical; a testable extension would track whether adoption of LLM tools by cognitive scientists is followed by more cross-disciplinary citations and fewer redundant constructs over the next decade.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This perspective/review argues that large language models (LLMs) can help cognitive science address five longstanding challenges: disciplinary silos, insufficient formalization, conceptual and measurement confusion, lack of generalizability, and neglect of ecological context. The authors organize the paper around these five areas, presenting illustrative applications (research maps, theory formalization, measurement taxonomies, integrated frameworks, contextualized representations), followed by a discussion of pitfalls (opacity, oversimplification, bias, data contamination, deskilling) and a call for open infrastructures and human oversight. The central claim is explicitly hedged: LLMs should complement rather than replace human expertise. The paper is a programmatic review rather than a new empirical study.
Significance. If the programmatic claim is accepted, the paper provides a useful organizational frame for a rapidly growing literature at the intersection of LLMs and cognitive science. Its strengths include a balanced treatment of risks, concrete examples for each proposed application area, and a constructive emphasis on open-weight models, interpretability, and validation as scientific priorities. The paper also explicitly lists outstanding questions, which is helpful for agenda-setting. However, the positive evidence is drawn from a small set of demonstrations, several of which are non-peer-reviewed preprints or the authors' own work, and the foundational assumption that embedding similarity tracks psychological/construct similarity is acknowledged but not critically examined in depth. The paper's value is therefore more synthesizing and cautionary than evidential.
major comments (3)
- [Section 2.3 / Table 1, row 3 / Figure 3] The measurement-taxonomy argument is load-bearing for the claim that LLMs can reduce conceptual and measurement confusion. However, the section treats semantic embedding proximity as a proxy for construct and measure equivalence without providing a benchmark or validation protocol. The Wulff and Mata [19] demonstration is described only as a 'sketch' that reduces a construct set by roughly 75%, and the paper defers validation to Box 3, Q1. Without evidence that these relabelings correspond to convergent/discriminant validity, improved predictive utility, or expert agreement, the claim risks reducing to the truism that LLMs can manipulate text. I recommend either tempering the wording to 'exploratory' or adding a concrete discussion of validation criteria and known failure cases.
- [Section 2.1 / Box 1 / Figure 2] The 'theory of mind' research map is presented as an example of how LLMs can map fragmented literatures, but no evidence is provided for the map's accuracy, stability, or interpretability. The historical narrative (e.g., the field 'originated from clusters in autism and child development') is derived from a cluster projection whose validity is not assessed. The manual labeling of clusters based on author keywords is a subjective step. For the first challenge in Table 1 to be convincing, the paper should at least note that such maps are exploratory hypotheses generators, not validated descriptions of the field, and ideally cite any reliability checks or human-evaluation results.
- [Sections 2.1–2.3 and reference list] Several of the central supporting examples are non-peer-reviewed preprints or come from the authors' own research program, including [12], [16], [49], [19], [63], [97], and [98]. This does not invalidate the review, but it is a load-bearing limitation because the paper's 'can help' claim rests on these selected demonstrations. I ask the authors to explicitly flag preprint status, distinguish their own work from independent replications, and note where independent verification is still needed. Without such signaling, a skeptical reader cannot separate established results from programmatic ambitions.
minor comments (4)
- [Box 3, Q1 and Q5] Q1 refers to 'cf. Box 2.1', but no such box exists; this should be 'Box 1'. Q5 refers to 'cf. Box 3', but it should refer to 'Box 2' (the box on open LLMs).
- [Figure 1 caption and Glossary] The words 'F ormal' and 'F ormal Models' appear with an unusual space. This looks like a typographical artifact and should be corrected.
- [Section 2.5] 'in-silico testing' should be 'in silico testing' (italicized, no hyphen), per standard usage.
- [Box 1 / Figure 2 caption] The caption says 'For more details see Thoma et al. [12]', but [12] is about behavioral reinforcement learning, not theory of mind. Either the theory-of-mind map is generated using the same method as [12], or a different reference/companion document is needed. Please clarify.
Circularity Check
No significant circularity: the paper is an argumentative review that relies on externally benchmarked examples and explicitly defers validation rather than presenting fitted outputs as predictions.
full rationale
The paper does not contain a derivation chain that reduces a claimed result to its inputs. It is a perspective/review whose central claim, that LLMs can serve as tools for a more integrative and cumulative cognitive science, is supported by a mix of independent external benchmarks and published applications. The self-citations (e.g., Wulff & Mata [19], [63]; Fulawka et al. [49]) are illustrative examples, not the sole or definitional basis for the claim. The measurement-taxonomy example is anchored to an external criterion: the paper states that the semantic representations 'reproduced empirical item-scale relationships well,' which is a benchmark outside the embedding construction, not a circular fit. The paper also explicitly flags the need for validation: Section 2.2 says 'these efforts require validation,' and Box 3 asks 'How can we best validate the accuracy and completeness of knowledge structures... automatically extracted by large language models?' Section 3 further acknowledges that LLMs 'may rely on statistical shortcuts or reflect biases in training data.' No equation or claimed predictive result is shown to be equivalent to a fitted parameter, and no uniqueness theorem or ansatz is imported from the authors' prior work. The presence of self-citation alone is not circularity, especially where the cited works are independently published and externally falsifiable. Thus no specific circular step can be exhibited, and the appropriate score is 0.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Semantic embeddings from LLMs capture meaningful psychological and conceptual structure.
- domain assumption LLM performance on textual tasks transfers to behavioral and cognitive prediction.
- domain assumption Literature structure inferred by embeddings is a faithful map of scientific knowledge.
- domain assumption Human researchers can effectively supervise and correct LLM outputs.
read the original abstract
Cognitive science faces ongoing challenges in research integration, formalization, conceptual clarity, and other areas, in part due to its multifaceted and interdisciplinary nature. Recent advances in artificial intelligence, particularly the development of language models, offer tools that may help to address these longstanding issues. Specifically, they can help map fragmented literatures, formalize verbal theories, identify overlap among constructs and measures, generate predictions across tasks, and extract cultural or ecological structure from naturalistic data. However, these opportunities come with risks, including oversimplification, opacity, deskilling, and bias. Taken together, we conclude that language models could serve as tools for a more integrative and cumulative cognitive science when used judiciously to complement, rather than replace, human agency.
Forward citations
Cited by 3 Pith papers
-
Psychological Competence as a Missing Dimension in AI Evaluation
Psychological competence—the capacity of human-facing AI to support appropriate user cognition, emotion, and decision-making—should be evaluated as a distinct interaction-level dimension beyond technical and safety be...
-
Psychological Competence as a Missing Dimension in AI Evaluation
AI evaluation should add 'psychological competence' — how well a system supports user reasoning, emotional stability, and autonomous decision-making — as a core dimension.
-
Taming the Centaur(s) with LAPITHS: a framework for a theoretically grounded interpretation of AI performances
LAPITHS shows that major claims for CENTAUR as a unified model of cognition lack theoretical or empirical support because similar results can be obtained from systems without the structural features associated with co...
Reference graph
Works this paper leans on
-
[1]
The mind’s new science: A history of the cognitive revolution
Gardner H. The mind’s new science: A history of the cognitive revolution. New York, USA: Basic Books; 1985
1985
-
[2]
What happened to cognitive science? Nature Human Behaviour
N´ u˜ nez R, Allen M, Gao R, Miller Rigoli C, Relaford-Doyle J, Semenuks A. What happened to cognitive science? Nature Human Behaviour. 2019;3(8):782–791. https://doi.org/10.1038/s41562-019-0626-2
-
[3]
Formalizing verbal theories: A tutorial by dialogue
Van Rooij I, Blokpoel M. Formalizing verbal theories: A tutorial by dialogue. Social Psychology. 2020;51(5):285–298. https://doi.org/10.1027/1864-9335/ a000428
-
[4]
How computational modeling can force theory building in psychological science
Guest O, Martin AE. How computational modeling can force theory building in psychological science. Perspectives on Psychological Science. 2021;16(4):789–802. https://doi.org/10.1177/1745691620970585
-
[5]
Defragmenting psychology
Anvari F, Alsalti T, Oehler LA, Hussey I, Elson M, Arslan RC. Defragmenting psychology. Nature Human Behaviour. 2025;9:836–839. https://doi.org/10.1038/ s41562-025-02138-0
2025
-
[6]
You can’t play 20 questions with nature and win: Projective comments on the papers of this symposium
Newell A. You can’t play 20 questions with nature and win: Projective comments on the papers of this symposium. In: Chase WG, editor. Visual Information Processing. New York, USA: Academic Press (Elsevier); 1973. p. 283–308
1973
-
[7]
An integrated theory of the mind
Anderson JR, Bothell D, Byrne MD, Douglass S, Lebiere C, Qin Y. An integrated theory of the mind. Psychological Review. 2004;111(4):1036–1060. https://doi. org/10.1037/0033-295X.111.4.1036
-
[8]
The weirdest people in the world? Behavioral and Brain Sciences
Henrich J, Heine SJ, Norenzayan A. The weirdest people in the world? Behavioral and Brain Sciences. 2010;33(2-3):61 – 83. https://doi.org/10.1017/ S0140525X0999152X
2010
-
[9]
Participant diver- sity is necessary to advance brain aging research
Wig GS, Klausner S, Chan MY, Sullins C, Rayanki A, Seale M. Participant diver- sity is necessary to advance brain aging research. Trends in Cognitive Sciences. 2024;28(2):92–96. https://doi.org/10.1016/j.tics.2023.12.004
-
[10]
Parallel distributed processing: explorations in the microstructure of cognition, vol
Rumelhart DE, McClelland JL, PDP Research Group C, editors. Parallel distributed processing: explorations in the microstructure of cognition, vol. 1: foundations. Cambridge, MA, USA: MIT Press; 1986
1986
-
[11]
The two disciplines of scientific psychology
Cronbach LJ. The two disciplines of scientific psychology. The American Psychol- ogist. 1957;12(11):671–684. https://doi.org/https://doi.org/10.1037/h0043943
-
[12]
Map- ping the landscape of behavioral reinforcement learning research
Thoma AI, Bolenz F, Tiede K, Yang Y, Palminteri S, Hertwig R, et al. Map- ping the landscape of behavioral reinforcement learning research. PsyArXiv. 2025;https://doi.org/10.31234/osf.io/6c2va v1. 18
-
[13]
Large language models surpass human experts in predicting neuroscience results
Luo X, Rechardt A, Sun G, Nejad KK, Y´ a˜ nez F, Yilmaz B, et al. Large language models surpass human experts in predicting neuroscience results. Nature Human Behaviour. 2024;9(2):305––315. https://doi.org/10.1038/s41562-024-02046-9
-
[14]
Addressing the theory crisis in psychology
Oberauer K, Lewandowsky S. Addressing the theory crisis in psychology. Psy- chonomic Bulletin & Review. 2019;26(5):1596–1618. https://doi.org/10.3758/ s13423-019-01645-2
2019
-
[15]
How to translate a verbal theory into a formal model
Smaldino PE. How to translate a verbal theory into a formal model. Social Psychology. 2020;51(4):207–218. https://doi.org/10.1027/1864-9335/a000425
-
[16]
Generating com- putational cognitive models using large language models
Rmus M, Jagadish AK, Mathony M, Ludwig T, Schulz E. Generating com- putational cognitive models using large language models. arXiv. 2025;https: //doi.org/10.48550/arXiv.2502.00879
-
[17]
theoraizer: AI-assisted theory construction
Waaijers M, Rosenbusch H, Van Lissa CJ, Roefs A, Borsboom D. theoraizer: AI-assisted theory construction. PsyArXiv. 2024 Aug;OSF. https://doi.org/10. 31234/osf.io/gu9yq
2024
-
[18]
From brain maps to cognitive ontologies: Infor- matics and the search for mental structure
Poldrack RA, Yarkoni T. From brain maps to cognitive ontologies: Infor- matics and the search for mental structure. Annual Review of Psychology. 2016;67(1):587–612. https://doi.org/10.1146/annurev-psych-122414-033729
-
[19]
Semantic embeddings reveal and address taxonomic incommensurability in psychological measurement
Wulff DU, Mata R. Semantic embeddings reveal and address taxonomic incommensurability in psychological measurement. Nature Human Behaviour. 2025;9(5):944–954. https://doi.org/10.1038/s41562-024-02089-y
-
[20]
A foundation model to predict and capture human cognition
Binz M, Akata E, Bethge M, Br¨ andle F, Callaway F, Coda-Forno J, et al. A foundation model to predict and capture human cognition. Nature. 2025;644(8078):1002–1009. https://doi.org/10.1038/s41586-025-09215-4
-
[21]
Towards a cognitive science of the human: Cross-cultural approaches and their urgency
Barrett HC. Towards a cognitive science of the human: Cross-cultural approaches and their urgency. Trends in Cognitive Sciences. 2020;24(8):620–638. https:// doi.org/10.1016/j.tics.2020.05.007
-
[22]
Building Knowledge-Guided Lexica to Model Cultural Variation
Havaldar S, Giorgi S, Rai S, Talhelm T, Guntuku SC, Ungar L. Building Knowledge-Guided Lexica to Model Cultural Variation. In: Duh K, Gomez H, Bethard S, editors. Proceedings of the 2024 Conference of the North Ameri- can Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Mexico City, Mexico: Ass...
2024
-
[23]
On the conversational persuasiveness of GPT-4
Salvi F. On the conversational persuasiveness of GPT-4. Nature Human Behaviour. 2025;9:1645–1653. https://doi.org/10.1038/s41562-025-02194-6. 19
-
[24]
The two cultures and the scientific revolution
Snow CP. The two cultures and the scientific revolution. New York, USA: Cambridge University Press; 1959
1959
-
[25]
Ending the reading wars: Reading acquisition from novice to expert
Castles A, Rastle K, Nation K. Ending the reading wars: Reading acquisition from novice to expert. Psychological Science in the Public Interest. 2018;19(1):5–51. https://doi.org/10.1177/1529100618772271
-
[26]
Risk preference shares the psychometric structure of major psychological traits
Frey R, Pedroni A, Mata R, Rieskamp J, Hertwig R. Risk preference shares the psychometric structure of major psychological traits. Scientific Advances. 2017;3(10):e1701381. https://doi.org/10.1126/sciadv.1701381
-
[27]
Risk preference: A view from psychology
Mata R, Frey R, Richter D, Schupp J, Hertwig R. Risk preference: A view from psychology. Journal of Economic Perspectives;32(2):155–172. https://doi.org/10. 1257/jep.32.2.155
-
[28]
Three gaps and what they may mean for risk preference
Hertwig R, Wulff DU, Mata R. Three gaps and what they may mean for risk preference. Philosophical Transactions of the Royal Society Series B, Biological Sciences. 2019;374(1766):20180140. https://doi.org/10.1098/rstb.2018.0140
arXiv 2019
-
[29]
In defense of disciplines: Interdisciplinarity and specialization in the research university
Jacobs JA. In defense of disciplines: Interdisciplinarity and specialization in the research university. Chicago, IL: University of Chicago Press; 2014
2014
-
[30]
Binz M, Alaniz S, Roskies A, Aczel B, Bergstrom CT, Allen C, et al. How should the advancement of large language models affect the practice of science? Proceedings of the National Academy of Sciences of the United States of America. 2025;122(5):e2401227121. https://doi.org/10.1073/pnas.2401227121
-
[31]
Eger S, Cao Y, D’Souza J, Geiger A, Greisinger C, Gross S, et al. Transforming science with large language models: A survey on AI-assisted scientific discovery, experimentation, content generation, and evaluation. arXiv. 2025;https://doi. org/10.48550/arXiv.2502.05151
-
[32]
Marshall IJ, Wallace BC. Toward systematic review automation: A practical guide to using machine learning tools in research synthesis. Systematic Reviews. 2019 Dec;8(1):163. https://doi.org/10.1186/s13643-019-1074-9
-
[33]
A digital future for the history of psychology? History of Psychology
Green CD. A digital future for the history of psychology? History of Psychology. 2016;19(3):209–219. https://doi.org/10.1037/hop0000012
-
[34]
LLMs4Synthesis: Leveraging large language models for scientific synthesis
Babaei Giglou H, D’Souza J, Auer S. LLMs4Synthesis: Leveraging large language models for scientific synthesis. Proceedings of the 24th ACM/IEEE Joint Con- ference on Digital Libraries. 2024 Dec;p. 1–12. https://doi.org/10.1145/3677389. 3702565
-
[35]
Exploring the applicability of large language models to citation context analysis
Nishikawa K, Koshiba H. Exploring the applicability of large language models to citation context analysis. Scientometrics. 2024;129:6751–6777. https://doi.org/ 10.1007/s11192-024-05142-9. 20
-
[36]
Identifying interdisciplinary emergence in the science of science: combination of network analysis and BERTopic
Kim K, Kogler DF, Maliphol S. Identifying interdisciplinary emergence in the science of science: combination of network analysis and BERTopic. Humanities and Social Sciences Communications. 2024;11(1):1–15. https://doi.org/10.1057/ s41599-024-03044-y
2024
-
[37]
An embedding approach for ana- lyzing the evolution of research topics with a case study on computer science subdomains
Taher Harikandeh SR, Aliakbary S, Taheri S. An embedding approach for ana- lyzing the evolution of research topics with a case study on computer science subdomains. Scientometrics. 2023;128(3):1567–1582. https://doi.org/10.1007/ s11192-023-04642-4
2023
-
[38]
Using sequences of life-events to predict human lives
Savcisens G, Eliassi-Rad T, Hansen LK, Mortensen LH, Lilleholt L, Rogers A, et al. Using sequences of life-events to predict human lives. Nature Computational Science. 2023 Dec;4(1):43–56. https://doi.org/10.1038/s43588-023-00573-5
-
[39]
Automating the practice of science: Opportunities, challenges, and impli- cations
Musslick S, Bartlett LK, Chandramouli SH, Dubova M, Gobet F, Griffiths TL, et al. Automating the practice of science: Opportunities, challenges, and impli- cations. Proceedings of the National Academy of Sciences of the United States of America. 2025;122(5):e2401238121. https://doi.org/10.1073/pnas.2401238121
-
[40]
Minimal coherence among varied theory of mind measures in childhood and adulthood
Warnell KR, Redcay E. Minimal coherence among varied theory of mind measures in childhood and adulthood. Cognition. 2019;191:103997. https://doi.org/10. 1016/j.cognition.2019.06.009
2019
-
[41]
Deconstructing and recon- structing theory of mind
Schaafsma SM, Pfaff DW, Spunt RP, Adolphs R. Deconstructing and recon- structing theory of mind. Trends in Cognitive Sciences. 2015;19(2):65–72. https: //doi.org/10.1016/j.tics.2014.11.007
-
[42]
Qwen3 Embedding: Advancing text embedding and reranking through foundation models
Zhang Y, Li M, Long D, Zhang X, Lin H, Yang B, et al. Qwen3 Embedding: Advancing text embedding and reranking through foundation models. arXiv. 2025;https://doi.org/10.48550/arXiv.2506.05176
-
[43]
Understanding how dimension reduc- tion tools work: An empirical approach to deciphering t-SNE, UMAP, TriMap, and PaCMAP for data visualization
Wang Y, Huang H, Rudin C, Shaposhnik Y. Understanding how dimension reduc- tion tools work: An empirical approach to deciphering t-SNE, UMAP, TriMap, and PaCMAP for data visualization. Journal of Machine Learning Research. 2021;22(201):1–73. https://doi.org/http://jmlr.org/papers/v22/20-1061.html
2021
-
[44]
Cognitive science is and should be pluralistic
Gentner D. Cognitive science is and should be pluralistic. Topics in Cognitive Science. 2019 Oct;11(4):884–891. https://doi.org/10.1111/tops.12459
-
[45]
Why summaries of research on psychological theories are often unin- terpretable
Meehl PE. Why summaries of research on psychological theories are often unin- terpretable. Psychological Reports. 1990;66(1):195–244. https://doi.org/10.2466/ PR0.66.1.195-244
1990
-
[46]
Michie S, Thomas J, Johnston M, Aonghusa PM, Shawe-Taylor J, Kelly MP, et al. The Human Behaviour-Change Project: Harnessing the power of artificial intelligence and machine learning for evidence synthesis and interpretation. Imple- mentation Science. 2017;12(1):121. https://doi.org/10.1186/s13012-017-0641-5. 21
-
[47]
Computational Models in Personality and Social Psychol- ogy
Read SJ, Monroe BM. Computational Models in Personality and Social Psychol- ogy. In: Sun R, editor. The Cambridge Handbook of Computational Cognitive Sciences. Cambridge, UK: Cambridge University Press; 2023. p. 795–835
2023
-
[48]
Using large-scale experiments and machine learning to discover theories of human decision-making
Peterson JC, Bourgin DD, Agrawal M, Reichman D, Griffiths TL. Using large-scale experiments and machine learning to discover theories of human decision-making. Science. 2021;372(6547):1209–1214. https://doi.org/10.1126/ science.abe2629
2021
-
[49]
Large language models accurately identify decision reasons in verbal reports
Fulawka K, Hertwig R, Wulff DU. Large language models accurately identify decision reasons in verbal reports. PsyArXiv. 2025;https://doi.org/10.31234/osf. io/yuzmw v1
doi:10.31234/osf 2025
-
[50]
Psychological measures aren’t tooth- brushes
Elson M, Hussey I, Alsalti T, Arslan RC. Psychological measures aren’t tooth- brushes. Communications Psychology. 2023;1(1):25. https://doi.org/10.1038/ s44271-023-00026-9
2023
-
[51]
The theory crisis in psychology: How to move forward
Eronen MI, Bringmann LF. The theory crisis in psychology: How to move forward. Perspectives on Psychological Science. 2021;16(4):779–788. https://doi.org/10. 1177/1745691620970586
2021
-
[52]
What is conceptual engineering and what should it be? Inquiry
Chalmers DJ. What is conceptual engineering and what should it be? Inquiry. 2020;68(9):2902–2919. https://doi.org/10.1080/0020174X.2020.1817141
arXiv 2020
-
[53]
The use of ontologies to accelerate the behavioral sciences: Promises and challenges
Sharp C, Kaplan RM, Strauman TJ. The use of ontologies to accelerate the behavioral sciences: Promises and challenges. Current Directions in Psychological Science. 2023;32(5):418–426. https://doi.org/10.1177/09637214231183917
-
[54]
A scoping review of ontologies related to human behaviour change
Norris E, Finnerty AN, Hastings J, Stokes G, Michie S. A scoping review of ontologies related to human behaviour change. Nature Human Behaviour. 2019;3(2):164–172. https://doi.org/10.1038/s41562-018-0511-4
-
[55]
A tool for addressing construct identity in literature reviews and meta-analyses
Larsen KR, Bong CH. A tool for addressing construct identity in literature reviews and meta-analyses. Mis Quarterly. 2016;40(3):529–552. https://doi.org/ 10.25300/misq/2016/40.3.01
-
[56]
Rosenbusch H, Wanders F, Pit IL. The Semantic Scale Network: An online tool to detect semantic overlap of psychological scales and prevent scale redundancies. Psychological Methods. 2020;25(3):380. https://doi.org/10.1037/met0000244
-
[57]
Language models accurately infer correlations between psychological items and scales from text alone
Hommel BE, Arslan RC. Language models accurately infer correlations between psychological items and scales from text alone. PsyArXiv. 2024;PsyArXiv. https: //doi.org/10.31234/osf.io/kjuce
-
[58]
Enhancing scale devel- opment: Pseudo factor analysis of language embedding similarity matrices
Guenole N, D’Urso ED, Samo A, Sun T, Haslbeck J. Enhancing scale devel- opment: Pseudo factor analysis of language embedding similarity matrices. PsyArXiv. 2025;https://doi.org/10.31234/osf.io/vf3se v2. 22
-
[59]
Generative psychometrics via AI-GENIE: Automatic item generation with network-integrated evaluation
Russell-Lasalandra L, Christensen A, Golino H. Generative psychometrics via AI-GENIE: Automatic item generation with network-integrated evaluation. PsyArXiv. 2025;https://doi.org/https://osf.io/preprints/psyarxiv/fgbj4 v2
2025
-
[60]
LLMs4SchemaDiscovery: A human-in-the-loop workflow for scientific schema mining with large language models
Sadruddin S, D’Souza J, Poupaki E, Watkins A, Babaei Giglou H, Rula A, et al. LLMs4SchemaDiscovery: A human-in-the-loop workflow for scientific schema mining with large language models. In: Curry E, editor. The Semantic Web. ESWC 2025. Lecture Notes in Computer Science. Cham: Springer; 2025. p. 244–261
2025
-
[61]
LLMs4OL: Large language models for ontology learning
Babaei Giglou H, D’Souza J, Auer S. LLMs4OL: Large language models for ontology learning. In: International Semantic Web Conference. Springer; 2023. p. 408–427
2023
-
[62]
LLMs4OM: Matching ontologies with large language models
Giglou HB, D’Souza J, Engel F, Auer S. LLMs4OM: Matching ontologies with large language models. arXiv. 2024;https://doi.org/10.48550/arXiv.2404.10317
-
[63]
Escaping the jingle–jangle jungle: Increasing concep- tual clarity in psychology using large language models
Wulff DU, Mata R. Escaping the jingle–jangle jungle: Increasing concep- tual clarity in psychology using large language models. Current Directions in Psychological Science. 2025;p. 09637214251382083. https://doi.org/10.1177/ 09637214251382083
2025
-
[64]
SOAR: An architecture for general intelligence
Laird JE, Newell A, Rosenbloom PS. SOAR: An architecture for general intelligence. Artificial Intelligence. 1987;33(1):1–64. https://doi.org/10.1016/ 0004-3702(87)90050-6
1987
-
[65]
Introduction to computational cognitive modeling
Sun R. Introduction to computational cognitive modeling. In: Sun R, editor. The Cambridge handbook of computational psychology. 1st ed. Cambridge, UK: Cam- bridge University Press; 2001. p. 3–19. Available from: https://www.cambridge. org/core/product/identifier/9780511816772%23c85741-ch1/type/book part
arXiv 2001
-
[66]
Byrne MD. Unified theories of cognition. WIREs Cognitive Science. 2012;3(4):431–438. https://doi.org/10.1002/wcs.1180
-
[67]
Choosing prediction over explanation in psychol- ogy: Lessons from machine learning
Yarkoni T, Westfall J. Choosing prediction over explanation in psychol- ogy: Lessons from machine learning. Perspectives on Psychological Science. 2017;12(6):1100–1122. https://doi.org/10.1177/1745691617693393
-
[68]
Kieval PH, Buckner C. “Captured” by centaur: Opaque predictions or process insights? Journal of Experimental Psychology: Animal Learning and Cognition. 2025 Sep;https://doi.org/10.1037/xan0000410
-
[69]
Centaur may have learned a shortcut that explains away psychological tasks
Xie H, Zhu JQ. Centaur may have learned a shortcut that explains away psychological tasks. PsyArXiv. 2025;https://doi.org/10.31234/osf.io/u7z4t v2
-
[70]
Large language models do not simulate human psychology
Schr¨ oder S, Morgenroth T, Kuhl U, Vaquet V, Paaßen B. Large language models do not simulate human psychology. arXiv. 2025;https://doi.org/10.48550/arXiv. 23 2508.06950
-
[71]
Cognitive modeling using artificial intelligence
Frank MC, Goodman ND. Cognitive modeling using artificial intelligence. Annual Review of Psychology. 2025 Sep 12;https://doi.org/10.31234/osf.io/wv7mg v1
-
[72]
OpinionGPT: Modelling explicit biases in instruction-tuned LLMs
Haller P, Aynetdinov A, Akbik A. OpinionGPT: Modelling explicit biases in instruction-tuned LLMs. In: Chang KW, Lee A, Rajani N, editors. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: System Demonstrations). Mexico City, Mexico: Association for Comp...
2024
-
[73]
Language model fine-tuning on scaled survey data for predicting distributions of public opinions
Suh J, Jahanparast E, Moon S, Kang M, Chang S. Language model fine-tuning on scaled survey data for predicting distributions of public opinions. arXiv. 2025;https://doi.org/10.48550/arXiv.2502.16761
-
[74]
Bhatia S, van Baal ST, Wang F, Walasek L. Computational analysis of 100 K choice dilemmas: Decision attributes, trade-off structures, and model-based pre- diction. Proceedings of the National Academy of Sciences of the United States of America. 2025;122(17):e2406489122. https://doi.org/10.1073/pnas.2406489122
-
[76]
Tosato T, Helbling S, Mantilla-Ramos YJ, Hegazy M, Tosato A, Lemay DJ, et al. Persistent instability in LLM’s personality measurements: Effects of scale, reasoning, and conversation history. arXiv. 2025;https://doi.org/10.48550/arXiv. 2509.13397
-
[77]
Missing the margins: A sys- tematic literature review on the demographic representativeness of LLMs
Sen I, Lutz M, Rogers E, Garcia D, Strohmaier M. Missing the margins: A sys- tematic literature review on the demographic representativeness of LLMs. In: Che W, Nabende J, Shutova E, Pilehvar MT, editors. Findings of the Associ- ation for Computational Linguistics: ACL 2025. Vienna, Austria: Association for Computational Linguistics; 2025. p. 24263–24289....
2025
-
[78]
Large language models that replace human participants can harmfully misportray and flatten identity groups
Wang A, Morgenstern J, Dickerson JP. Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence. 2025;7(3):400–411. https://doi.org/10.1038/ s42256-025-00986-z
2025
-
[79]
AI surrogates and illusions of generalizability in cognitive science
Crockett MJ, Messeri L. AI surrogates and illusions of generalizability in cognitive science. Trends in Cognitive Sciences. 2025;p. S1364661325002517. https://doi. org/10.1016/j.tics.2025.09.012. 24
-
[80]
Reclaiming AI as a theoretical tool for cognitive science
Van Rooij I, Guest O, Adolfi F, de Haan R, Kolokolova A, Rich P. Reclaiming AI as a theoretical tool for cognitive science. Computational Brain & Behavior. 2024;7(4):616–636. https://doi.org/10.1007/s42113-024-00217-5
-
[81]
Why an overreliance on AI-driven modelling is bad for science
Narayanan A, Kapoor S. Why an overreliance on AI-driven modelling is bad for science. Nature. 2025;640(8058):312–314. https://doi.org/10.1038/ d41586-025-01067-2
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.