Pith. sign in

REVIEW 3 major objections 4 minor 3 cited by

This review argues that large language models can address cognitive science's chronic problems—fragmented literatures, vague theories, construct redundancy, narrow models, and missing context—when deployed as tools that complement rather th

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 06:51 UTC pith:LFD4JQZT

load-bearing objection A balanced, well-scoped review that frames LLMs as tools for cognitive science; the 'can help' claim is programmatic and honestly hedged, but the evidence base is thin in places. the 3 major comments →

arxiv 2511.00206 v3 pith:LFD4JQZT submitted 2025-10-31 cs.AI cs.CL

Addressing Longstanding Challenges in Cognitive Science with Language Models

classification cs.AI cs.CL
keywords large language modelscognitive scienceknowledge synthesistheory formalizationmeasurement taxonomiesjingle-jangle fallaciesfoundation modelsecological validity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper's thesis is that the cognitive sciences' most persistent failures—disciplinary silos, informal verbal theories, a proliferation of overlapping constructs and measures, task-bound models that do not generalize, and neglect of cultural and ecological context—are tractable with current language-model tools. It surveys five application classes: semantic research maps, LLM-assisted translation of verbal theories into executable models, embedding-based measurement taxonomies, foundation models such as Centaur that predict behavior across tasks, and contextualized representations from naturalistic data. The paper grounds each suggestion in recent demonstrations, including a pipeline that generated computational models matching or beating established benchmarks and an embedding analysis that condensed a personality construct set by roughly three-quarters. A sympathetic reader comes away with a concrete agenda: use LLMs as instruments for integration and formalization while keeping human oversight, interpretability, and openness as constraints. The payoff, if the thesis holds, is a cognitive science that accumulates rather than fragments.

Core claim

The paper's central claim is programmatic: large language models can function as instruments for a more integrative, cumulative cognitive science. Five chronic challenges are identified—fragmentation, insufficient formalization, conceptual and measurement confusion, lack of generalizability, and neglect of context—and for each the paper assembles evidence that LLMs already provide useful traction. LLM embeddings of article titles and abstracts produce research maps that expose thematic clusters, temporal development, and cross-field bridges; LLM-assisted pipelines translate verbal theories into executable code and even generate new computational models that can match or exceed established on

What carries the argument

The load-bearing mechanism is the semantic embedding: texts, questionnaire items, construct labels, and task descriptions are converted into numerical vectors in a shared high-dimensional space where proximity encodes similarity in meaning. That single representation carries most of the argument—it is what lets research maps reveal conceptual links, lets measurement taxonomies flag jingle–jangle fallacies, and lets a foundation model such as Centaur (a transformer trained on behavioral datasets to predict choices across tasks) treat disparate experimental paradigms as tokens in one language. Around this core, the paper assembles two further mechanisms: LLM generation pipelines that convert v

Load-bearing premise

The entire proposal depends on the assumption that the semantic structure LLMs learn from text faithfully tracks the conceptual and behavioral structure of the mind—an assumption the paper itself flags as an open validation question.

What would settle it

Take a preregistered set of established personality scales and ask whether LLM embeddings of their items recover the scales' known factor structure out-of-sample; if the embeddings systematically place items from different constructs closer than items from the same construct, the measurement-taxonomy claim is falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Research maps built from LLM embeddings can expose conceptual and methodological links between subfields that citation-based tools miss, giving researchers a practical route out of disciplinary silos.
  • LLM-assisted formalization can turn vague verbal theories into executable models with testable predictions; at least one reported pipeline produced models that matched or outperformed established computational models across decision making, learning, planning, and working memory.
  • Embedding-based measurement taxonomies can detect redundant constructs and relabel or eliminate them; the paper reports a demonstration that condensed a hypothesized personality construct set by roughly 75 percent.
  • Multitask foundation models such as Centaur offer a platform for prediction across many cognitive tasks and generalization to unseen task structures, a direct alternative to one-model-per-phenomenon research.
  • LLM analysis of naturalistic data can bring ecological, cultural, and individual variation into cognitive modeling—for instance, by extracting decision attributes and trade-offs from over 100,000 real-world choice dilemmas.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If embedding-based mapping matures, the binding constraint shifts from discovering redundant constructs to governing how they get consolidated; the same tools that expose overlap could be used to run transparent, versioned conceptual-revision processes rather than one-off pruning exercises.
  • The evidence that text-trained LLMs predict behavior on nonverbal tasks suggests a testable conjecture the paper does not state: a large share of human task behavior is driven by linguistically accessible task features. Comparing LLM predictions on verbal versus purely visuospatial task variants would bound that share.
  • The paper's open-infrastructure argument implies a concrete discipline-level experiment: teams using open-weight models should, in aggregate, produce more independently replicable measurement taxonomies and predictions than teams using closed APIs, because contamination and version drift can be audited.
  • The paper's contrasting dystopian and utopian futures imply that the real bottleneck is institutional, not technical; a testable extension would track whether adoption of LLM tools by cognitive scientists is followed by more cross-disciplinary citations and fewer redundant constructs over the next decade.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This perspective/review argues that large language models (LLMs) can help cognitive science address five longstanding challenges: disciplinary silos, insufficient formalization, conceptual and measurement confusion, lack of generalizability, and neglect of ecological context. The authors organize the paper around these five areas, presenting illustrative applications (research maps, theory formalization, measurement taxonomies, integrated frameworks, contextualized representations), followed by a discussion of pitfalls (opacity, oversimplification, bias, data contamination, deskilling) and a call for open infrastructures and human oversight. The central claim is explicitly hedged: LLMs should complement rather than replace human expertise. The paper is a programmatic review rather than a new empirical study.

Significance. If the programmatic claim is accepted, the paper provides a useful organizational frame for a rapidly growing literature at the intersection of LLMs and cognitive science. Its strengths include a balanced treatment of risks, concrete examples for each proposed application area, and a constructive emphasis on open-weight models, interpretability, and validation as scientific priorities. The paper also explicitly lists outstanding questions, which is helpful for agenda-setting. However, the positive evidence is drawn from a small set of demonstrations, several of which are non-peer-reviewed preprints or the authors' own work, and the foundational assumption that embedding similarity tracks psychological/construct similarity is acknowledged but not critically examined in depth. The paper's value is therefore more synthesizing and cautionary than evidential.

major comments (3)
  1. [Section 2.3 / Table 1, row 3 / Figure 3] The measurement-taxonomy argument is load-bearing for the claim that LLMs can reduce conceptual and measurement confusion. However, the section treats semantic embedding proximity as a proxy for construct and measure equivalence without providing a benchmark or validation protocol. The Wulff and Mata [19] demonstration is described only as a 'sketch' that reduces a construct set by roughly 75%, and the paper defers validation to Box 3, Q1. Without evidence that these relabelings correspond to convergent/discriminant validity, improved predictive utility, or expert agreement, the claim risks reducing to the truism that LLMs can manipulate text. I recommend either tempering the wording to 'exploratory' or adding a concrete discussion of validation criteria and known failure cases.
  2. [Section 2.1 / Box 1 / Figure 2] The 'theory of mind' research map is presented as an example of how LLMs can map fragmented literatures, but no evidence is provided for the map's accuracy, stability, or interpretability. The historical narrative (e.g., the field 'originated from clusters in autism and child development') is derived from a cluster projection whose validity is not assessed. The manual labeling of clusters based on author keywords is a subjective step. For the first challenge in Table 1 to be convincing, the paper should at least note that such maps are exploratory hypotheses generators, not validated descriptions of the field, and ideally cite any reliability checks or human-evaluation results.
  3. [Sections 2.1–2.3 and reference list] Several of the central supporting examples are non-peer-reviewed preprints or come from the authors' own research program, including [12], [16], [49], [19], [63], [97], and [98]. This does not invalidate the review, but it is a load-bearing limitation because the paper's 'can help' claim rests on these selected demonstrations. I ask the authors to explicitly flag preprint status, distinguish their own work from independent replications, and note where independent verification is still needed. Without such signaling, a skeptical reader cannot separate established results from programmatic ambitions.
minor comments (4)
  1. [Box 3, Q1 and Q5] Q1 refers to 'cf. Box 2.1', but no such box exists; this should be 'Box 1'. Q5 refers to 'cf. Box 3', but it should refer to 'Box 2' (the box on open LLMs).
  2. [Figure 1 caption and Glossary] The words 'F ormal' and 'F ormal Models' appear with an unusual space. This looks like a typographical artifact and should be corrected.
  3. [Section 2.5] 'in-silico testing' should be 'in silico testing' (italicized, no hyphen), per standard usage.
  4. [Box 1 / Figure 2 caption] The caption says 'For more details see Thoma et al. [12]', but [12] is about behavioral reinforcement learning, not theory of mind. Either the theory-of-mind map is generated using the same method as [12], or a different reference/companion document is needed. Please clarify.

Circularity Check

0 steps flagged

No significant circularity: the paper is an argumentative review that relies on externally benchmarked examples and explicitly defers validation rather than presenting fitted outputs as predictions.

full rationale

The paper does not contain a derivation chain that reduces a claimed result to its inputs. It is a perspective/review whose central claim, that LLMs can serve as tools for a more integrative and cumulative cognitive science, is supported by a mix of independent external benchmarks and published applications. The self-citations (e.g., Wulff & Mata [19], [63]; Fulawka et al. [49]) are illustrative examples, not the sole or definitional basis for the claim. The measurement-taxonomy example is anchored to an external criterion: the paper states that the semantic representations 'reproduced empirical item-scale relationships well,' which is a benchmark outside the embedding construction, not a circular fit. The paper also explicitly flags the need for validation: Section 2.2 says 'these efforts require validation,' and Box 3 asks 'How can we best validate the accuracy and completeness of knowledge structures... automatically extracted by large language models?' Section 3 further acknowledges that LLMs 'may rely on statistical shortcuts or reflect biases in training data.' No equation or claimed predictive result is shown to be equivalent to a fitted parameter, and no uniqueness theorem or ansatz is imported from the authors' prior work. The presence of self-citation alone is not circularity, especially where the cited works are independently published and externally falsifiable. Thus no specific circular step can be exhibited, and the appropriate score is 0.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No fitted parameters or new entities are introduced. The load-bearing assumptions are domain-level: that LLM representations and predictions are valid proxies for psychological structure and behavior, and that human oversight can mitigate LLM risks. These are explicit or implicit in the paper's argument and are only partially acknowledged as open questions.

axioms (4)
  • domain assumption Semantic embeddings from LLMs capture meaningful psychological and conceptual structure.
    The proposed tools (research maps, measurement taxonomies, formalization) assume LLM representations reflect true relations among constructs, measures, and findings. Acknowledged as needing validation in Box 3 Q1 and Section 2.2 ('these efforts require validation').
  • domain assumption LLM performance on textual tasks transfers to behavioral and cognitive prediction.
    The integrated-framework and contextualized-representation sections assume text-trained models can predict human behavior (e.g., Centaur [20]). The paper notes brittleness [68–70] but does not establish general validity.
  • domain assumption Literature structure inferred by embeddings is a faithful map of scientific knowledge.
    Research maps embed titles/abstracts and assume proximity reflects conceptual similarity, with clusters manually labeled. No external validation is reported in this paper.
  • domain assumption Human researchers can effectively supervise and correct LLM outputs.
    The recommendation that LLMs complement rather than replace humans assumes effective human oversight is feasible despite opacity, deskilling, and automation bias concerns raised in Section 3.

pith-pipeline@v1.3.0-alltime-deepseek · 16659 in / 7683 out tokens · 69212 ms · 2026-08-04T06:51:17.364383+00:00 · methodology

0 comments
read the original abstract

Cognitive science faces ongoing challenges in research integration, formalization, conceptual clarity, and other areas, in part due to its multifaceted and interdisciplinary nature. Recent advances in artificial intelligence, particularly the development of language models, offer tools that may help to address these longstanding issues. Specifically, they can help map fragmented literatures, formalize verbal theories, identify overlap among constructs and measures, generate predictions across tasks, and extract cultural or ecological structure from naturalistic data. However, these opportunities come with risks, including oversimplification, opacity, deskilling, and bias. Taken together, we conclude that language models could serve as tools for a more integrative and cumulative cognitive science when used judiciously to complement, rather than replace, human agency.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Psychological Competence as a Missing Dimension in AI Evaluation

    cs.AI 2026-07 conditional novelty 5.0

    Psychological competence—the capacity of human-facing AI to support appropriate user cognition, emotion, and decision-making—should be evaluated as a distinct interaction-level dimension beyond technical and safety be...

  2. Psychological Competence as a Missing Dimension in AI Evaluation

    cs.AI 2026-07 accept novelty 5.0

    AI evaluation should add 'psychological competence' — how well a system supports user reasoning, emotional stability, and autonomous decision-making — as a core dimension.

  3. Taming the Centaur(s) with LAPITHS: a framework for a theoretically grounded interpretation of AI performances

    cs.AI 2026-04 unverdicted novelty 5.0

    LAPITHS shows that major claims for CENTAUR as a unified model of cognition lack theoretical or empirical support because similar results can be obtained from systems without the structural features associated with co...

Reference graph

Works this paper leans on

98 extracted references · 32 canonical work pages · cited by 2 Pith papers · 1 internal anchor

  1. [1]

    The mind’s new science: A history of the cognitive revolution

    Gardner H. The mind’s new science: A history of the cognitive revolution. New York, USA: Basic Books; 1985

  2. [2]

    What happened to cognitive science? Nature Human Behaviour

    N´ u˜ nez R, Allen M, Gao R, Miller Rigoli C, Relaford-Doyle J, Semenuks A. What happened to cognitive science? Nature Human Behaviour. 2019;3(8):782–791. https://doi.org/10.1038/s41562-019-0626-2

  3. [3]

    Formalizing verbal theories: A tutorial by dialogue

    Van Rooij I, Blokpoel M. Formalizing verbal theories: A tutorial by dialogue. Social Psychology. 2020;51(5):285–298. https://doi.org/10.1027/1864-9335/ a000428

  4. [4]

    How computational modeling can force theory building in psychological science

    Guest O, Martin AE. How computational modeling can force theory building in psychological science. Perspectives on Psychological Science. 2021;16(4):789–802. https://doi.org/10.1177/1745691620970585

  5. [5]

    Defragmenting psychology

    Anvari F, Alsalti T, Oehler LA, Hussey I, Elson M, Arslan RC. Defragmenting psychology. Nature Human Behaviour. 2025;9:836–839. https://doi.org/10.1038/ s41562-025-02138-0

  6. [6]

    You can’t play 20 questions with nature and win: Projective comments on the papers of this symposium

    Newell A. You can’t play 20 questions with nature and win: Projective comments on the papers of this symposium. In: Chase WG, editor. Visual Information Processing. New York, USA: Academic Press (Elsevier); 1973. p. 283–308

  7. [7]

    An integrated theory of the mind

    Anderson JR, Bothell D, Byrne MD, Douglass S, Lebiere C, Qin Y. An integrated theory of the mind. Psychological Review. 2004;111(4):1036–1060. https://doi. org/10.1037/0033-295X.111.4.1036

  8. [8]

    The weirdest people in the world? Behavioral and Brain Sciences

    Henrich J, Heine SJ, Norenzayan A. The weirdest people in the world? Behavioral and Brain Sciences. 2010;33(2-3):61 – 83. https://doi.org/10.1017/ S0140525X0999152X

  9. [9]

    Participant diver- sity is necessary to advance brain aging research

    Wig GS, Klausner S, Chan MY, Sullins C, Rayanki A, Seale M. Participant diver- sity is necessary to advance brain aging research. Trends in Cognitive Sciences. 2024;28(2):92–96. https://doi.org/10.1016/j.tics.2023.12.004

  10. [10]

    Parallel distributed processing: explorations in the microstructure of cognition, vol

    Rumelhart DE, McClelland JL, PDP Research Group C, editors. Parallel distributed processing: explorations in the microstructure of cognition, vol. 1: foundations. Cambridge, MA, USA: MIT Press; 1986

  11. [11]

    The two disciplines of scientific psychology

    Cronbach LJ. The two disciplines of scientific psychology. The American Psychol- ogist. 1957;12(11):671–684. https://doi.org/https://doi.org/10.1037/h0043943

  12. [12]

    Map- ping the landscape of behavioral reinforcement learning research

    Thoma AI, Bolenz F, Tiede K, Yang Y, Palminteri S, Hertwig R, et al. Map- ping the landscape of behavioral reinforcement learning research. PsyArXiv. 2025;https://doi.org/10.31234/osf.io/6c2va v1. 18

  13. [13]

    Large language models surpass human experts in predicting neuroscience results

    Luo X, Rechardt A, Sun G, Nejad KK, Y´ a˜ nez F, Yilmaz B, et al. Large language models surpass human experts in predicting neuroscience results. Nature Human Behaviour. 2024;9(2):305––315. https://doi.org/10.1038/s41562-024-02046-9

  14. [14]

    Addressing the theory crisis in psychology

    Oberauer K, Lewandowsky S. Addressing the theory crisis in psychology. Psy- chonomic Bulletin & Review. 2019;26(5):1596–1618. https://doi.org/10.3758/ s13423-019-01645-2

  15. [15]

    How to translate a verbal theory into a formal model

    Smaldino PE. How to translate a verbal theory into a formal model. Social Psychology. 2020;51(4):207–218. https://doi.org/10.1027/1864-9335/a000425

  16. [16]

    Generating com- putational cognitive models using large language models

    Rmus M, Jagadish AK, Mathony M, Ludwig T, Schulz E. Generating com- putational cognitive models using large language models. arXiv. 2025;https: //doi.org/10.48550/arXiv.2502.00879

  17. [17]

    theoraizer: AI-assisted theory construction

    Waaijers M, Rosenbusch H, Van Lissa CJ, Roefs A, Borsboom D. theoraizer: AI-assisted theory construction. PsyArXiv. 2024 Aug;OSF. https://doi.org/10. 31234/osf.io/gu9yq

  18. [18]

    From brain maps to cognitive ontologies: Infor- matics and the search for mental structure

    Poldrack RA, Yarkoni T. From brain maps to cognitive ontologies: Infor- matics and the search for mental structure. Annual Review of Psychology. 2016;67(1):587–612. https://doi.org/10.1146/annurev-psych-122414-033729

  19. [19]

    Semantic embeddings reveal and address taxonomic incommensurability in psychological measurement

    Wulff DU, Mata R. Semantic embeddings reveal and address taxonomic incommensurability in psychological measurement. Nature Human Behaviour. 2025;9(5):944–954. https://doi.org/10.1038/s41562-024-02089-y

  20. [20]

    A foundation model to predict and capture human cognition

    Binz M, Akata E, Bethge M, Br¨ andle F, Callaway F, Coda-Forno J, et al. A foundation model to predict and capture human cognition. Nature. 2025;644(8078):1002–1009. https://doi.org/10.1038/s41586-025-09215-4

  21. [21]

    Towards a cognitive science of the human: Cross-cultural approaches and their urgency

    Barrett HC. Towards a cognitive science of the human: Cross-cultural approaches and their urgency. Trends in Cognitive Sciences. 2020;24(8):620–638. https:// doi.org/10.1016/j.tics.2020.05.007

  22. [22]

    Building Knowledge-Guided Lexica to Model Cultural Variation

    Havaldar S, Giorgi S, Rai S, Talhelm T, Guntuku SC, Ungar L. Building Knowledge-Guided Lexica to Model Cultural Variation. In: Duh K, Gomez H, Bethard S, editors. Proceedings of the 2024 Conference of the North Ameri- can Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Mexico City, Mexico: Ass...

  23. [23]

    On the conversational persuasiveness of GPT-4

    Salvi F. On the conversational persuasiveness of GPT-4. Nature Human Behaviour. 2025;9:1645–1653. https://doi.org/10.1038/s41562-025-02194-6. 19

  24. [24]

    The two cultures and the scientific revolution

    Snow CP. The two cultures and the scientific revolution. New York, USA: Cambridge University Press; 1959

  25. [25]

    Ending the reading wars: Reading acquisition from novice to expert

    Castles A, Rastle K, Nation K. Ending the reading wars: Reading acquisition from novice to expert. Psychological Science in the Public Interest. 2018;19(1):5–51. https://doi.org/10.1177/1529100618772271

  26. [26]

    Risk preference shares the psychometric structure of major psychological traits

    Frey R, Pedroni A, Mata R, Rieskamp J, Hertwig R. Risk preference shares the psychometric structure of major psychological traits. Scientific Advances. 2017;3(10):e1701381. https://doi.org/10.1126/sciadv.1701381

  27. [27]

    Risk preference: A view from psychology

    Mata R, Frey R, Richter D, Schupp J, Hertwig R. Risk preference: A view from psychology. Journal of Economic Perspectives;32(2):155–172. https://doi.org/10. 1257/jep.32.2.155

  28. [28]

    Three gaps and what they may mean for risk preference

    Hertwig R, Wulff DU, Mata R. Three gaps and what they may mean for risk preference. Philosophical Transactions of the Royal Society Series B, Biological Sciences. 2019;374(1766):20180140. https://doi.org/10.1098/rstb.2018.0140

  29. [29]

    In defense of disciplines: Interdisciplinarity and specialization in the research university

    Jacobs JA. In defense of disciplines: Interdisciplinarity and specialization in the research university. Chicago, IL: University of Chicago Press; 2014

  30. [30]

    How should the advancement of large language models affect the practice of science? Proceedings of the National Academy of Sciences of the United States of America

    Binz M, Alaniz S, Roskies A, Aczel B, Bergstrom CT, Allen C, et al. How should the advancement of large language models affect the practice of science? Proceedings of the National Academy of Sciences of the United States of America. 2025;122(5):e2401227121. https://doi.org/10.1073/pnas.2401227121

  31. [31]

    Transforming science with large language models: A survey on AI-assisted scientific discovery, experimentation, content generation, and evaluation

    Eger S, Cao Y, D’Souza J, Geiger A, Greisinger C, Gross S, et al. Transforming science with large language models: A survey on AI-assisted scientific discovery, experimentation, content generation, and evaluation. arXiv. 2025;https://doi. org/10.48550/arXiv.2502.05151

  32. [32]

    Toward systematic review automation: A practical guide to using machine learning tools in research synthesis

    Marshall IJ, Wallace BC. Toward systematic review automation: A practical guide to using machine learning tools in research synthesis. Systematic Reviews. 2019 Dec;8(1):163. https://doi.org/10.1186/s13643-019-1074-9

  33. [33]

    A digital future for the history of psychology? History of Psychology

    Green CD. A digital future for the history of psychology? History of Psychology. 2016;19(3):209–219. https://doi.org/10.1037/hop0000012

  34. [34]

    LLMs4Synthesis: Leveraging large language models for scientific synthesis

    Babaei Giglou H, D’Souza J, Auer S. LLMs4Synthesis: Leveraging large language models for scientific synthesis. Proceedings of the 24th ACM/IEEE Joint Con- ference on Digital Libraries. 2024 Dec;p. 1–12. https://doi.org/10.1145/3677389. 3702565

  35. [35]

    Exploring the applicability of large language models to citation context analysis

    Nishikawa K, Koshiba H. Exploring the applicability of large language models to citation context analysis. Scientometrics. 2024;129:6751–6777. https://doi.org/ 10.1007/s11192-024-05142-9. 20

  36. [36]

    Identifying interdisciplinary emergence in the science of science: combination of network analysis and BERTopic

    Kim K, Kogler DF, Maliphol S. Identifying interdisciplinary emergence in the science of science: combination of network analysis and BERTopic. Humanities and Social Sciences Communications. 2024;11(1):1–15. https://doi.org/10.1057/ s41599-024-03044-y

  37. [37]

    An embedding approach for ana- lyzing the evolution of research topics with a case study on computer science subdomains

    Taher Harikandeh SR, Aliakbary S, Taheri S. An embedding approach for ana- lyzing the evolution of research topics with a case study on computer science subdomains. Scientometrics. 2023;128(3):1567–1582. https://doi.org/10.1007/ s11192-023-04642-4

  38. [38]

    Using sequences of life-events to predict human lives

    Savcisens G, Eliassi-Rad T, Hansen LK, Mortensen LH, Lilleholt L, Rogers A, et al. Using sequences of life-events to predict human lives. Nature Computational Science. 2023 Dec;4(1):43–56. https://doi.org/10.1038/s43588-023-00573-5

  39. [39]

    Automating the practice of science: Opportunities, challenges, and impli- cations

    Musslick S, Bartlett LK, Chandramouli SH, Dubova M, Gobet F, Griffiths TL, et al. Automating the practice of science: Opportunities, challenges, and impli- cations. Proceedings of the National Academy of Sciences of the United States of America. 2025;122(5):e2401238121. https://doi.org/10.1073/pnas.2401238121

  40. [40]

    Minimal coherence among varied theory of mind measures in childhood and adulthood

    Warnell KR, Redcay E. Minimal coherence among varied theory of mind measures in childhood and adulthood. Cognition. 2019;191:103997. https://doi.org/10. 1016/j.cognition.2019.06.009

  41. [41]

    Deconstructing and recon- structing theory of mind

    Schaafsma SM, Pfaff DW, Spunt RP, Adolphs R. Deconstructing and recon- structing theory of mind. Trends in Cognitive Sciences. 2015;19(2):65–72. https: //doi.org/10.1016/j.tics.2014.11.007

  42. [42]

    Qwen3 Embedding: Advancing text embedding and reranking through foundation models

    Zhang Y, Li M, Long D, Zhang X, Lin H, Yang B, et al. Qwen3 Embedding: Advancing text embedding and reranking through foundation models. arXiv. 2025;https://doi.org/10.48550/arXiv.2506.05176

  43. [43]

    Understanding how dimension reduc- tion tools work: An empirical approach to deciphering t-SNE, UMAP, TriMap, and PaCMAP for data visualization

    Wang Y, Huang H, Rudin C, Shaposhnik Y. Understanding how dimension reduc- tion tools work: An empirical approach to deciphering t-SNE, UMAP, TriMap, and PaCMAP for data visualization. Journal of Machine Learning Research. 2021;22(201):1–73. https://doi.org/http://jmlr.org/papers/v22/20-1061.html

  44. [44]

    Cognitive science is and should be pluralistic

    Gentner D. Cognitive science is and should be pluralistic. Topics in Cognitive Science. 2019 Oct;11(4):884–891. https://doi.org/10.1111/tops.12459

  45. [45]

    Why summaries of research on psychological theories are often unin- terpretable

    Meehl PE. Why summaries of research on psychological theories are often unin- terpretable. Psychological Reports. 1990;66(1):195–244. https://doi.org/10.2466/ PR0.66.1.195-244

  46. [46]

    The Human Behaviour-Change Project: Harnessing the power of artificial intelligence and machine learning for evidence synthesis and interpretation

    Michie S, Thomas J, Johnston M, Aonghusa PM, Shawe-Taylor J, Kelly MP, et al. The Human Behaviour-Change Project: Harnessing the power of artificial intelligence and machine learning for evidence synthesis and interpretation. Imple- mentation Science. 2017;12(1):121. https://doi.org/10.1186/s13012-017-0641-5. 21

  47. [47]

    Computational Models in Personality and Social Psychol- ogy

    Read SJ, Monroe BM. Computational Models in Personality and Social Psychol- ogy. In: Sun R, editor. The Cambridge Handbook of Computational Cognitive Sciences. Cambridge, UK: Cambridge University Press; 2023. p. 795–835

  48. [48]

    Using large-scale experiments and machine learning to discover theories of human decision-making

    Peterson JC, Bourgin DD, Agrawal M, Reichman D, Griffiths TL. Using large-scale experiments and machine learning to discover theories of human decision-making. Science. 2021;372(6547):1209–1214. https://doi.org/10.1126/ science.abe2629

  49. [49]

    Large language models accurately identify decision reasons in verbal reports

    Fulawka K, Hertwig R, Wulff DU. Large language models accurately identify decision reasons in verbal reports. PsyArXiv. 2025;https://doi.org/10.31234/osf. io/yuzmw v1

  50. [50]

    Psychological measures aren’t tooth- brushes

    Elson M, Hussey I, Alsalti T, Arslan RC. Psychological measures aren’t tooth- brushes. Communications Psychology. 2023;1(1):25. https://doi.org/10.1038/ s44271-023-00026-9

  51. [51]

    The theory crisis in psychology: How to move forward

    Eronen MI, Bringmann LF. The theory crisis in psychology: How to move forward. Perspectives on Psychological Science. 2021;16(4):779–788. https://doi.org/10. 1177/1745691620970586

  52. [52]

    What is conceptual engineering and what should it be? Inquiry

    Chalmers DJ. What is conceptual engineering and what should it be? Inquiry. 2020;68(9):2902–2919. https://doi.org/10.1080/0020174X.2020.1817141

  53. [53]

    The use of ontologies to accelerate the behavioral sciences: Promises and challenges

    Sharp C, Kaplan RM, Strauman TJ. The use of ontologies to accelerate the behavioral sciences: Promises and challenges. Current Directions in Psychological Science. 2023;32(5):418–426. https://doi.org/10.1177/09637214231183917

  54. [54]

    A scoping review of ontologies related to human behaviour change

    Norris E, Finnerty AN, Hastings J, Stokes G, Michie S. A scoping review of ontologies related to human behaviour change. Nature Human Behaviour. 2019;3(2):164–172. https://doi.org/10.1038/s41562-018-0511-4

  55. [55]

    A tool for addressing construct identity in literature reviews and meta-analyses

    Larsen KR, Bong CH. A tool for addressing construct identity in literature reviews and meta-analyses. Mis Quarterly. 2016;40(3):529–552. https://doi.org/ 10.25300/misq/2016/40.3.01

  56. [56]

    The Semantic Scale Network: An online tool to detect semantic overlap of psychological scales and prevent scale redundancies

    Rosenbusch H, Wanders F, Pit IL. The Semantic Scale Network: An online tool to detect semantic overlap of psychological scales and prevent scale redundancies. Psychological Methods. 2020;25(3):380. https://doi.org/10.1037/met0000244

  57. [57]

    Language models accurately infer correlations between psychological items and scales from text alone

    Hommel BE, Arslan RC. Language models accurately infer correlations between psychological items and scales from text alone. PsyArXiv. 2024;PsyArXiv. https: //doi.org/10.31234/osf.io/kjuce

  58. [58]

    Enhancing scale devel- opment: Pseudo factor analysis of language embedding similarity matrices

    Guenole N, D’Urso ED, Samo A, Sun T, Haslbeck J. Enhancing scale devel- opment: Pseudo factor analysis of language embedding similarity matrices. PsyArXiv. 2025;https://doi.org/10.31234/osf.io/vf3se v2. 22

  59. [59]

    Generative psychometrics via AI-GENIE: Automatic item generation with network-integrated evaluation

    Russell-Lasalandra L, Christensen A, Golino H. Generative psychometrics via AI-GENIE: Automatic item generation with network-integrated evaluation. PsyArXiv. 2025;https://doi.org/https://osf.io/preprints/psyarxiv/fgbj4 v2

  60. [60]

    LLMs4SchemaDiscovery: A human-in-the-loop workflow for scientific schema mining with large language models

    Sadruddin S, D’Souza J, Poupaki E, Watkins A, Babaei Giglou H, Rula A, et al. LLMs4SchemaDiscovery: A human-in-the-loop workflow for scientific schema mining with large language models. In: Curry E, editor. The Semantic Web. ESWC 2025. Lecture Notes in Computer Science. Cham: Springer; 2025. p. 244–261

  61. [61]

    LLMs4OL: Large language models for ontology learning

    Babaei Giglou H, D’Souza J, Auer S. LLMs4OL: Large language models for ontology learning. In: International Semantic Web Conference. Springer; 2023. p. 408–427

  62. [62]

    LLMs4OM: Matching ontologies with large language models

    Giglou HB, D’Souza J, Engel F, Auer S. LLMs4OM: Matching ontologies with large language models. arXiv. 2024;https://doi.org/10.48550/arXiv.2404.10317

  63. [63]

    Escaping the jingle–jangle jungle: Increasing concep- tual clarity in psychology using large language models

    Wulff DU, Mata R. Escaping the jingle–jangle jungle: Increasing concep- tual clarity in psychology using large language models. Current Directions in Psychological Science. 2025;p. 09637214251382083. https://doi.org/10.1177/ 09637214251382083

  64. [64]

    SOAR: An architecture for general intelligence

    Laird JE, Newell A, Rosenbloom PS. SOAR: An architecture for general intelligence. Artificial Intelligence. 1987;33(1):1–64. https://doi.org/10.1016/ 0004-3702(87)90050-6

  65. [65]

    Introduction to computational cognitive modeling

    Sun R. Introduction to computational cognitive modeling. In: Sun R, editor. The Cambridge handbook of computational psychology. 1st ed. Cambridge, UK: Cam- bridge University Press; 2001. p. 3–19. Available from: https://www.cambridge. org/core/product/identifier/9780511816772%23c85741-ch1/type/book part

  66. [66]

    Unified theories of cognition

    Byrne MD. Unified theories of cognition. WIREs Cognitive Science. 2012;3(4):431–438. https://doi.org/10.1002/wcs.1180

  67. [67]

    Choosing prediction over explanation in psychol- ogy: Lessons from machine learning

    Yarkoni T, Westfall J. Choosing prediction over explanation in psychol- ogy: Lessons from machine learning. Perspectives on Psychological Science. 2017;12(6):1100–1122. https://doi.org/10.1177/1745691617693393

  68. [68]

    Captured

    Kieval PH, Buckner C. “Captured” by centaur: Opaque predictions or process insights? Journal of Experimental Psychology: Animal Learning and Cognition. 2025 Sep;https://doi.org/10.1037/xan0000410

  69. [69]

    Centaur may have learned a shortcut that explains away psychological tasks

    Xie H, Zhu JQ. Centaur may have learned a shortcut that explains away psychological tasks. PsyArXiv. 2025;https://doi.org/10.31234/osf.io/u7z4t v2

  70. [70]

    Large language models do not simulate human psychology

    Schr¨ oder S, Morgenroth T, Kuhl U, Vaquet V, Paaßen B. Large language models do not simulate human psychology. arXiv. 2025;https://doi.org/10.48550/arXiv. 23 2508.06950

  71. [71]

    Cognitive modeling using artificial intelligence

    Frank MC, Goodman ND. Cognitive modeling using artificial intelligence. Annual Review of Psychology. 2025 Sep 12;https://doi.org/10.31234/osf.io/wv7mg v1

  72. [72]

    OpinionGPT: Modelling explicit biases in instruction-tuned LLMs

    Haller P, Aynetdinov A, Akbik A. OpinionGPT: Modelling explicit biases in instruction-tuned LLMs. In: Chang KW, Lee A, Rajani N, editors. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: System Demonstrations). Mexico City, Mexico: Association for Comp...

  73. [73]

    Language model fine-tuning on scaled survey data for predicting distributions of public opinions

    Suh J, Jahanparast E, Moon S, Kang M, Chang S. Language model fine-tuning on scaled survey data for predicting distributions of public opinions. arXiv. 2025;https://doi.org/10.48550/arXiv.2502.16761

  74. [74]

    Computational analysis of 100 K choice dilemmas: Decision attributes, trade-off structures, and model-based pre- diction

    Bhatia S, van Baal ST, Wang F, Walasek L. Computational analysis of 100 K choice dilemmas: Decision attributes, trade-off structures, and model-based pre- diction. Proceedings of the National Academy of Sciences of the United States of America. 2025;122(17):e2406489122. https://doi.org/10.1073/pnas.2406489122

  75. [76]

    Persistent instability in LLM’s personality measurements: Effects of scale, reasoning, and conversation history

    Tosato T, Helbling S, Mantilla-Ramos YJ, Hegazy M, Tosato A, Lemay DJ, et al. Persistent instability in LLM’s personality measurements: Effects of scale, reasoning, and conversation history. arXiv. 2025;https://doi.org/10.48550/arXiv. 2509.13397

  76. [77]

    Missing the margins: A sys- tematic literature review on the demographic representativeness of LLMs

    Sen I, Lutz M, Rogers E, Garcia D, Strohmaier M. Missing the margins: A sys- tematic literature review on the demographic representativeness of LLMs. In: Che W, Nabende J, Shutova E, Pilehvar MT, editors. Findings of the Associ- ation for Computational Linguistics: ACL 2025. Vienna, Austria: Association for Computational Linguistics; 2025. p. 24263–24289....

  77. [78]

    Large language models that replace human participants can harmfully misportray and flatten identity groups

    Wang A, Morgenstern J, Dickerson JP. Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence. 2025;7(3):400–411. https://doi.org/10.1038/ s42256-025-00986-z

  78. [79]

    AI surrogates and illusions of generalizability in cognitive science

    Crockett MJ, Messeri L. AI surrogates and illusions of generalizability in cognitive science. Trends in Cognitive Sciences. 2025;p. S1364661325002517. https://doi. org/10.1016/j.tics.2025.09.012. 24

  79. [80]

    Reclaiming AI as a theoretical tool for cognitive science

    Van Rooij I, Guest O, Adolfi F, de Haan R, Kolokolova A, Rich P. Reclaiming AI as a theoretical tool for cognitive science. Computational Brain & Behavior. 2024;7(4):616–636. https://doi.org/10.1007/s42113-024-00217-5

  80. [81]

    Why an overreliance on AI-driven modelling is bad for science

    Narayanan A, Kapoor S. Why an overreliance on AI-driven modelling is bad for science. Nature. 2025;640(8058):312–314. https://doi.org/10.1038/ d41586-025-01067-2

Showing first 80 references.