Pith. sign in

REVIEW 9 cited by

Turning large language models into cognitive models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03917 v1 pith:UGX2MIBF submitted 2023-06-06 cs.CL cs.AIcs.LG

Turning large language models into cognitive models

classification cs.CL cs.AIcs.LG
keywords modelscognitivelargelanguagebehaviorfinetuninghumanrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large language models are powerful systems that excel at many tasks, ranging from translation to mathematical reasoning. Yet, at the same time, these models often show unhuman-like characteristics. In the present paper, we address this gap and ask whether large language models can be turned into cognitive models. We find that -- after finetuning them on data from psychological experiments -- these models offer accurate representations of human behavior, even outperforming traditional cognitive models in two decision-making domains. In addition, we show that their representations contain the information necessary to model behavior on the level of individual subjects. Finally, we demonstrate that finetuning on multiple tasks enables large language models to predict human behavior in a previously unseen task. Taken together, these results suggest that large, pre-trained models can be adapted to become generalist cognitive models, thereby opening up new research directions that could transform cognitive psychology and the behavioral sciences as a whole.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

    cs.CL 2024-10 unverdicted novelty 8.0

    ErrorRadar is a new benchmark of 2,500 multimodal K-12 math problems for MLLM error step identification and categorization, where GPT-4o trails human experts by ~10%.

  2. Simulating Word Suggestion Usage in Mobile Typing to Guide Intelligent Text Entry Design

    cs.HC 2026-02 unverdicted novelty 7.0

    WSTypist is a new RL-based simulation model that reproduces human-like word suggestion strategies, individual differences, and adaptation to design changes in mobile text entry.

  3. Thinking Under Uncertainty: Evidence Use and Information-Seeking in Language Models

    cs.LG 2026-07 conditional novelty 6.0

    Inference-time thinking strengthens value-guided choice and cuts noise in LLMs, but does not produce UCB-like or stronger Thompson-like information-seeking on a controlled bandit task.

  4. Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

    cs.CL 2026-07 conditional novelty 6.0

    Combining language-model generation with rule-based selection reproduces several pragmatic phenomena, but the language models only worked reliably as idea generators, not as judges of formal linguistic properties.

  5. Mixture of Cognitive Experts in Large Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0

    Routing CV experts into atomic evidence then Bloom-staged verbalization improves LVLM benchmarks and yields measurable query-conditioned reasoning traces.

  6. Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games

    cs.AI 2026-06 unverdicted novelty 6.0

    Equation-to-Behavior Prompting lets large LLMs match cognitive models like Bayesian updating in persuasion games; RL training cuts small-model belief error by 26.5% and improves diverse training outcomes by 2.5-12%.

  7. From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets

    econ.GN 2025-09 conditional novelty 6.0

    LLM experts in credence goods markets reduce efficiency and consumer surplus unless liability or transparent prosocial objectives operate, and expert delegation with transparent objectives can outperform human-only markets.

  8. Teaching Values to Machines: Simulating Human-Like Behavior in LLMs

    cs.AI 2026-05 unverdicted novelty 5.0

    Value-prompted LLMs align with human value structures and value-behavior relationships, and incorporating human value distributions improves population-level simulations.

  9. Can You Trick the Grader? Adversarial Persuasion of LLM Judges

    cs.CL 2025-08 conditional novelty 5.0

    Strategically inserted persuasive sentences inflate LLM judges' scores for incorrect math solutions across six benchmarks and fourteen models, but the study lacks length-matched controls separating rhetoric from lengt...