REVIEW 30 cited by
Machine Psychology
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) show increasingly advanced emergent capabilities and are being incorporated across various societal domains. Understanding their behavior and reasoning abilities therefore holds significant importance. We argue that a fruitful direction for research is engaging LLMs in behavioral experiments inspired by psychology that have traditionally been aimed at understanding human cognition and behavior. In this article, we highlight and summarize theoretical perspectives, experimental paradigms, and computational analysis techniques that this approach brings to the table. It paves the way for a "machine psychology" for generative artificial intelligence (AI) that goes beyond performance benchmarks and focuses instead on computational insights that move us toward a better understanding and discovery of emergent abilities and behavioral patterns in LLMs. We review existing work taking this approach, synthesize best practices, and highlight promising future directions. We also highlight the important caveats of applying methodologies designed for understanding humans to machines. We posit that leveraging tools from experimental psychology to study AI will become increasingly valuable as models evolve to be more powerful, opaque, multi-modal, and integrated into complex real-world settings.
Forward citations
Cited by 30 Pith papers
-
The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment
LLM yes-no bias on moral dilemmas is an order-plus-lexical surface artifact, not a moral shift; models have a nearly format-invariant graded stance that the standard binary readout confounds.
-
The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals
Task conditioning suppresses safety-critical signal reporting in language and vision models that unconstrained versions report at higher rates, creating an inattentional gap that decouples benchmark safety from real-w...
-
Strategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory
Frontier LLMs survive and often thrive in evolutionary Prisoner's Dilemma tournaments, and each model family shows a distinct, context-dependent cooperation fingerprint.
-
Natural Language Processing Psychometrics
LLM persona questionnaire responses yield interpretable text features that transfer, with modest accuracy, to classifying depression in real human speech.
-
Social Pressure Breaks Majority Voting in LLM Safety Panels
Shared wrong-label peer messages drive LLM safety-review panels to a 100% false-alarm rate, because each reviewer adopts the push toward 'unsafe' and majority voting then amplifies the individual over-flagging.
-
Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
Post-training Qwen3.5-122B-A10B on 363 office workflow tasks improved SWE-Bench Pro pass@1 by 5.8 points, with trajectory analysis attributing the gain to four general goal-directed behaviors.
-
From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs
SAE features recovered from matched high–low trait behaviors can be steered to bidirectionally shift situational personality expression and produce human-like social benefit–cost patterns in an 8B LLM.
-
Mixture of Cognitive Experts in Large Vision-Language Models
Routing CV experts into atomic evidence then Bloom-staged verbalization improves LVLM benchmarks and yields measurable query-conditioned reasoning traces.
-
Some Large Language Models Exhibit Consistent Risk Attitudes
Most of six LLMs show stable, cross-domain risk attitudes—consistent mappings from perceived risk to decisions—though these cluster in a narrower range than human risk preferences.
-
Large language models replicate and predict human cooperation across experiments in game theory
Llama-3.1-8B with a multi-step reasoning-and-filter prompt reproduces human cooperation rates across 121 dyadic games (MSD=0.031, r=0.89), outperforming Nash-equilibrium predictions (MSD=0.096, r=0.78).
-
The Mechanistic Emergence of Symbol Grounding in Language Models
Symbol grounding emerges in Transformers and state-space models through middle-layer 'aggregate' attention heads that connect environmental cues to words, but not in unidirectional LSTMs.
-
Quantifying Data Contamination in Psychometric Evaluations of LLMs
Across 21 LLMs and four standard psychology questionnaires, models recognize the items, know which trait each item measures, and can choose responses to hit a specified target score.
-
From Monolingual to Bilingual: Investigating Language Conditioning in Large Language Models for Psycholinguistic Tasks
Prompted language identity changes both the outputs and the internal layer representations of Llama-3.3-70B and Qwen2.5-72B on sound symbolism and word valence tasks.
-
Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues
In a fixed set of Japanese AI-to-AI counseling dialogues, expert ratings favored a structured prompt over a minimal one, while automated LLM ratings were reproducible but systematically more lenient.
-
Large Language Models are Near-Optimal Decision-Makers with a Non-Human Learning Behavior
Across uncertainty, risk, and set-shifting tasks, LLMs generally outperformed humans and neared optimality while exhibiting distinctly non-human decision-making processes.
-
Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities
An adversarial-agent framework adapted from human decision-making research reveals that GPT-4, Gemini-1.5, and DeepSeek-V3 are more rigid and easily manipulated in bandit and trust-game tasks than GPT-3.5 or humans.
-
Memorization and Knowledge Injection in Gated LLMs
MEGa injects episodic memories into separate gated LoRA adapters selected by embedding similarity, mitigating catastrophic forgetting and enabling recall, QA, and compositional questions on two datasets.
-
How Personality Traits Shape LLM Risk-Taking Behaviour
Using direct certainty-equivalent questions, the authors find GPT-4o behaves close to risk-neutral and that Openness-related personality prompts shift its risk parameters in a human-like direction, while GPT-4-Turbo d...
-
Kernels of Selfhood: GPT-4o shows humanlike patterns of cognitive consistency moderated by free choice
GPT-4o's ratings of Putin moved toward the valence of an essay it wrote, and this shift grew when the model was given an illusory free choice about the essay.
-
The "LLM World of Words" English free association norms generated by large language models
A new dataset of 3+ million free association responses from three LLMs, matched to human norms, with validation showing human-like semantic priming and gender bias patterns.
-
Humans are more gullible than LLMs in believing common psychological myths
Four LLMs believed 8 to 24 percent of 50 psychological myths, versus 51 to 63 percent for human students, and RAG generally, but not always, lowered belief rates.
-
A Conceptual Framework for AI Capability Evaluations
A descriptive conceptual framework with seven elements (target, task, subject, inputs, instance, measurement, result analysis) for systematizing analysis of AI capability evaluations.
-
Adapting to LLMs: How Insiders and Outsiders Reshape Scientific Knowledge Production
Researchers outside core AI fields became markedly more application-oriented, transdisciplinary, and socially accountable in their LLM-era papers, while AI insiders mainly responded by diversifying collaborations.
-
Evolutionary ecology of words
Words as organisms in an AI-judged battle royale evolve toward semantically 'strong' animal names, showing diverse and sometimes punctuated dynamics.
-
Psychologically Enhanced AI Agents
MBTI personality prompts measurably change how LLM agents write stories and play strategic games, with self-reflection before communication supporting cooperative behavior.
-
AI Awareness
A review arguing that AI awareness is a measurable, four-dimensional functional capacity (metacognition, self, social, situational) that current LLMs partially exhibit and that both improves AI and creates safety risks.
-
Evaluating Personality Traits in Large Language Models: Insights from Psychological Questionnaires
Across five questionnaires, five LLMs consistently self-report high Agreeableness, Openness, and Conscientiousness and low Neuroticism, but reported trait dominance is sensitive to how questionnaire scales are combined.
-
A Probabilistic WxChallenge Proposal
Two optional WxChallenge games let players bet confidence credits on ensemble-based thresholds or bins, with scores based on information gain over the baseline.
-
LLM-based Human Simulations Have Not Yet Been Reliable
Current LLM-based human simulations are not yet reliable; the paper reviews why and proposes a validation framework to improve consistency with real human behavior.
-
A Survey on Human-Centric LLMs
A review that sorts existing evidence on how well large language models imitate individual human skills and collective social dynamics into one taxonomy.
Discussion (0). Continue with ORCID to comment.