REVIEW 6 cited by
CogBench: a large language model walks into a psychology lab
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have significantly advanced the field of artificial intelligence. Yet, evaluating them comprehensively remains challenging. We argue that this is partly due to the predominant focus on performance metrics in most benchmarks. This paper introduces CogBench, a benchmark that includes ten behavioral metrics derived from seven cognitive psychology experiments. This novel approach offers a toolkit for phenotyping LLMs' behavior. We apply CogBench to 35 LLMs, yielding a rich and diverse dataset. We analyze this data using statistical multilevel modeling techniques, accounting for the nested dependencies among fine-tuned versions of specific LLMs. Our study highlights the crucial role of model size and reinforcement learning from human feedback (RLHF) in improving performance and aligning with human behavior. Interestingly, we find that open-source models are less risk-prone than proprietary models and that fine-tuning on code does not necessarily enhance LLMs' behavior. Finally, we explore the effects of prompt-engineering techniques. We discover that chain-of-thought prompting improves probabilistic reasoning, while take-a-step-back prompting fosters model-based behaviors.
Forward citations
Cited by 6 Pith papers
-
From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models
A 14-domain, 91-subskill, three-layer cognitive taxonomy organizes LLM capability research, and a 15,934-paper mapping shows attention concentrated in Language-Semantic Competence and Reasoning while social, moral, an...
-
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
LLMs struggle to use passive observations for reverse engineering, but active intervention improves performance, largely through the process of generating queries rather than the data obtained.
-
H2HTalk: Evaluating Large Language Models as Emotional Companion
H2HTalk is a new 4,650-scenario benchmark that scores LLM emotional companions on dialogue, memory, and itinerary planning, and finds models struggle with implicit needs and long-horizon memory.
-
DinoCompanion: An Attachment-Theory Informed Multimodal Robot for Emotionally Responsive Child-AI Interaction
A child-facing robot trained with an attachment-theory-informed preference optimization is claimed to beat GPT-4o and Gemini-2.5-Pro on a new ten-competency benchmark, though evaluation and derivation issues undermine...
-
Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents
Shapley-Coop asks LLM agents to negotiate prices for contributions based on Shapley value reasoning, improving cooperation and reward fairness in three multi-agent tasks.
-
AI Agent Behavioral Science
AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.
Discussion (0). Sign in to comment.