Pith. sign in

REVIEW 6 cited by

CogBench: a large language model walks into a psychology lab

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.18225 v1 pith:7CMMHPTE submitted 2024-02-28 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords llmsbehaviorcogbenchmodelshumanlanguagelargemetrics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have significantly advanced the field of artificial intelligence. Yet, evaluating them comprehensively remains challenging. We argue that this is partly due to the predominant focus on performance metrics in most benchmarks. This paper introduces CogBench, a benchmark that includes ten behavioral metrics derived from seven cognitive psychology experiments. This novel approach offers a toolkit for phenotyping LLMs' behavior. We apply CogBench to 35 LLMs, yielding a rich and diverse dataset. We analyze this data using statistical multilevel modeling techniques, accounting for the nested dependencies among fine-tuned versions of specific LLMs. Our study highlights the crucial role of model size and reinforcement learning from human feedback (RLHF) in improving performance and aligning with human behavior. Interestingly, we find that open-source models are less risk-prone than proprietary models and that fine-tuning on code does not necessarily enhance LLMs' behavior. Finally, we explore the effects of prompt-engineering techniques. We discover that chain-of-thought prompting improves probabilistic reasoning, while take-a-step-back prompting fosters model-based behaviors.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A 14-domain, 91-subskill, three-layer cognitive taxonomy organizes LLM capability research, and a 15,934-paper mapping shows attention concentrated in Language-Semantic Competence and Reasoning while social, moral, an...

  2. Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLMs struggle to use passive observations for reverse engineering, but active intervention improves performance, largely through the process of generating queries rather than the data obtained.

  3. H2HTalk: Evaluating Large Language Models as Emotional Companion

    cs.CL 2025-07 conditional novelty 5.0 of 10

    H2HTalk is a new 4,650-scenario benchmark that scores LLM emotional companions on dialogue, memory, and itinerary planning, and finds models struggle with implicit needs and long-horizon memory.

  4. DinoCompanion: An Attachment-Theory Informed Multimodal Robot for Emotionally Responsive Child-AI Interaction

    cs.AI 2025-06 reject novelty 5.0 of 10

    A child-facing robot trained with an attachment-theory-informed preference optimization is claimed to beat GPT-4o and Gemini-2.5-Pro on a new ten-competency benchmark, though evaluation and derivation issues undermine...

  5. Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM Agents

    cs.MA 2025-06 reject novelty 5.0 of 10

    Shapley-Coop asks LLM agents to negotiate prices for contributions based on Shapley value reasoning, improving cooperation and reward fairness in three multi-agent tasks.

  6. AI Agent Behavioral Science

    q-bio.NC 2025-06 conditional novelty 4.0 of 10

    AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.

Pith tools