Pith. sign in

REVIEW 12 cited by

ToMBench: Benchmarking Theory of Mind in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15052 v2 pith:H2F2R2KI submitted 2024-02-23 cs.CL cs.AI

ToMBench: Benchmarking Theory of Mind in Large Language Models

classification cs.CL cs.AI
keywords llmstombenchevaluationmindtheoryabilitieslanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Theory of Mind (ToM) is the cognitive capability to perceive and ascribe mental states to oneself and others. Recent research has sparked a debate over whether large language models (LLMs) exhibit a form of ToM. However, existing ToM evaluations are hindered by challenges such as constrained scope, subjective judgment, and unintended contamination, yielding inadequate assessments. To address this gap, we introduce ToMBench with three key characteristics: a systematic evaluation framework encompassing 8 tasks and 31 abilities in social cognition, a multiple-choice question format to support automated and unbiased evaluation, and a build-from-scratch bilingual inventory to strictly avoid data leakage. Based on ToMBench, we conduct extensive experiments to evaluate the ToM performance of 10 popular LLMs across tasks and abilities. We find that even the most advanced LLMs like GPT-4 lag behind human performance by over 10% points, indicating that LLMs have not achieved a human-level theory of mind yet. Our aim with ToMBench is to enable an efficient and effective evaluation of LLMs' ToM capabilities, thereby facilitating the development of LLMs with inherent social intelligence.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs

    cs.CL 2026-07 conditional novelty 6.0

    HSS-Synth generates 230k instruction-tuning samples for 14 humanities/social-science fields and reports state-of-the-art fine-tuning results on 16 benchmarks.

  2. Mental World Modeling

    cs.CL 2026-07 conditional novelty 6.0

    Coupling physical and mental state in a world model, with target-specific observations and joint transitions, is necessary to predict human decisions across eight LLM backends on a process-annotated benchmark.

  3. Pigeonholing: how bad prompts hurt models, causing collapse and mistakes

    cs.CL 2026-06 conditional novelty 6.0

    Unintentionally bad contexts (user suggestions or prior wrong assistant answers) cause LLMs to repeat errors, lose diversity, and flip stances, worsening with turns; RLVR on synthetic errors recovers 43–60% of the drop.

  4. Social World Model for Lifelong Social Intelligence

    cs.AI 2026-06 unverdicted novelty 6.0

    The Social World Model supplies a five-dimension decomposition and closed-loop training loop that lets a 7B open model match Gemini 3 Flash on social metrics while showing zero forgetting on ASCENT-Bench.

  5. Reinforcing Human Behavior Simulation via Verbal Feedback

    cs.LG 2026-05 unverdicted novelty 6.0

    DITTO uses RL with verbal feedback to train LLMs for human behavior simulation, reporting 36% average gains over base models and outperforming GPT-5.4 on 6 of 10 SOUL benchmark tasks.

  6. Are you with me? A Framework for Detecting Mental Model Discrepancies in Task-Based Team Dialogues

    cs.AI 2026-05 unverdicted novelty 6.0

    A new framework identifies four mental model discrepancy types in team dialogues and demonstrates they carry predictive signals for future misalignments via uniform-weighted historical counts.

  7. Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation

    cs.AI 2026-04 unverdicted novelty 6.0

    SiPeR improves recommendation accuracy and response quality in situated conversations by estimating scene transitions and performing Bayesian inverse inference with multimodal LLMs.

  8. AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum

    cs.AI 2026-04 unverdicted novelty 6.0

    AIT Academy introduces a tripartite curriculum for AI agents across natural science, humanities, and social science domains, with reported gains of 15.9 points in security and 7 points in social reasoning under specif...

  9. Pigeonholing: how bad prompts hurt models, causing collapse and mistakes

    cs.CL 2026-06 unverdicted novelty 5.0

    Bad contexts in LLM conversations cause error repetition, mode collapse, and opinion flipping with 38-40% performance drops that worsen over turns, mitigated by RLVR trained with synthetic errors.

  10. OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind

    cs.AI 2026-05 unverdicted novelty 5.0

    OSCToM uses RL-guided generation with an extended DSL and surrogate models to create nested belief conflict tasks, raising FANToM accuracy from 0.2% to 76% while being 6x more efficient.

  11. LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue

    cs.CL 2025-09 reject novelty 5.0

    LLMs can imitate mental-state annotation in team dialogue but systematically err on spatial reasoning and prosodic cues, per a six-dialogue CReST pilot.

  12. When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?

    cs.HC 2025-10 conditional novelty 3.0

    Researchers' claims of AI theory of mind are really about behavioral prediction, so AI evaluation should shift from isolated cognitive tests to human-AI interaction.