Pith. sign in

REVIEW 7 cited by

LLMs achieve adult human performance on higher-order theory of mind tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.18870 v2 pith:TP3I2DZR submitted 2024-05-29 cs.AI cs.CLcs.HC

LLMs achieve adult human performance on higher-order theory of mind tasks

classification cs.AI cs.CLcs.HC
keywords humanllmsperformanceadulthigher-ordermindtheoryadult-level
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

This paper examines the extent to which large language models (LLMs) have developed higher-order theory of mind (ToM); the human ability to reason about multiple mental and emotional states in a recursive manner (e.g. I think that you believe that she knows). This paper builds on prior work by introducing a handwritten test suite -- Multi-Order Theory of Mind Q&A -- and using it to compare the performance of five LLMs to a newly gathered adult human benchmark. We find that GPT-4 and Flan-PaLM reach adult-level and near adult-level performance on ToM tasks overall, and that GPT-4 exceeds adult performance on 6th order inferences. Our results suggest that there is an interplay between model size and finetuning for the realisation of ToM abilities, and that the best-performing LLMs have developed a generalised capacity for ToM. Given the role that higher-order ToM plays in a wide range of cooperative and competitive human behaviours, these findings have significant implications for user-facing LLM applications.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts

    cs.CL 2025-06 conditional novelty 7.0

    PuzzleWorld benchmark reveals state-of-the-art AI models solve only 18% of complex puzzlehunt problems with 40% stepwise accuracy, matching novices but trailing enthusiasts, while fine-tuning on traces yields modest gains.

  2. We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

    cs.AI 2024-07 accept novelty 7.0

    WE-MATH benchmark reveals most LMMs rely on rote memorization for visual math while GPT-4o has shifted toward knowledge generalization.

  3. Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration

    cs.CL 2025-06 conditional novelty 6.0

    Tomcat, an LLM agent using few-shot chain-of-thought or commonsense prompting, matches human performance on intent accuracy, action optimality, and planning optimality in a dynamic collaborative task.

  4. From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning

    cs.LG 2026-06 unverdicted novelty 5.0

    Thinking-RFT improves Theory of Mind accuracy by 6% over SFT on shortcut-free datasets, with 10% gains on higher-order reasoning and better generalization to new domains.

  5. A Survey of Large Language Models for Perception and Measurement of Human Psychology

    cs.CY 2026-05 unverdicted novelty 5.0

    A survey proposing a three-pillar framework to evaluate LLMs as tools for measuring latent psychological constructs and reviewing applications in personality and mental health.

  6. OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind

    cs.AI 2026-05 unverdicted novelty 5.0

    OSCToM uses RL-guided generation with an extended DSL and surrogate models to create nested belief conflict tasks, raising FANToM accuracy from 0.2% to 76% while being 6x more efficient.

  7. Network Effects and Agreement Drift in LLM Debates

    cs.SI 2026-04 unverdicted novelty 4.0

    LLM agents in controlled network debates show agreement drift toward specific opinion positions, requiring separation of structural effects from LLM biases before using them as human behavioral proxies.