Pith. sign in

REVIEW 4 cited by

OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.06044 v3 pith:CFGRQ7U5 submitted 2024-02-08 cs.AI cs.CL

classification cs.AIcs.CL
keywords mentalstatescharactersn-tomopentompsychologicalquestionsworld
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural Theory-of-Mind (N-ToM), machine's ability to understand and keep track of the mental states of others, is pivotal in developing socially intelligent agents. However, prevalent N-ToM benchmarks have several shortcomings, including the presence of ambiguous and artificial narratives, absence of personality traits and preferences, a lack of questions addressing characters' psychological mental states, and limited diversity in the questions posed. In response to these issues, we construct OpenToM, a new benchmark for assessing N-ToM with (1) longer and clearer narrative stories, (2) characters with explicit personality traits, (3) actions that are triggered by character intentions, and (4) questions designed to challenge LLMs' capabilities of modeling characters' mental states of both the physical and psychological world. Using OpenToM, we reveal that state-of-the-art LLMs thrive at modeling certain aspects of mental states in the physical world but fall short when tracking characters' mental states in the psychological world.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study

    cs.AI 2025-06 conditional novelty 6.0 of 10

    ChatGPT 4o and o1-mini chose more risk-averse lottery options than real respondents in Sydney, Hong Kong, Dhaka, and Nanjing; o1-mini was closer to humans, and Chinese prompts widened the gap.

  2. The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.

  3. From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Attention heads in multimodal LLMs linearly encode agents' beliefs, and steering those heads along probe-derived directions improves first- and second-order belief accuracy on the new GridToM benchmark.

  4. H2HTalk: Evaluating Large Language Models as Emotional Companion

    cs.CL 2025-07 conditional novelty 5.0 of 10

    H2HTalk is a new 4,650-scenario benchmark that scores LLM emotional companions on dialogue, memory, and itinerary planning, and finds models struggle with implicit needs and long-horizon memory.

Pith tools