Pith. sign in

REVIEW 8 cited by

FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15421 v3 pith:UVNIMJFA submitted 2023-10-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords benchmarkfantomllmsmindmodelsreasoningtheoryanswering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Theory of mind (ToM) evaluations currently focus on testing models using passive narratives that inherently lack interactivity. We introduce FANToM, a new benchmark designed to stress-test ToM within information-asymmetric conversational contexts via question answering. Our benchmark draws upon important theoretical requisites from psychology and necessary empirical considerations when evaluating large language models (LLMs). In particular, we formulate multiple types of questions that demand the same underlying reasoning to identify illusory or false sense of ToM capabilities in LLMs. We show that FANToM is challenging for state-of-the-art LLMs, which perform significantly worse than humans even with chain-of-thought reasoning or fine-tuning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting

    cs.CL 2026-06 unverdicted novelty 6.5 of 10

    PerspectiveGap benchmark shows LLMs achieve only 14.9% average pass rate on multi-agent orchestration prompting tasks, with GPT-5.5 at 62%.

  3. Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Reasoning-enabled LLMs show more robust performance on Theory of Mind tests under prompt and task perturbations, supporting a robustness-based reading of recent gains.

  4. The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.

  5. Beyond Context to Cognitive Appraisal: Emotion Reasoning as a Theory of Mind Benchmark for Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LLMs are moderately accurate at emotion reasoning in a new appraisal-based theory-of-mind benchmark but rely on System-1-like heuristics and struggle to link specific appraisals to emotions.

  6. SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SocialMaze is a six-task benchmark that claims to evaluate LLM social reasoning along deep reasoning, dynamic interaction, and information uncertainty dimensions.

  7. Does It Make Sense to Speak of Introspection in Large Language Models?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The authors argue that an untrained large language model inferring its own sampling temperature from the style of its own output qualifies as a minimal, consciousness-free form of introspection.

  8. Intentionally Unintentional: GenAI Exceptionalism and the First Amendment

    cs.CY 2025-06 conditional novelty 4.0 of 10

    The paper argues that GenAI outputs are not First Amendment-protected speech because the models have no communicative intent, so users have no speech right to receive them and regulators need not meet strict scrutiny.

Pith tools