Pith. sign in

REVIEW 14 cited by

FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15421 v3 pith:UVNIMJFA submitted 2023-10-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords benchmarkfantomllmsmindmodelsreasoningtheoryanswering
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Theory of mind (ToM) evaluations currently focus on testing models using passive narratives that inherently lack interactivity. We introduce FANToM, a new benchmark designed to stress-test ToM within information-asymmetric conversational contexts via question answering. Our benchmark draws upon important theoretical requisites from psychology and necessary empirical considerations when evaluating large language models (LLMs). In particular, we formulate multiple types of questions that demand the same underlying reasoning to identify illusory or false sense of ToM capabilities in LLMs. We show that FANToM is challenging for state-of-the-art LLMs, which perform significantly worse than humans even with chain-of-thought reasoning or fine-tuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. R^3-VQA: "Read the Room" by Video Social Reasoning

    cs.CV 2025-05 conditional novelty 7.0 of 10

    R3-VQA is a new real-world video benchmark on which the best tested model, GPT-4o, scores 83% on generated questions but only 54% on human-written ones, while humans score 91% and 80%.

  3. PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting

    cs.CL 2026-06 unverdicted novelty 6.5 of 10

    PerspectiveGap benchmark shows LLMs achieve only 14.9% average pass rate on multi-agent orchestration prompting tasks, with GPT-5.5 at 62%.

  4. LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A new open-source library and benchmark, xRouteBench, evaluates LLM routers on a shared cost-aware protocol across text, memory, vision, time-series, and personalized tasks.

  5. Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Reasoning-enabled LLMs show more robust performance on Theory of Mind tests under prompt and task perturbations, supporting a robustness-based reading of recent gains.

  6. The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.

  7. Beyond Context to Cognitive Appraisal: Emotion Reasoning as a Theory of Mind Benchmark for Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LLMs are moderately accurate at emotion reasoning in a new appraisal-based theory-of-mind benchmark but rely on System-1-like heuristics and struggle to link specific appraisals to emotions.

  8. SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SocialMaze is a six-task benchmark that claims to evaluate LLM social reasoning along deep reasoning, dynamic interaction, and information uncertainty dimensions.

  9. Position: Theory of Mind Benchmarks are Broken for Large Language Models

    cs.AI 2024-12 conditional novelty 6.0 of 10

    The paper proposes that LLM theory-of-mind evaluation should measure functional adaptation to partners, not just literal prediction of their behavior, and shows the two can diverge sharply in simple games.

  10. Does It Make Sense to Speak of Introspection in Large Language Models?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    The authors argue that an untrained large language model inferring its own sampling temperature from the style of its own output qualifies as a minimal, consciousness-free form of introspection.

  11. Code Simulation as a Proxy for High-order Tasks in Large Language Models

    cs.LG 2025-02 conditional novelty 5.0 of 10

    LLM performance on naturalistic reasoning tasks tracks performance on equivalent Python code simulation, but the effect is partly driven by pattern matching and memorization rather than faithful execution.

  12. Boundless Socratic Learning with Language Games

    cs.AI 2024-11 conditional novelty 5.0 of 10

    A position paper claiming recursive self-improvement in a closed language-only system can reach arbitrary capability, and proposing language games as the mechanism.

  13. Multi-ToM: Evaluating Multilingual Theory of Mind Capabilities in Large Language Models

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A multilingual and culturally adapted Theory of Mind benchmark for seven languages, with evaluations of six LLMs showing lower performance in low-resource languages and accuracy drops when cultural details are added.

  14. Intentionally Unintentional: GenAI Exceptionalism and the First Amendment

    cs.CY 2025-06 conditional novelty 4.0 of 10

    The paper argues that GenAI outputs are not First Amendment-protected speech because the models have no communicative intent, so users have no speech right to receive them and regulators need not meet strict scrutiny.

Pith tools