REVIEW 8 cited by
FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Theory of mind (ToM) evaluations currently focus on testing models using passive narratives that inherently lack interactivity. We introduce FANToM, a new benchmark designed to stress-test ToM within information-asymmetric conversational contexts via question answering. Our benchmark draws upon important theoretical requisites from psychology and necessary empirical considerations when evaluating large language models (LLMs). In particular, we formulate multiple types of questions that demand the same underlying reasoning to identify illusory or false sense of ToM capabilities in LLMs. We show that FANToM is challenging for state-of-the-art LLMs, which perform significantly worse than humans even with chain-of-thought reasoning or fine-tuning.
Forward citations
Cited by 8 Pith papers
-
Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex
Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.
-
PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting
PerspectiveGap benchmark shows LLMs achieve only 14.9% average pass rate on multi-agent orchestration prompting tasks, with GPT-5.5 at 62%.
-
Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
Reasoning-enabled LLMs show more robust performance on Theory of Mind tests under prompt and task perturbations, supporting a robustness-based reading of recent gains.
-
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
A new interactive language-game benchmark shows LLMs lag behind simple word-embedding baselines and that newer reasoning models regress on theory-of-mind tasks.
-
Beyond Context to Cognitive Appraisal: Emotion Reasoning as a Theory of Mind Benchmark for Large Language Models
LLMs are moderately accurate at emotion reasoning in a new appraisal-based theory-of-mind benchmark but rely on System-1-like heuristics and struggle to link specific appraisals to emotions.
-
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
SocialMaze is a six-task benchmark that claims to evaluate LLM social reasoning along deep reasoning, dynamic interaction, and information uncertainty dimensions.
-
Does It Make Sense to Speak of Introspection in Large Language Models?
The authors argue that an untrained large language model inferring its own sampling temperature from the style of its own output qualifies as a minimal, consciousness-free form of introspection.
-
Intentionally Unintentional: GenAI Exceptionalism and the First Amendment
The paper argues that GenAI outputs are not First Amendment-protected speech because the models have no communicative intent, so users have no speech right to receive them and regulators need not meet strict scrutiny.
Discussion (0). Sign in to comment.