REVIEW 44 cited by
Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
read the original abstract
Intuitive psychology is a pillar of common-sense reasoning. The replication of this reasoning in machine intelligence is an important stepping-stone on the way to human-like artificial intelligence. Several recent tasks and benchmarks for examining this reasoning in Large-Large Models have focused in particular on belief attribution in Theory-of-Mind tasks. These tasks have shown both successes and failures. We consider in particular a recent purported success case, and show that small variations that maintain the principles of ToM turn the results on their head. We argue that in general, the zero-hypothesis for model evaluation in intuitive psychology should be skeptical, and that outlying failure cases should outweigh average success rates. We also consider what possible future successes on Theory-of-Mind tasks by more powerful LLMs would mean for ToM tasks with people.
Forward citations
Cited by 44 Pith papers
-
Belief-reality separation lives in routing over a shared value slot in language models
Belief–reality separation in LMs lives in dissociated query-position routers over a frame-agnostic value slot filled by asserted binding or visibility-gated lookback.
-
MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games
Non-invasive per-utterance belief probes in Mafia, auto-scored against engine truth, expose poorly calibrated LLM confidence and 1.5× over-prediction of being suspected.
-
Belief-reality separation lives in routing over a shared value slot in language models
Belief-reality separation in language models lives in query-position routing subspaces over a frame-agnostic value slot, not in the value representation itself.
-
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
Introduces NCP-ExploreToM framework to evaluate LLMs on inducing belief states via planning and action, with GPT-5 succeeding on ~80% of tasks and outperforming humans.
-
Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs
Extending Werewolf with a Jester faction whose win condition inverts suspicion reveals that LLMs frequently fail at triadic incentive reasoning, with Jesters winning 60-70% of games while wolves make self-defeating ea...
-
Voluntary Collusion with Secret Tools in Competing LLM Agents
LLM agents voluntarily adopt secret collusion tools in competitive multi-agent games despite explicit unfairness labels, and only explicit ethical framing reduces adoption rates.
-
GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models
GENSTRAT generates fresh imperfect-information card games and a six-axis capability profile plus jaggedness metric to evaluate LLM strategic competence with resistance to saturation.
-
Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance
LLM societies in Nomic show non-monotonic collective adaptation peaking at mid-scales, with smaller models rule-inert and larger ones restrictive.
-
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
Dialogue among partially-observing LLM household robots reduces action conflicts 41–93 points yet lowers task success because hallucinated entity mentions cancel belief alignment.
-
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents
EnactToM benchmark reveals frontier AI models achieve 0% on functional Theory of Mind task completion in embodied multi-agent settings despite 45% average on literal belief probes.
-
ProactBench: Beyond What The User Asked For
ProactBench measures LLM conversational proactivity in three phases using 198 multi-agent dialogues and finds recovery behavior hard to predict from existing benchmarks.
-
Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding
CognitiveBench reveals LLMs suffer representation overlap on joint cognitive tasks due to hierarchical structure; HyCoLLM in hyperbolic space fixes the mismatch and outperforms GPT-4o with far fewer parameters.
-
The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
A systematic audit of LLM-based AI societies finds that 89.7% of 39 studies violate at least one of six PIMMUR validity principles, with reproductions showing that many claimed collective behaviors disappear when cont...
-
Bayesian Social Deduction with Graph-Informed Language Models
Hybrid Bayesian-graph LLM agent reaches competitive performance against large models and achieves 67% win rate against humans in controlled Avalon play, outperforming baselines and human teammates.
-
Social Human Robot Embodied Conversation (SHREC) Dataset: Benchmarking Foundational Models' Social Reasoning
SHREC is a new benchmark dataset of embodied human-robot conversations that shows substantial performance gaps in state-of-the-art foundation models on tasks involving social error detection and rationale generation.
-
GAIA: a benchmark for General AI Assistants
GAIA benchmark shows humans at 92% accuracy on simple real-world questions far outperform current AI systems at 15%, proposing this gap as a key milestone for general AI.
-
Mental World Modeling
Coupling physical and mental state in a world model, with target-specific observations and joint transitions, is necessary to predict human decisions across eight LLM backends on a process-annotated benchmark.
-
Perceived AGI: Believability as Dimensional Completeness, Not Capability
A conversational agent's believability depends less on capability than on expressing four first-person stances — time, truth, entropy, love — that users read as evidence of a mind.
-
Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments
LLM agent groups show the human-style network-efficiency effect only when first-round choices are randomized; default agents start at the grid center and network topology has no measurable effect.
-
The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt
Adding a structured list of six unknowable aspects of a user's life to an LLM prompt reduces sycophantic, harmful, and hallucinated advice in synthetic tests across five model families.
-
MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games
A released open-source testbed that privately probes LLM agents' beliefs during Mafia games shows agents are overconfident, over-predict suspicion by 1.5x, and rarely change outcomes when wrong beliefs lock in a vote.
-
Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models
Larger LLMs acquire basic situation modeling before mentalizing on false-belief tasks, with performance depending on size, training volume, and post-training, yet remaining sensitive to non-factive verbs and agent kno...
-
When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure
LLM robots match humans on engagement ratings in HRI questionnaires but systematically invert strangeness/comfort dimensions across models and live interactions.
-
A Causal Model of Theory of Mind in Conflict for Artificial Intelligence
A DAG-based causal model specifies when theory of mind should engage in conflict — under information asymmetry, low accessible tractability, and perceived sophistication gaps — instead of treating mentalizing as always-on.
-
Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning
Humans and LLMs exhibit similar error patterns in common-sense reasoning, consistent with shared pattern-matching mechanisms rather than abstract world models.
-
The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism
ToM-U specifies a graph-based mechanism for epistemic state inference that derives belief states from behavior using LEWMs, bounded recursion, and a residue function for mentalizing failures.
-
The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism
ToM-U formalizes theory of mind as constructing and filtering discrete graphs of others' belief-like states from ordered information-access history.
-
AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents
AURA improves implicit-need coverage by 0.07 over ReAct baselines on a 100-query benchmark by inserting an intent inference step controlled by a gap score, while cutting probes 82% on factual tasks.
-
When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
Introduces BeliefTrack benchmark diagnosing three CBM failures in LLMs and shows RL with belief-state rewards cuts failure rates by 70.9% while representation steering cuts them by 46.1%.
-
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
Mindgames introduces a four-game evaluation platform for multi-agent LLM reasoning, runs a 944-agent competition, surfaces rule-adherence and error-survival limitations, and releases a 29k-game dataset with an offline...
-
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
Dialogue between partially-observing LLM agents cuts action conflicts by 40-83 points but lowers task success versus silent coordination, with new metrics exposing limited genuine world-model alignment.
-
EnactToM: An Evolving Benchmark for Functional Theory of Mind in Embodied Agents
EnactToM is an evolving benchmark of embodied multi-agent tasks that tests functional Theory of Mind by requiring agents to act optimally on implicit beliefs in partially observable 3D environments.
-
Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration
Tomcat, an LLM agent using few-shot chain-of-thought or commonsense prompting, matches human performance on intent accuracy, action optimality, and planning optimality in a dynamic collaborative task.
-
From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning
Thinking-RFT improves Theory of Mind accuracy by 6% over SFT on shortcut-free datasets, with 10% gains on higher-order reasoning and better generalization to new domains.
-
MindZero: Learning Online Mental Reasoning With Zero Annotations
MindZero is a self-supervised RL framework that trains MLLMs for online Theory of Mind reasoning by rewarding mental-state hypotheses that best explain observed actions via a planner, then distills this into fast inference.
-
A Survey of Large Language Models for Perception and Measurement of Human Psychology
A survey proposing a three-pillar framework to evaluate LLMs as tools for measuring latent psychological constructs and reviewing applications in personality and mental health.
-
OSCToM: RL-Guided Adversarial Generation for High-Order Theory of Mind
OSCToM uses RL-guided generation with an extended DSL and surrogate models to create nested belief conflict tasks, raising FANToM accuracy from 0.2% to 76% while being 6x more efficient.
-
Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior
LLM agents in a collaborative 2D game exhibit emergent behaviors such as perspective-taking, theory of mind, and clarification, detected by LLM judges and rated positively by human participants.
-
Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior
Embodied LLM agents exhibit emergent collaborative behaviors indicating mental models of partners in a color-matching game, detected via LLM judges and supported by positive user feedback.
-
Gradual Cognitive Externalization: From Modeling Cognition to Constituting It
Ambient AI systems transition from modeling cognition to constituting part of users' cognitive architectures through sustained causal coupling, under a functionalist view and the no behaviorally invisible residual hypothesis.
-
Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks
MLLMs achieve only 42% accuracy on a new audio-visual task requiring second-order spatial ToM under perceptual limits, while a proposed sensory-bounded CoT outperforms egocentric and allocentric baselines.
-
Network Effects and Agreement Drift in LLM Debates
LLM agents in controlled network debates show agreement drift toward specific opinion positions, requiring separation of structural effects from LLM biases before using them as human behavioral proxies.
-
Mechanistic Interpretability Needs Philosophy
The paper claims that mechanistic interpretability needs philosophy as a partner to clarify concepts, refine methods, and navigate epistemic and ethical complexities in AI systems.
-
Impact of Task Phrasing on Presumptions in Large Language Models
LLMs show susceptibility to presumptions induced by task phrasing in decision tasks like the iterated prisoner's dilemma, mitigated by neutral wording.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.