An argument paper reframes LLM explainability as an embodied, situated practice based on Dourish and enactivist cognition, identifying ontological obstacles in internal explanations and advocating affordance-based designs.
arXiv preprint arXiv:2504.01698 , year=
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4representative citing papers
Improvements in LLM Theory of Mind on static benchmarks do not reliably improve performance in dynamic, first-person human-AI interactions across goal-oriented and experience-oriented tasks.
Thinking-RFT improves Theory of Mind accuracy by 6% over SFT on shortcut-free datasets, with 10% gains on higher-order reasoning and better generalization to new domains.
MindZero is a self-supervised RL framework that trains MLLMs for online Theory of Mind reasoning by rewarding mental-state hypotheses that best explain observed actions via a planner, then distills this into fast inference.
citing papers explorer
-
Embodied Explainability and Ontological Obstacles: Why We Struggle to Explain the Answers of Large Language Models (LLMs)
An argument paper reframes LLM explainability as an embodied, situated practice based on Dourish and enactivist cognition, identifying ontological obstacles in internal explanations and advocating affordance-based designs.
-
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations
Improvements in LLM Theory of Mind on static benchmarks do not reliably improve performance in dynamic, first-person human-AI interactions across goal-oriented and experience-oriented tasks.
-
From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning
Thinking-RFT improves Theory of Mind accuracy by 6% over SFT on shortcut-free datasets, with 10% gains on higher-order reasoning and better generalization to new domains.
-
MindZero: Learning Online Mental Reasoning With Zero Annotations
MindZero is a self-supervised RL framework that trains MLLMs for online Theory of Mind reasoning by rewarding mental-state hypotheses that best explain observed actions via a planner, then distills this into fast inference.