REVIEW 4 major objections 4 minor 18 references
When researchers say an AI has a theory of mind, they are describing behavioral prediction and pattern matching, not genuine mental states; the paper argues isolated ToM tests ask the wrong question and should be replaced by studies of huma
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 12:37 UTC pith:FJH7WT34
load-bearing objection Useful reorientation toward interaction-level evaluation, but the simulation/genuine distinction is asserted, not defended; the practical point survives without it. the 4 major comments →
When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking About?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the current discourse conflates simulation with experience. When LLMs pass false-belief and other ToM tasks, they are matching surface patterns from massive human discourse about mental states; the paper calls this behavioral mimicry, distinct from human ToM, which develops through embodied interaction and is entangled with desires, motivations, and hot/cold cognition. To support this, the paper points to evidence such as performance collapse when task wording is obfuscated, the poor transfer from explicit ToM inference to implicit behavioral judgment in GPT-4, and the prevalence of belief-only measures, with 75.5% of ToM evaluations ignoring emotions, desir
What carries the argument
The argument turns on a distinction between simulation and experience, and on a mutual-theory-of-mind approach as the alternative evaluative lens. The simulation-experience distinction does the negative work: it says that behavioral success on false-belief tasks, because it lacks embodied, motivated, social grounding, cannot count as genuine ToM. The mutual-theory-of-mind approach does the positive work: it models human-AI interaction as three iteratively shaping elements—interpretation, feedback, and mutuality—in which humans anthropomorphize the AI and the AI statistically models the human, without requiring genuine mental states on either side. The dynamic, team-level emphasis is the pape
Load-bearing premise
The argument depends on assuming that real theory of mind requires embodied experience and motivational states, so a disembodied program that merely predicts behavior cannot be credited with it; if a purely behavioral or functionalist definition of ToM is correct, the central distinction collapses.
What would settle it
A concrete study would match two AI systems on ToM benchmark scores and then have humans collaborate with them on a shared task. If the higher-scoring system consistently produced better team outcomes, or if static scores predicted interaction quality better than human understanding of the agent did, the paper's central claim that isolated ToM tests are irrelevant would be called into question. Alternatively, an LLM that displayed genuinely motivation-driven errors—mistakes caused by desires overriding logic rather than by edge cases in training data—would challenge the claim that AI lacks mot
If this is right
- ToM benchmark scores on LLMs should be interpreted as measures of pattern reproduction, not evidence of mental-state attribution; reporting them as cognition overstates what is known.
- Static false-belief tasks cannot predict how an AI will behave in human-AI collaboration, so benchmark results should be treated as weak evidence for deployment.
- Evaluation of AI social capability should move to interactive, dynamic tasks measuring mutual adaptation, feedback, and team outcomes.
- Improving the human's understanding of the AI may matter more for team performance than improving the AI's independent ToM capability.
- Design efforts should aim at supporting mutual adaptation given human anthropomorphism and AI statistical modeling, rather than chasing higher ToM test scores.
Where Pith is reading between the lines
- If the paper's distinction between simulation and experience is accepted, a functionalist account of ToM—where any system that predicts and explains behavior counts—is the main rival view; the paper does not refute it, but its argument implies that adopting functionalism would make the AI-ToM question purely behavioral.
- The mutual-ToM approach suggests a concrete testable extension: compare team performance under conditions where static ToM scores are high but human understanding of the agent is low; the paper's position predicts the latter dominates.
- A natural measurement extension would apply dynamic, time-series measures of team states to human-AI interaction, treating mutual understanding as an emergent property of the interaction rather than a property of either agent.
- One might extend the argument to other cognitive attributes: if ToM is not validly testable in isolation for LLMs, the same critique applies to claims of AI reasoning, planning, or creativity based on static benchmarks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that when researchers attribute theory of mind (ToM) or mental models to LLMs, they are actually describing behavioral prediction and statistical pattern matching, not genuine mental states. The authors critique the use of isolated human cognitive tests for LLMs, citing benchmark validity problems, training-data dependence, and the gap between explicit ToM inference and implicit application. They propose a shift toward studying mutual ToM in human-AI interaction, drawing on the Mutual ToM framework of Wang and Goel (2022) and dynamic team-level measurements such as those of Grimm et al. (2023). The paper is explicitly a position piece, not an empirical study, and its contribution is meant to be a reorientation of evaluation practice.
Significance. If the argument is accepted, the paper would redirect ToM evaluation away from decontextualized benchmark scores and toward interaction-level constructs, which is practically relevant given the rapid proliferation of ToM benchmarks for LLMs. The paper usefully surfaces construct-validity concerns (Wang et al., 2025), highlights empirical discrepancies (Gu et al., 2024; Zhang et al., 2024), and offers a positive agenda for human-AI team research. However, the central distinction between simulation and genuine ToM is asserted rather than defended, and the paper engages only minimally with the philosophy-of-mind literature. The strength of the paper is therefore largely conditional on whether the authors can ground their definitional premise; as it stands, the main claim is defensible but not fully established.
major comments (4)
- [§2, para. 1; Abstract] The central claim—that LLM ToM success is 'behavioral mimicry' rather than 'genuine mental state attribution'—rests on an asserted definition of genuine ToM as grounded in embodied experience and intrinsic motivational states. The sentence 'In a narrow technical sense, AI systems that track beliefs and predict behaviors could be said to have “ToM”' grants the functionalist reading, then retracts it without argument. Standard functionalist accounts in philosophy and cognitive science characterize ToM as the capacity to attribute mental states in order to predict and explain behavior, independent of implementation. The paper never engages these accounts. This is load-bearing: if functionalism is correct, the simulation/genuine distinction dissolves, and the conclusion in §4 that researchers are 'observing sophisticated behavioral prediction, not genuine mental state attribution' is unsuppo
- [§1, paras. 1 and 3] The manuscript makes several unsupported empirical assertions that are used to support the core thesis. In §1, 'if you listen to an authentic interaction between two humans and an interaction between a human and an AI agent, you can tell a difference' is offered without any data or protocol. Later in §1, 'LLM performance plummets dramatically' with obfuscated names and the claim that self-verification makes performance worse are stated without citations or effect sizes. These are not peripheral illustrations; they purport to show that LLM behavior is pattern matching rather than reasoning. Please provide evidence or reframe these as conjectures with concrete tests that could confirm or falsify them.
- [§2, para. 3] The discussion of 'emergence' conflates ToM with phenomenal consciousness. The Zhu et al. (2024) result—that LLMs internally represent different agents' beliefs—is countered by saying that structured representations do not prove 'consciousness exists.' But a functionalist account of ToM does not require consciousness; internal belief representations are exactly the kind of state a functionalist would count as a mental model. The paper needs to separate the question 'does the LLM have ToM?' from 'is the LLM conscious?' and argue why consciousness is necessary for genuine ToM. Without this, the dismissal of representational evidence is a non sequitur.
- [§3, Mutual ToM] The positive proposal is under-specified. The paper says we should test 'system dynamics' and 'mutual adaptation' instead of isolated ToM tests, but it does not say what variables should be measured, what experimental designs would be appropriate, or what pattern of results would count as evidence for mutual ToM. The reference to Grimm et al. (2023) gives an example of dynamical measurement, but the mapping from team-resilience measures to mutual-ToM constructs is not drawn. This matters because the paper's second central claim—that the field should shift to interaction-level evaluation—needs an operationalizable framework, otherwise it risks replacing one vague construct with another.
minor comments (4)
- [Title page] Typo: 'Univeristy' should be 'University.' Also the abstract contains an incomplete sentence: 'the entire testing paradigm may be flawed in applying individual human cognitive tests to AI systems, but assessing human cognition directly in the moment of human-AI interaction' appears to be missing a phrase such as 'and we should instead be assessing.'
- [References [3] and §1] The text cites 'Kanishk et al. (2023)' for BigToM, but the reference [3] is Gandhi, Fränken, Gerstenberg, and Goodman. Please correct the in-text citation to match the reference list.
- [§1, para. 1] The characterization of Kosinski (2023) as arguing that 'individual artificial neurons in LLMs function like Chinese rooms' is not an accurate summary of that paper, which argues that ToM may have spontaneously emerged in LLMs. Please verify the attribution.
- [§1, Sally-Anne test] The Sally-Anne false-belief test is attributed to Scassellati (2001), which is a dissertation on robot ToM. The original task is from Baron-Cohen, Leslie, and Frith (1985). Please cite the primary source or clarify why Scassellati is the relevant reference.
Circularity Check
No circular derivation; self-citation is peripheral and non-load-bearing.
full rationale
This is a position paper and literature synthesis rather than a derivation with equations, fitted parameters, or predictions. The central claim—that researchers often use 'ToM' to mean behavioral prediction and that LLMs simulate rather than possess embodied ToM—is advanced as a philosophical/definitional argument, supported by citations to external benchmarks (e.g., Gu et al., Strachan et al., Wang et al.). No step reduces by construction to its inputs: the SimpleToM accuracy figures (90%, 50%, 15%) are external evidence, not fitted parameters, and no benchmark result is renamed as a prediction. The only author-overlapping reference is Grimm et al. 2023, used as an example of dynamic team measurement ('human-autonomy team (HAT) resilience studies employing dynamic measurements (e.g., Grimm et al., 2023) can serve as a valuable reference'), and it is not load-bearing for the paper's ToM thesis. The disputed premise that genuine ToM requires embodied experience and intrinsic motivational states is asserted rather than defended; that is an argumentative weakness (a correctness/characterization risk), not circularity under the definitions used here, because the paper does not derive a quantitative prediction from that premise. Honest finding: no significant circularity.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Human ToM is grounded in embodied experience, motivated reasoning, and genuine understanding.
- domain assumption LLMs are 'n-gram models on steroids' performing 'universal approximate retrieval' rather than reasoning.
- domain assumption Static third-person ToM tests lack construct validity for predicting dynamic human-AI interaction.
- domain assumption The Mutual ToM framework is the appropriate alternative unit of analysis.
- domain assumption Behavioral success on ToM tasks is not evidence of mental state attribution.
read the original abstract
When researchers claim AI systems possess ToM or mental models, they are fundamentally discussing behavioral predictions and bias corrections rather than genuine mental states. This position paper argues that the current discourse conflates sophisticated pattern matching with authentic cognition, missing a crucial distinction between simulation and experience. While recent studies show LLMs achieving human-level performance on ToM laboratory tasks, these results are based only on behavioral mimicry. More importantly, the entire testing paradigm may be flawed in applying individual human cognitive tests to AI systems, but assessing human cognition directly in the moment of human-AI interaction. I suggest shifting focus toward mutual ToM frameworks that acknowledge the simultaneous contributions of human cognition and AI algorithms, emphasizing the interaction dynamics, instead of testing AI in isolation.
Reference graph
Works this paper leans on
-
[1]
Chen, Z., Wu, J., Zhou, J., Wen, B., Bi, G., Jiang, G., Cao, Y ., Hu, M., Lai, Y ., Xiong, Z., & Huang, M. (2024). ToMBench: Benchmarking theory of mind in large language models.arXiv preprint arXiv:2402.15052
Pith/arXiv arXiv 2024
-
[2]
G., Xia, T., Mao, H., Thumiger, N., Desai, A., Stoica, I., Klimovic, A., Neubig, G., & Gonzalez, J
Cuadron, A., Li, D., Ma, W., Wang, X., Wang, Y ., Zhuang, S., Liu, S., Schroeder, L. G., Xia, T., Mao, H., Thumiger, N., Desai, A., Stoica, I., Klimovic, A., Neubig, G., & Gonzalez, J. E. (2025). The danger of overthinking: Examining the reasoning-action dilemma in agentic tasks.arXiv preprint arXiv:2502.08235. 3 When Researchers Say Mental Model/Theory o...
Pith/arXiv arXiv 2025
-
[3]
P., Gerstenberg, T., & Goodman, N
Gandhi, K., Fränken, J. P., Gerstenberg, T., & Goodman, N. (2023). Understanding social reasoning in language models with language models. InAdvances in Neural Information Processing Systems(V ol. 36, pp. 13518–13529)
2023
-
[4]
A., Gorman, J
Grimm, D. A., Gorman, J. C., Cooke, N. J., Demir, M., & McNeese, N. J. (2023). Dynamical measurement of team resilience.Journal of Cognitive Engineering and Decision Making, 17(4), 351–382
2023
-
[5]
Gu, Y ., Tafjord, O., Kim, H., Moore, J., Bras, R. L., Clark, P., & Choi, Y . (2024). SimpleTom: Exposing the gap between explicit tom inference and implicit tom application in LLMs.arXiv preprint arXiv:2410.13648
arXiv 2024
-
[6]
Kambhampati, S. (2024). Can large language models reason and plan?Annals of the New York Academy of Sciences, 1534(1), 15–18
2024
-
[7]
Kosinski, M. (2023). Theory of mind may have spontaneously emerged in large language models.arXiv preprint arXiv:2302.02083
Pith/arXiv arXiv 2023
-
[8]
Kunda, Z. (1990). The case for motivated reasoning.Psychological Bulletin, 108(3), 480–498
1990
-
[9]
Premack, D., & Woodruff, G. (1978). Does the chimpanzee have a theory of mind?Behavioral and Brain Sciences, 1(4), 515–526
1978
-
[10]
Scassellati, B. M. (2001).F oundations for a Theory of Mind for a Humanoid Robot(Doctoral dissertation, Massachusetts Institute of Technology)
2001
-
[11]
Sellars, W. (1956). Empiricism and the philosophy of mind.Minnesota Studies in the Philosophy of Science, 1(19), 253–329
1956
-
[12]
Strachan, J. W. A., Albergo, D., Borghini, G., Pansardi, O., Scaliti, E., Gupta, S., Saxena, K., Rufo, A., Panzeri, S., Manzi, G., Graziano, M. S. A., & Becchio, C. (2024). Testing theory of mind in large language models and humans.Nature Human Behaviour, 8(7), 1285–1295
2024
-
[13]
Wang, Q., & Goel, A. K. (2022). Mutual theory of mind for human-AI communication.arXiv preprint arXiv:2210.03842
Pith/arXiv arXiv 2022
-
[14]
Wang, Q., Zhou, X., Sap, M., Forlizzi, J., & Shen, H. (2025). Rethinking theory of mind benchmarks for LLMs: Towards a user-centered perspective.arXiv preprint arXiv:2504.10839
Pith/arXiv arXiv 2025
-
[15]
M., & Liu, D
Wellman, H. M., & Liu, D. (2004). Scaling of theory-of-mind tasks.Child Development, 75(2), 523–541
2004
-
[16]
Wilf, A., Lee, S. S., Liang, P. P., & Morency, L. P. (2023). Think twice: Perspective-taking improves large language models’ theory-of-mind capabilities.arXiv preprint arXiv:2311.10227
Pith/arXiv arXiv 2023
-
[17]
Zhang, S., Wang, X., Zhang, W., Chen, Y ., Gao, L., Wang, D., Zhang, W., Wang, X., & Wen, Y . (2024). Mutual theory of mind in human-AI collaboration: An empirical study with LLM-driven AI agents in a real-time shared workspace task.arXiv preprint arXiv:2409.08811
Pith/arXiv arXiv 2024
-
[18]
Zhu, W., Zhang, Z., & Wang, Y . (2024). Language models represent beliefs of self and others.arXiv preprint arXiv:2402.18496. 4
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.