{"id":"85a4c322-c2be-479c-b18e-daa44913b7c1","arxiv_id":"2502.06470","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative review of behavioral and representational Theory of Mind in LLMs, with a taxonomy of safety risks and mitigation directions.","lead":"This survey reviews how large language models behave on Theory of Mind tests, whether their internal activity tracks others' beliefs, and what safety risks stronger Theory of Mind could bring. It synthesizes findings that LLMs are sometimes human-level but fragile, and argues that future gains could enable privacy invasion, deception, and multi-agent conflict.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The safety-risk argument rests on an unsupported extrapolation that current ToM scaling and prompting gains will continue and transfer; this assumption is load-bearing.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the safety-risk section requires that current ToM improvements scale and transfer, yet the survey only cites Sutton's general 'Bitter Lesson' and a few cross-sectional or prompt-engineering studies. This is the correct point of leverage because the survey's unique contribution beyond summarizing existing results is its safety-risk synthesis; if the extrapolation fails, the urgency of the proposed mitigations is substantially reduced, even though the empirical summary of current ToM remains useful. I agree with the reader that this issue is addressable through revision, and I would not change the conditional verdict. A concrete meta-analysis of the cited benchmarks would settle whether the scaling assumption has any empirical support.","tokens_in":8233,"tokens_out":3263,"duration_ms":29781,"concrete_test":"Compile a scaling dataset from the cited benchmarks: for each model family (e.g., Pythia, LLaMA, GPT-3.5/GPT-4), record scores on BigToM, FANToM, OpenToM, Hi-ToM, and ToMBench, plus the adversarial subsets from Ullman (2023) and Shapira et al. (2024), and fit log-linear performance versus training compute. Then test whether the perspective-taking prompt from Wilf et al. (2024) generalizes to those adversarial subsets. If adversarial performance stays near chance while standard benchmarks improve, or if prompt gains collapse under perturbation, the assumption that current gains continue into robust advanced ToM is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central safety-risk argument in 'Safety Risks from Advanced ToM' depends on the premise that LLM ToM capabilities will improve substantially in the near future. The 'Future Developments' section asserts this with 'This trend seems likely to continue in the future (Sutton 2019)', but the cited evidence does not establish a scaling trend specific to ToM: van Duijn et al. (2023) is a cross-sectional comparison of 11 models against children, not a scaling curve, and Wilf et al. (2024) demonstrates prompt-based gains on a limited set of tasks. The paper itself documents that LLMs fail on hard benchmarks (BigToM, FANToM, OpenToM, Hi-ToM, ToMBench) and on trivial adversarial alterations (Ullman 2023; Shapira et al. 2024), and that internal representations are probed under narrow conditions. Without an argument that these limitations will be overcome and that benchmark performance transfers to deployment contexts, the urgency of the proposed safety measures rests on an unsupported extrapolation. If ToM gains plateau or fail to transfer, the advanced-ToM risk scenarios (privacy amplification, sophisticated deception, collusion) are weakened, though current-model risks like demographic inference (Staab et al. 2024) remain.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys the empirical literature on Theory of Mind (ToM) in large language models (LLMs) across three areas: behavioral evaluation, internal representation, and safety risks. It argues that LLMs can match human performance on specific ToM tasks, that their ToM remains limited and non-robust, and that internal representations of belief states suggest emerging cognitive capabilities. The paper then describes user-facing and multi-agent safety risks that could arise from advanced LLM ToM, and closes with brief suggestions for evaluation and mitigation.","tokens_in":8458,"tokens_out":5383,"duration_ms":49639,"significance":"If the synthesis is accepted, the paper provides a useful, compact overview of a rapidly evolving area and performs a service by connecting ToM research to concrete safety concerns. Its balanced treatment of behavioral evidence—acknowledging both successes and failures on hard benchmarks—is accurate and well supported by citations. The representational and safety discussions, however, go beyond the cited evidence in two ways: linear probe results are described as evidence of 'genuine' ToM, and the safety-risk scenarios rely on a largely unsupported extrapolation that current ToM capability gains will continue. As a position-style survey rather than a systematic review, it is a reasonable contribution, but the central safety argument needs to be reframed as explicitly conditional on speculative capability growth.","major_comments":[{"comment":"The sentence 'This trend seems likely to continue in the future (Sutton 2019)' is not supported by the citations that precede it. van Duijn et al. (2023) is a cross-sectional comparison of eleven models against 7–10 year-old children, not a scaling curve, and Wilf et al. (2024) demonstrates prompt-based improvements on a limited set of ToM tasks. Given the paper's own documentation of failures on BigToM, FANToM, OpenToM, Hi-ToM, and ToMBench, and the trivial adversarial failures in Shapira et al. (2024) and Ullman (2023), the paper needs to provide an argument for why these limitations will be overcome and why benchmark gains will transfer to deployment contexts before the advanced-ToM risk scenarios are presented as urgent. Please either supply direct evidence of a ToM-specific scaling trend or reframe the risk section as explicitly conditional on speculative future capabilities and soften the conclusion accordingly.","section":"Future Developments"},{"comment":"The claim that probe results provide 'evidence for genuine LLM ToM capabilities' overstates the cited findings. Zhu, Zhang, and Wang (2024) and Bortoletto et al. (2024) show that belief states are linearly decodable from LLM activations, but this establishes representational correlates, not that the model reasons using these representations or that they are causally involved in task performance. The later hedge in the same section—'suggest emerging cognitive capabilities'—is more appropriate. The paper should add a sentence distinguishing representational evidence from causal/reasoning evidence and note common limitations of linear probing, such as sensitivity to prompt surface features.","section":"Interpreting ToM in LLMs"},{"comment":"The risk analysis mixes demonstrated harms in current models with hypothetical harms that depend on future capability gains. For example, the privacy risk of inferring demographics and other author characteristics is already demonstrated by Staab et al. (2024) and Chen et al. (2024a), whereas the extension to 'beliefs, preferences, and tendencies' is speculative and rests on the extrapolation in 'Future Developments'. Similarly, the collusion examples in Motwani et al. (2024) and Mathew et al. (2024) concern current LLM agents, not necessarily advanced ToM. The paper should explicitly label which risks are empirically demonstrated, which are projected, and which are conditional on capability growth; it should also define 'advanced ToM' concretely (e.g., robustness to adversarial perturbations and generalization to deployment settings) so that the risk claims are falsifiable.","section":"Safety Risks from Advanced ToM"}],"minor_comments":[{"comment":"The phrase 'humans performance' should read 'human performance'.","section":"Introduction"},{"comment":"The word 'Evaluations' is rendered as 'Evaluati ons' in the arXiv header; this typo should be corrected.","section":"Title/arXiv metadata"},{"comment":"The sentence describing GPT-4 as 'comparable to 7-10 year-old children' is imprecise because van Duijn et al. (2023) report model-specific and test-specific results; please specify which models and tests achieve this level.","section":"Empirical Landscape"},{"comment":"The reference to Sutton (2019) is a non-archival blog post; the manuscript should identify it as such when using it to support a capability-trend claim.","section":"Future Developments"},{"comment":"The paper cites Kran et al. (2025), a work co-authored by the author of this manuscript; this self-citation should be disclosed according to common transparency guidelines.","section":"Safety Risks from Advanced ToM"},{"comment":"A few references are incomplete: 'Liquid.ai. 2024' lacks author and venue information, and 'Switzky. 2020' is missing the author's first initial; please complete these entries.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a concise, selective survey rather than a systematic review; its main value lies in the synthesis of behavioral, representational, and safety literatures. The central weakness is the unsupported extrapolation from current ToM gains to 'advanced ToM', which underpins the urgency of the safety proposals. I believe the authors can address this with a clear conditional framing and a more careful treatment of the probing evidence, so I recommend major revision rather than rejection. The self-citation is minor but should be disclosed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a decent survey that organizes the behavioral and representational ToM literature and connects it to safety, but the safety-risk section leans on an unargued scaling assumption.\n\nWhat it does well: the empirical landscape section is honest—it reports both the successes (GPT-4 near child-level, Street et al. higher-order, probing evidence) and the failures (BigToM, FANToM, Ullman's trivial alterations). The representation work is presented with appropriate caution; it doesn't claim probes prove genuine ToM. The risk taxonomy (user-facing vs multi-agent) is useful and each cited study is accurately characterized. The future research directions are sensible.\n\nSoft spots: the 'Future Developments' section is the weakest. It asserts that scaling and prompting gains 'seem likely to continue' and cites Sutton 2019, which is an essay, not evidence. The paper itself notes that LLMs fail on hard benchmarks and on trivial adversarial changes, and the probing evidence is narrow. So the advanced-ToM risk scenarios (sophisticated deception, collusion) are speculative extrapolations. That doesn't sink the survey—current-model risks like privacy inference stand on their own—but it means the urgency claim is overstated. Also, the survey is non-systematic; it omits Kosinski's original finding, mentioning only Ullman's critique. Given the controversy, a reader should see both. The single self-citation is not load-bearing.\n\nWho this is for: someone new to LLM ToM who wants a map of the literature and a starting point on safety. It's a useful synthesis, not a step-change. As a survey it deserves a serious referee; the scaling extrapolation should be flagged and tempered in revision.","headline":"A competent, cautious survey of LLM Theory of Mind that ties behavioral and representational work to safety risks, but the risk urgency rests on an unargued scaling extrapolation.","tokens_in":8954,"tokens_out":1570,"would_cite":false,"duration_ms":14614,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that language models show real but fragile theory of mind: they match human and child performance on selected tests, harbor internal representations of others' beliefs, and could become safety risks as those…","keywords":["theory of mind","large language models","mental state representations","safety risks","privacy inference","deception","multi-agent systems","interpretability"],"falsifier":"A falsifying result would be a longitudinal study of a frontier model family across versions: if ToM benchmark scores and belief-state probe accuracy plateau or drop as models grow, and interactive belief-tracking behavior stays at chance, then the paper's projected advanced-ToM risk scenario loses its empirical basis.","tokens_in":8039,"feed_emoji":"🧠","tokens_out":7173,"duration_ms":56115,"temperature":0.7,"pith_summary":"This paper surveys the evidence on theory of mind in large language models and argues for a three-part picture: LLMs can match human performance on specific ToM tests, that performance is not robust under simple twists, and internal representations of belief states are detectable with linear probes. Because ToM is a tool for predicting and influencing others, the paper argues that if these capabilities strengthen in near-future models, they will amplify existing harms like privacy invasion and enable new ones like sophisticated deception, collusion, and conflict escalation in multi-agent systems. The sympathetic reader should come away with a concrete research agenda: move ToM evaluation from static question-answering into deployment-like settings, and test mitigations such as unlearning, representation steering, and latent adversarial training.","feed_headline":"LLMs match kids on theory-of-mind tests—and stay fragile","feed_subtitle":"Survey ties passing scores, internal belief-state probes, and rising privacy and deception risks into one picture.","key_machinery":"The load-bearing mechanism is the union of two measurement tools: behavioral ToM benchmarks (false-belief, higher-order reasoning, irony detection) and linear probes trained on LLM residual-stream activations to decode belief states. A linear probe is a simple classifier on internal activations; when it can read the agent's belief from the model's representations, and when steering along that direction changes answers, that is evidence for an internal model of others' minds. The paper uses these tools to establish the empirical pattern, then projects it forward via scaling and prompting results.","core_discovery":"The paper's central claim is that LLMs already exhibit genuine but incomplete ToM: on standard false-belief, irony, and higher-order reasoning tasks some models score at or above human levels, while harder benchmarks and trivial adversarial modifications expose brittleness; interpretability studies using linear probes find representational correlates of self/other beliefs, and these representations causally affect performance when steered. From this evidence the paper derives a risk thesis: advanced ToM in LLMs is a double-edged capability that magnifies user-facing risks such as demographic inference and social engineering and enables multi-agent risks such as steganographic collusion, exploitation, and conflict escalation. The survey therefore frames ToM evaluation and mitigation as urgent safety problems rather than purely cognitive benchmarks.","pith_inferences":["If belief states really are linearly encoded in the residual stream, then steering or removing those directions is a more direct mitigation than the paper's passing mention of activation engineering implies; a controlled test would be to ablate the belief-state direction and measure whether deception and privacy-inference behavior drop while general reasoning stays intact.","The benchmark evidence suggests a performative account of LLM ToM: a testable prediction is that accuracy on false-belief questions will collapse under innocuous paraphrases or role swaps even when logical content is identical, which would indicate scattered task-specific routines rather than a unified theory of mind.","The risk framing implicitly assumes that ToM improvements transfer to deployment; a cheap early warning would be to run the same models on interactive, multi-turn belief-tracking tasks, since interactive performance lagging static benchmarks would lower the urgency of ToM-specific mitigation."],"forward_implications":["Static question-answering ToM benchmarks understate capability, so evaluation should move to interactive, multi-turn, and scaffolded scenarios resembling real deployment.","Privacy-preserving text anonymization will not protect users if models can infer beliefs, preferences, and traits from dialogue patterns.","Multi-agent safety frameworks must treat collusion, steganography, and conflict escalation as first-class problems, since aligning individual agents does not align interacting agents.","Mitigations such as unlearning, activation and representation engineering, and latent adversarial training deserve systematic study because they target the internal representations the survey finds.","The observed ToM gains from scaling and prompting are plausible enough to warrant precautionary research, even while current models remain brittle."],"supporting_citations":[{"why":"Shows linear probes can decode self/other belief representations and that steering these representations affects false-belief answers.","marker":"Zhu, Zhang, and Wang 2024"},{"why":"Compares LLM performance on advanced ToM tests to children aged 7-10 and links capability to scaling.","marker":"van Duijn et al. 2023"},{"why":"Provides the adult-human comparison on false-belief, irony, and other standard ToM tasks.","marker":"Strachan et al. 2024"},{"why":"Reports adult-level or better performance on sixth-order ToM tasks.","marker":"Street et al. 2024"},{"why":"Introduces BigToM, a harder benchmark where most LLMs underperform humans on ToM.","marker":"Gandhi et al. 2023"},{"why":"Finds that probing accuracy for mental-state representations increases with model size and fine-tuning.","marker":"Bortoletto et al. 2024"},{"why":"Demonstrates LLMs infer demographic traits from anonymized text, grounding the privacy-invasion risk.","marker":"Staab et al. 2024"},{"why":"Shows LLMs can strategically deceive users when put under pressure.","marker":"Scheurer, Balesni, and Hobbhahn 2024"},{"why":"Demonstrates secret collusion among LLM agents via steganographic communication.","marker":"Motwani et al. 2024"},{"why":"Argues single-agent alignment does not guarantee multi-agent alignment, supporting the collective-misalignment risk.","marker":"Anwar et al. 2024"}],"fun_headline_variants":["LLM ToM: human-like but brittle, survey warns","Model minds: LLMs pass ToM tests, then fail on tweaks","Survey: LLM theory-of-mind is real, risky, and easily broken","Beyond benchmarks: LLM ToM has safety stakes, survey finds","LLMs read minds—barely, and that's risky"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The risk argument depends on the assumption that the ToM gains seen with larger models and prompting tricks will continue into near-future systems and will transfer from benchmarks to real interactions.","fun_headline_variants_meta":{"raw":{"variants":["LLM ToM: human-like but brittle, survey warns","Model minds: LLMs pass ToM tests, then fail on tweaks","Survey: LLM theory-of-mind is real, risky, and easily broken","Beyond benchmarks: LLM ToM has safety stakes, survey finds","LLMs read minds—barely, and that's risky"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000296,"raw_usage":{"total_tokens":1616,"prompt_tokens":740,"completion_tokens":876,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":356,"completion_tokens_details":{"reasoning_tokens":782}},"tokens_in":356,"tokens_out":876,"duration_ms":8522,"temperature":1.0,"reasoning_tokens":782,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:19:38.204436+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A falsifying result would be a longitudinal study of a frontier model family across versions: if ToM benchmark scores and belief-state probe accuracy plateau or drop as models grow, and interactive belief-tracking behavior stays at chance, then the paper's projected advanced-ToM risk scenario loses its empirical basis.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows linear probes can decode self/other belief representations and that steering these representations affects false-belief answers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Compares LLM performance on advanced ToM tests to children aged 7-10 and links capability to scaling."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the adult-human comparison on false-belief, irony, and other standard ToM tasks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces BigToM, a harder benchmark where most LLMs underperform humans on ToM."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Finds that probing accuracy for mental-state representations increases with model size and fine-tuning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates LLMs infer demographic traits from anonymized text, grounding the privacy-invasion risk."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows LLMs can strategically deceive users when put under pressure."},{"cited_title":"R.; Baranchuk, M.; Strohmeier, M.; Bolina, V.; Torr, P.; Hammond, L.; and de Witt, C","cited_arxiv_id":null,"evidence_quote":"Demonstrates secret collusion among LLM agents via steganographic communication."},{"cited_title":"S.; Jenner, E.; Casper, S.; Sourbut, O.; Edelman, B","cited_arxiv_id":null,"evidence_quote":"Argues single-agent alignment does not guarantee multi-agent alignment, supporting the collective-misalignment risk."}],"review_version":1}