{"id":"a9d5ce8f-ee66-40c2-9e4c-fc6b3ad4d08a","arxiv_id":"2505.00026","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of recent story-based Theory of Mind benchmarks and enhancement strategies for large language models, organized by mental state coverage and method type.","lead":"This paper surveys how Theory of Mind is measured and improved in large language models. It reviews recent story-based benchmarks and enhancement methods, serving as a reference for researchers in the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first broad survey' novelty claim is not supported by any reported search protocol and is vulnerable to prior overlapping surveys; if it falls, the contribution weakens to an incremental synthesis.","rationale":"Reading in good faith, the survey is a coherent, well-organized synthesis and its taxonomy could be genuinely useful. The most load-bearing risk is not that some individual benchmark or method is missing, but that the headline novelty claim—'first broad survey' of both evaluation and enhancement—is a factual assertion with no reported search methodology. The reader's weakest assumption (representativeness of coverage) overlaps with this, but my concern is more specifically about the uniqueness claim. The concrete test is inexpensive and decisive: a reproducible literature search before the paper's revision date either surfaces a competing survey or does not. If it does, the paper should be accepted conditionally with the novelty claim softened; if it does not, the central claim stands. I therefore keep the reader's CONDITIONAL verdict unchanged rather than moving to ACCEPT or REJECT.","tokens_in":21347,"tokens_out":3403,"duration_ms":34188,"concrete_test":"Use the arXiv API or Google Scholar with a documented query, e.g., title/abstract containing 'theory of mind' AND 'large language model' AND ('survey' OR 'review'), restricted to papers published before 2025-08-25 (the v2 date). Manually screen the results for any prior work that reviews both evaluation benchmarks and enhancement methods for LLM ToM. If one or more such surveys exist, the Section 1 'first' wording must be softened; if none exist, the novelty claim survives this check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution statement (Section 1, 'Broad Survey') is that it is 'the first broad survey that addresses both the evaluation and enhancement of LLMs' ToM capabilities.' This is a factual novelty claim, not a conceptual argument. The paper reports no systematic literature search: no query strings, databases, inclusion/exclusion criteria, or search date are given in the Introduction or Limitations. The Limitations explicitly narrow the scope to story-based benchmarks, exclude purely spatial scenarios such as BIB, and relegate interactive benchmarks mostly to an appendix, so the 'first broad survey' assertion is effectively a claim about coverage of the whole field. Given the fast-moving 2023-2025 arXiv literature, it is plausible that a prior or concurrent survey already covers both assessment and enhancement. If so, the 'first' claim is false, and the main contribution becomes 'a useful survey' rather than a gap-filling first. The reader's coverage concern is related but slightly different; I would pin the risk specifically on the unsupported novelty assertion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey paper reviews theory of mind (ToM) in large language models from two angles: evaluation benchmarks and enhancement strategies. For evaluation, it focuses on story-based benchmarks from 2023-2024, both text-only and multimodal, describing their design, mental state coverage, and evolution. For enhancement, it categorizes recent methods into those relying solely on prompting and those incorporating additional techniques such as fine-tuning or inverse planning, then proposes future research directions. The paper claims to be the first broad survey covering both assessment and enhancement of LLMs' ToM capabilities.","tokens_in":21534,"tokens_out":2674,"duration_ms":29265,"significance":"If its coverage is representative and its taxonomy sound, this survey would be a useful quick reference for researchers entering the area: it organizes eleven benchmarks in a comparative table, summarizes enhancement approaches at a glance, and flags open problems such as higher-order reasoning, multimodal evaluation, and active/agentic ToM. The mental-state coverage table and the clear separation of prompt-based and fine-tuning-based methods are helpful organizational devices. However, the paper's value depends on its novelty and coverage claims, and its 'assessment' component is descriptive rather than evaluative, containing no performance numbers or critical synthesis of empirical findings.","major_comments":[{"comment":"The claim that this is 'the first broad survey that addresses both the evaluation and enhancement of LLMs' ToM capabilities' is not supported by any reported literature search protocol. The paper does not state which databases were queried, what search strings or inclusion/exclusion criteria were used, or the date of the search. The Limitations paragraph additionally narrows the coverage to story-based benchmarks, explicitly excluding purely spatial scenarios such as BIB and relegating interactive benchmarks to an appendix. Given these restrictions, the 'broad survey' and 'first' assertions are stronger than the evidence provided. Please either report a systematic search and justify the coverage decisions, or temper the novelty claim by situating the paper relative to existing surveys (beyond Ma et al. 2023b).","section":"Section 1, Contribution 'Broad Survey' and Limitations"},{"comment":"The paper's abstract and title promise an assessment of LLMs' ToM capabilities, but the body never reports any evaluation results, model-family comparisons, or even a summary of accuracy findings across the benchmarks and enhancement methods. For example, Section 4 opens by stating that 'most evaluations... highlight the limitations of LLMs' without citing any specific numbers or aggregated conclusions, and the benchmark descriptions in Section 3 contain no performance data. A survey of this kind does not need a full meta-analysis, but it should at least synthesize the directional findings (e.g., order effects, false-belief vs. true-belief gaps, differences between text-only and multimodal settings). Without this, the 'assessment' component is a catalog rather than an assessment.","section":"Sections 3-4 and the Abstract"},{"comment":"The paper treats 'goals' in MMToM-QA and MuMA-ToM as equivalent to 'intentions' in ATOMS. This is a reasonable decision, but it is also a substantive interpretive choice that affects the mental-state coverage table, and it is stated only in a table footnote. Please justify this mapping, or at least acknowledge that it is an assumption, because a reader comparing Table 1 across benchmarks could otherwise overinterpret the coverage comparison.","section":"Section 2 and Table 1"}],"minor_comments":[{"comment":"The phrase 'in-depth analysis' appears in both the abstract and the contributions list; consider varying the wording to avoid redundancy.","section":"Abstract and Section 1"},{"comment":"The appendix heading 'Abilities in Theory of Mind Space (A TOMS)' contains a spacing error; it should be 'ATOMS'.","section":"Section 2 and Appendix B"},{"comment":"The sentence 'These task, with adjustments to elements like the characters, the container, or the objects involved, forms the basis...' has a subject-verb agreement error ('These task' and 'forms').","section":"Appendix A.3"},{"comment":"In the SYMBOLICTOM panel, the notation 'BBob,Alice' appears without clear subscript formatting; please use a consistent notation such as B_Bob,Alice or define it in the caption.","section":"Section 4.1, Figure 2"},{"comment":"Table 3 omits all citations 'due to width constraints.' Since the table is a key reference comparison, please restore the citations or provide a companion table with full references.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"This is a survey with no new experiments or datasets. The main risk is the unsupported 'first broad survey' novelty claim, which should be verified against the recent literature. If the authors can substantiate or soften that claim and add a minimal empirical synthesis, the paper would be acceptable as a reference resource for the community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful survey, well organized, and mostly accurate. The 'first broad survey' claim is the one thing I'd push back on, but it's a fixable overstatement rather than a load-bearing flaw.\n\nWhat's new: the pairing of benchmark review with enhancement-strategy review in one place, with a mental-state coverage table (Table 1) and a method comparison table (Table 2). That genuinely helps someone entering the area. The benchmark descriptions are detailed and the trend analysis (order, generation, context, questions, mental states) is sensible. The enhancement section covers SYMBOLICTOM, SIMTOM, PercepToM, TIMETOM, ToM-LM, BIP-ALM, LIMP in enough depth to convey how they relate. The paper also flags real limitations of each method, which is more than many surveys do.\n\nSoft spots: the novelty claim is not supported. No search protocol, no inclusion criteria, no databases or dates. The Limitations section explicitly excludes spatial benchmarks like BIB and relegates interactive work to an appendix, so calling it the 'first broad survey' of both evaluation and enhancement is a stretch given how fast this literature moves. I don't know of a direct competitor that covers the same 2023–2025 ground, so I'd frame it as 'first to our knowledge' with the scope caveat—which the paper mostly does, but the contribution bullet says 'broad survey' without the caveat. Also, the 'assessment' component is descriptive: no performance numbers, so it's more a map than an evaluation. That's fine for a survey, but the word 'assessment' in the title oversells it slightly. The interactive benchmarks and pre-LLM methods are relegated to an appendix; that's a scope decision, not a flaw, but it undercuts the 'broad' claim further. The self-citations (Qin et al. 2023, Parmar et al. 2024) are not load-bearing; they're context and a future-direction suggestion, so no circularity problem.\n\nIf you're new to LLM ToM and want a structured entry point, this paper is worth reading. It won't change your research agenda if you're already in the area, but it's reliable as a reference. I'd send it to a serious referee mostly to check whether the 'first' claim holds and to push for a methodology note on literature selection. That's a light revision, not a rejection.\n\nRecommendation: accept for review; it's a solid reference with an overstatement that the authors can fix.","headline":"A competent, well-organized survey of LLM ToM benchmarks and enhancement methods whose only real flaw is an overstated 'first broad survey' novelty claim and no search protocol.","tokens_in":22013,"tokens_out":645,"would_cite":true,"duration_ms":7808,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims to be the first broad survey covering both the evaluation and the enhancement of theory of mind in large language models.","keywords":["theory of mind","large language models","benchmarks","evaluation","enhancement strategies","prompt engineering","mental states","survey"],"falsifier":"A reader could falsify the first-survey claim by finding a published survey from before 2025 that already covers both evaluation and enhancement of theory of mind in large language models. The taxonomy's completeness could be tested by rerunning the review with purely spatial benchmarks counted in; if the claimed trends—conversation-based contexts and multimodal expansion—survive the addition, the map holds, and if not, the selection was not representative.","tokens_in":21183,"feed_emoji":"🧠","tokens_out":7621,"duration_ms":72990,"temperature":0.7,"pith_summary":"This survey tries to organize a fast-moving research area: do large language models really have a theory of mind, and if not, what can be done about it? The paper's central claim is that it is the first broad survey to cover both evaluation and enhancement of LLMs' theory-of-mind capabilities, bringing the two strands under one taxonomy. For evaluation, it reviews story-based benchmarks from 2023–2024, sorting them by text-only versus multimodal input and by which mental states they cover. For enhancement, it separates prompt-only strategies from methods that add fine-tuning, symbolic reasoning, or inverse planning. A sympathetic reader would take the contribution to be the map itself: a structured comparison that makes it easy to see how benchmarks and improvement methods have evolved and where gaps remain.","feed_headline":"One survey maps both tests and fixes for LLM theory of mind","feed_subtitle":"It organizes benchmarks by mental states and enhancement tactics by method, showing where models still fall short.","key_machinery":"The paper's organizing device is the ATOMS inventory of seven mental states—beliefs, intentions, desires, emotions, knowledge, percepts, and non-literal communications—which it uses as a common yardstick to compare every benchmark. A second load-bearing device is the notion of \"order\" of belief attribution, from first-order (what a character believes) up to fourth-order (what one character thinks another believes about a third's belief), which lets the survey measure benchmark difficulty and track evolution. On the enhancement side, the central distinction is prompt-only methods versus methods that add fine-tuning, model checking, or inverse planning. These axes together carry the survey's main work: turning a scattered literature into comparison tables and trend statements.","core_discovery":"On its own terms, the paper establishes a structured map of current theory-of-mind research on LLMs. It claims that story-based benchmarks have rapidly evolved from first-order belief questions on synthetic narratives toward higher-order beliefs (up to fourth order), multi-turn conversational contexts, broader mental states, and multimodal household scenarios, and it supports this with a comparison of thirteen benchmarks. It further claims that enhancement strategies fall into two families: prompt-only methods—belief graphs, perspective-taking, perception-aware context extraction, temporal belief-state chains—and methods that add fine-tuning or external machinery such as semantic parsing with a model checker and inverse multi-agent planning. The paper's conclusion is that despite these benchmarks and methods, LLMs still lack dependable theory-of-mind abilities, and consistent assessment remains difficult because theory of mind cannot be captured by a limited set of questions.","pith_inferences":["The comparison table implies a convergence that the paper does not state: both benchmark design and enhancement methods are gravitating to the same bottleneck—tracking what each character perceives and when—so a diagnostic benchmark that isolates perception errors could predict which enhancement method will help.","The paper's observation that methods are mostly tested on multiple-choice formats, if taken further, suggests that open-ended evaluation would likely compress the performance differences between prompt-only and fine-tuned methods.","A natural extension the authors leave implicit: apply prompt-only techniques like perspective-taking and temporal belief-state chains to emotions, desires, and non-literal communication, using a benchmark with full mental-state coverage, to test whether the methods generalize beyond beliefs."],"forward_implications":["Researchers can use the mental-state and order columns to choose benchmarks that match the capability they want to test, instead of relying on popularity.","The documented trend implies that future benchmarks will continue toward conversational and multimodal settings, so methods tested only on narrative multiple-choice questions may not transfer.","Because prompt-only enhancement methods are pipelines, their ceiling is set by the first perception-tracking step; improving that step should improve downstream answers across methods.","The paper's reading of the evidence says single-benchmark accuracies should not be read as proof of general theory of mind; robust evaluation needs multiple mental states and question formats.","Story-based passive benchmarks are, by the paper's account, insufficient; evaluating LLMs as active agents in interactive settings is a stated next step."],"supporting_citations":[{"why":"Provides the earlier benchmark review and situated-evaluation framing that this survey extends.","marker":"(Ma et al., 2023b)"},{"why":"Supplies the ATOMS taxonomy of seven mental states used to categorize every benchmark.","marker":"(Beaudoin et al., 2020)"},{"why":"Defines ToMi, the widely used story-based benchmark and the origin of story accuracy, which most enhancement methods test on.","marker":"(Le et al., 2019)"},{"why":"Introduces HI-TOM, the benchmark with up to fourth-order belief questions that anchors the survey's \"orders\" trend.","marker":"(Wu et al., 2023)"},{"why":"Introduces FANTOM, the conversation-based benchmark that defines illusory ToM and shifts evaluation from narratives to dialogue.","marker":"(Kim et al., 2023)"},{"why":"Introduces TOMBENCH, built from scratch and covering nearly all ATOMS mental states, used to argue that benchmarks are expanding in mental-state coverage.","marker":"(Chen et al., 2024)"},{"why":"Proposes SYMBOLICTOM, the multi-character belief-graph prompting method that anchors the prompt-only enhancement family.","marker":"(Sclar et al., 2023)"},{"why":"Proposes SIMTOM, the perspective-taking prompt strategy that the survey presents as a second prompt-only route to improving ToM.","marker":"(Wilf et al., 2024)"},{"why":"Proposes BIP-ALM and the MMToM-QA multimodal benchmark, anchoring the fine-tuning and inverse-planning enhancement family.","marker":"(Jin et al., 2024)"},{"why":"Introduces MuMA-ToM and LIMP, the multimodal multi-agent benchmark and inverse multi-agent planning method that extend the survey's multimodal and higher-order coverage.","marker":"(Shi et al., 2025)"}],"fun_headline_variants":["LLM theory of mind: benchmarks evolve, enhancements fall short","Survey maps ToM tests and fixes, finds models still lacking","From belief chains to model checkers: LLM ToM still unreliable","LLMs can't read minds yet: survey of ToM benchmarks and boosts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the benchmarks and enhancement methods selected for review fairly represent the field; the paper does not document a systematic search protocol or inclusion criteria, so if important work—for example, spatial-scenario benchmarks it explicitly excludes—were weighed equally, the taxonomy and conclusions could shift.","fun_headline_variants_meta":{"raw":{"variants":["LLM theory of mind: benchmarks evolve, enhancements fall short","Survey maps ToM tests and fixes, finds models still lacking","From belief chains to model checkers: LLM ToM still unreliable","LLMs can't read minds yet: survey of ToM benchmarks and boosts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00085,"raw_usage":{"total_tokens":3644,"prompt_tokens":842,"completion_tokens":2802,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":2726}},"tokens_in":458,"tokens_out":2802,"duration_ms":21963,"temperature":1.0,"reasoning_tokens":2726,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:06:19.575945+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could falsify the first-survey claim by finding a published survey from before 2025 that already covers both evaluation and enhancement of theory of mind in large language models. The taxonomy's completeness could be tested by rerunning the review with purely spatial benchmarks counted in; if the claimed trends—conversation-based contexts and multimodal expansion—survive the addition, the map holds, and if not, the selection was not representative.","supporting_citations":[],"review_version":1}