{"id":"54c7b8bf-b27a-45c8-8b5c-dcb4f833b671","arxiv_id":"2506.09450","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A new benchmark, UniToMBench, is proposed for evaluating Theory of Mind in LLMs, but its evaluation results are mixed and do not substantiate the claimed improvements.","lead":"This paper introduces UniToMBench, a benchmark that merges two existing Theory of Mind test suites and adds over 1,000 new hand-written scenarios with multi-turn and evolving story formats. The authors report evaluation results on five LLMs, but the data do not consistently support their claim that perspective-taking improves ToM performance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed improvement from SimToM is not supported by the paper's own numbers: results are mixed, many tasks decline, and no repeated runs, error bars, or significance tests are reported.","rationale":"The reader's weakest assumption is dataset construct validity, namely the absence of human baselines, inter-annotator agreement, and comparison with established psychometric tests. I agree that this is serious, but the more load-bearing point for the paper's stated contribution is that the experimental design cannot support the causal improve claim even if the dataset were perfect. The paper's own tables show inconsistent and frequently negative SimToM effects; without repeated runs and paired significance tests, the reported numbers cannot distinguish systematic gains from stochastic variation. This is especially acute because the conclusion's wording overstates the data: on the custom evolving-story task, all models except GPT-4o decline, and on the TOMBENCH tasks GPT-4o-Mini declines on most tasks. A benchmark could still be useful as a descriptive instrument, but the paper frames UniToMBench as improving ToM, so the unsupported effect claim is central. I recommend keeping the reader's REJECT verdict, with the caveat that the rejection is driven by lack of statistical evidence rather than solely by dataset-validation worries.","tokens_in":12230,"tokens_out":4767,"duration_ms":49577,"concrete_test":"Obtain the 1,025 custom items and the TOMBENCH items, run each baseline and SimToM condition K=10 times under the paper's temperature, and apply McNemar's exact test per model and task to the paired item-level correctness matrix. If the majority of task comparisons are not significant at p<0.05 or the point estimates are negative, the enhancement claim fails. Report the resulting effect sizes and confidence intervals instead of single percentages.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is that the central enhancement claim is not established by the reported experiments. Section 5 reports a single set of accuracy percentages at temperature 0.7, with no repeated trials, confidence intervals, or significance tests. On the custom dataset (Table 6), SimToM leaves Evolving Story accuracy essentially unchanged for GPT-4o (89.9 to 90.1) and lowers it for GPT-4o-Mini (88.6 to 85.1), GPT-3.5 (69.5 to 55.8), Llama 3 8B (82.2 to 74.0), and Gemma 2 27B (79.5 to 73.0). On Multi-Interaction tasks, only GPT-4o-Mini improves (55.8 to 65.8); GPT-4o drops from 71.2 to 63.0. The TOMBENCH results in Table 1 are also inconsistent: GPT-4o-Mini declines on seven of eight tasks with SimToM, and GPT-3.5 swings from 34.7 to 10.0 on Unexpected Outcome while jumping from 24.3 to 68.9 on Hinting. Thus the conclusion that integrating multi-interaction tasks, evolving stories, and the SimToM framework enhances LLM performance is not a robust reading of the data. A few-point gap in either direction is within plausible sampling noise at temperature 0.7, and several point estimates contradict the claim. The benchmark's construct validity is a separate concern, but even taking the scenarios at face value, the improvement hypothesis is untested.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces UniToMBench, a benchmark that augments TOMBENCH with 1,025 hand-written multi-interaction and evolving-story scenarios, and evaluates five LLMs with and without the SimToM perspective-taking prompting framework. The authors report accuracy percentages on both TOMBENCH tasks and their custom tasks, and conclude that integrating multi-interaction tasks, evolving stories, and SimToM enhances LLM performance in detecting irony, sarcasm, and hidden emotions. The paper includes a public code repository and uses established TOMBENCH items for comparability.","tokens_in":12578,"tokens_out":5765,"duration_ms":62882,"significance":"If the benchmark and the SimToM integration were validated, UniToMBench could be a useful resource for studying Theory of Mind in LLMs: it combines a custom multi-turn dataset with an existing benchmark, evaluates multiple open and proprietary models, and releases code publicly. However, the central empirical claim is not supported by the reported numbers, and the dataset's validity as a Theory-of-Mind instrument is unestablished. As it stands, the paper does not provide reliable evidence that SimToM improves ToM performance or that the custom scenarios measure mental-state reasoning rather than literal comprehension.","major_comments":[{"comment":"The conclusion that SimToM 'enhances LLM performance' is contradicted by the paper's own results. GPT-4o's False Belief Task accuracy drops from 89.5 to 79.5, its Multi-Interaction accuracy drops from 71.2 to 63.0, GPT-3.5-Turbo's Evolving Stories accuracy drops from 69.5 to 55.8, and GPT-4o-Mini declines on seven of eight TOMBENCH tasks (e.g., False Belief Task 72.0 to 67.0, Strange Story Task 79.4 to 69.0, Hinting Test 74.8 to 68.9). With single runs at temperature 0.7 (Section 4.2), no confidence intervals, and no significance tests, the few improvements that do appear are within plausible sampling noise. The paper needs repeated runs, error bars, and paired statistical tests before any improvement claim can be evaluated.","section":"Section 5, Tables 1 and 6"},{"comment":"The construct validity of the 1,025 custom scenarios is not established. There is no human baseline, no inter-annotator agreement, and no psychometric validation. The 'cake' example is ambiguous: Tom prefers chocolate but thinks strawberry is easier, and the question 'Which cake would Tom choose?' does not specify whether he decides as a tie-breaker or for his own preference, so answer B is not uniquely determined by the story. The 'Evolving Story' example about Aaron's photography mostly tests literal comprehension of a sequence, not mental-state reasoning. Without evidence that the gold answers are unambiguous and that the tasks require ToM, the accuracy numbers in Table 6 cannot be interpreted as measures of Theory of Mind.","section":"Section 4.1, Dataset Composition"},{"comment":"The paper lists 'consistency across runs' as one of the three evaluation metrics, but no consistency data are reported anywhere in the paper. All models are queried once at temperature 0.7 (Section 4.2), which makes it impossible to assess reproducibility. Either the metric should be removed or the corresponding experiments and results should be added.","section":"Section 4.3, Evaluation Metrics"},{"comment":"The narrative overstates the findings. The Discussion says SimToM-enhanced models 'excel in maintaining consistency across multi-turn interactions,' but GPT-4o's Multi-Interaction accuracy falls from 71.2 to 63.0 and Llama 3 8B's falls from 74.4 to 68.5. The Conclusion claims improvements in detecting irony and sarcasm, yet the TOMBENCH Non-Literal Communication (Irony/Sarcasm) results are mixed or unreported in the main tables. The claims should be scaled back to match the data.","section":"Sections 6 and 7, Discussion and Conclusion"}],"minor_comments":[{"comment":"There are typos in section headings: 'Evaulating ToM Capabilities' and 'Methods for Enchancing ToM' should be 'Evaluating' and 'Enhancing'; 'SimTom' in the related-work paragraph should be 'SimToM'.","section":"Section 2 headings"},{"comment":"The caption says 'The averages of the performance scores from all tests are presented,' but the table shows individual task rows and no average row; please clarify what averages, if any, are reported.","section":"Table 1 caption"},{"comment":"The categories 'Intentions Explanations' appear twice in each table (e.g., 77%/71%/28% and 83.1%/70.8%/76.4% for GPT models), with no distinction between the two entries; this appears to be a labeling error and should be corrected.","section":"Appendix A, Tables 2-5"},{"comment":"The text says 'All models were queried at a temperature of 0.7 for consistency,' but a temperature above zero makes outputs stochastic; without seeds or multiple runs, temperature does not by itself ensure consistency.","section":"Section 4.2, Experimental Setup"},{"comment":"The 'Goal' paragraph is set as an unformatted heading; it should be formatted consistently with the rest of the section structure.","section":"Section 3, Goal"}],"recommendation":"reject","confidential_remarks":"The paper's central claim is contradicted by its own tables, and the dataset lacks basic validity evidence such as human baselines and item-level justification. The authors could potentially resubmit a very different paper with repeated runs, significance testing, and human validation, but the current manuscript does not meet the bar for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a benchmark paper whose central claim is undercut by its own results. The custom dataset is a real contribution, but the evidence that integrating SimToM 'enhances' ToM performance is mixed at best and mostly negative.\n\nWhat's actually new: 1,025 hand-written scenarios in two formats—multi-interaction dialogues and evolving stories—that address a genuine gap in static one-shot ToM benchmarks. That is a useful piece of work, and the authors correctly identify in Section 3 that SimToM can overgeneralize and hurt performance. The prose is clear, and the public code link helps.\n\nThe soft spots are serious. Section 5 reports single runs at temperature 0.7 with no error bars, repeated trials, or significance tests. The stress-test note is right: the direction of the effect is not consistently positive. On the custom set, SimToM essentially doesn't change GPT-4o's evolving-story score (89.9 vs 90.1) and lowers it for every other model except GPT-4o-Mini's multi-interaction score (55.8 to 65.8). On TOMBENCH, GPT-4o's false belief drops from 89.5 to 79.5, GPT-4o-Mini declines on seven of eight tasks, and GPT-3.5's unexpected outcome crashes from 34.7 to 10.0. The conclusion says integration enhances detection of irony, sarcasm, and hidden emotions, but the supporting tables tell a different story. The paper's own Section 3 lists reasons SimToM can hurt; the abstract and conclusion ignore those caveats.\n\nThe other soft spot is construct validity. There is no human baseline, no inter-annotator agreement, and the provided cake example is a preference judgment, not a clear belief or emotion inference. That makes it hard to know what the custom accuracy numbers measure. This is a secondary concern relative to the main claim, but it matters if the benchmark is to be trusted.\n\nWho this is for: people building or evaluating ToM benchmarks. The dataset could be a starting point, and the task formats are worth discussing. But as submitted, the improvement claim is not demonstrated. I'd send it to a serious referee, because the resource itself deserves scrutiny and the field needs more diverse ToM tasks, but I'd expect either major revision or rejection on the evidence.","headline":"A useful new ToM dataset, but the paper's central improvement claim is contradicted by its own tables.","tokens_in":13103,"tokens_out":2318,"would_cite":false,"duration_ms":26103,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"UniToMBench integrates perspective-taking to improve and assess Theory of Mind in LLMs.","keywords":["Theory of Mind","LLM evaluation","perspective-taking","SimToM","TOMBENCH","social cognition","multi-agent reasoning","benchmark design"],"falsifier":"Present the same 1,025 multiple-choice scenarios to human adults and to models after every mental-state verb has been replaced with a neutral verb, and after swapping which character knows which fact; if adult scores are at ceiling and model accuracy is unchanged by the swaps, the items measure literal comprehension, not Theory of Mind.","tokens_in":12081,"feed_emoji":"🧠","tokens_out":7579,"duration_ms":80058,"temperature":0.7,"pith_summary":"UniToMBench is a Theory of Mind (ToM) benchmark that combines perspective-taking prompts, multi-turn dialogues, and evolving story scenarios. The paper argues that this combination—SimToM's viewpoint-taking prompting, TOMBENCH-style task categories, and 1,025 hand-written multiple-choice items—improves language models' detection of irony, sarcasm, and hidden emotions, and gives a fuller read of their social reasoning than static single-turn tests. On the reported evaluations, GPT-4o and GPT-4o Mini stay above 80% on many emotion and belief tasks, while knowledge-based tasks vary widely, and the smaller open models usually lose accuracy when perspective-taking is added. The value, if the benchmark works as described, is a single testbed for both improving and measuring machine social cognition.","feed_headline":"UniToMBench boosts LLM scores on irony, sarcasm, and hidden emotions","feed_subtitle":"The unified benchmark pairs perspective-taking prompts with 1,025 evolving-story scenarios to test social cognition in language models.","key_machinery":"The load-bearing mechanism is SimToM (Simulation Theory of Mind) prompting: the model is first prompted to take a specific character's perspective, and only then answers the ToM question. UniToMBench embeds that mechanism in two custom scenario families—500 multi-interaction tasks that track mental states across conversational turns, and 525 evolving-story tasks that track motivation and relationships across time—while keeping TOMBENCH's eight task categories for comparability. The evaluation procedure measures task-completion accuracy, error attribution, and run-to-run consistency at temperature 0.7.","core_discovery":"The central claim is that adding the SimToM perspective-taking framework to a unified benchmark changes measured Theory-of-Mind performance in a task-dependent way, and that the change is informative. Specifically, the paper claims that multi-interaction tasks and evolving stories—supported by SimToM prompting—improve LLM performance in detecting irony, sarcasm, and hidden emotions, while static or knowledge-heavy tasks like faux-pas recognition and some belief-tracking tests remain weak points. The evidence is organized as accuracy over 1,025 custom scenarios plus TOMBENCH categories, with error attribution and consistency as secondary metrics. The authors present the benchmark as a tool to both stimulate and evaluate social cognition, not as a claim that models have genuine ToM.","pith_inferences":["Editorial inference: without a human baseline on the same 1,025 items, the 80%+ scores cannot be read as human-level ToM; collecting adult norms would calibrate the benchmark.","Editorial inference: the cake example is solvable from stated preference rather than belief tracking, so a rewording that requires false-belief reasoning would test whether these items measure ToM at all.","Editorial inference: swapping which character knows which fact in the custom scenarios would show whether SimToM-enhanced models track distinct mental states or exploit story templates."],"forward_implications":["SimToM should be applied selectively, not universally: it raised GPT-3.5 Turbo's Hinting Test score from 24.3% to 68.9%, but it lowered accuracy for Llama 3 8B and Gemma 2 27B on evolving-story and multi-interaction tasks.","Static ToM tests understate model ability: models that appear weak on one-shot false-belief questions can show better belief tracking when tested across evolving dialogues, and vice versa.","The 80%-plus emotion and belief scores of GPT-4o and GPT-4o Mini, paired with wide variation on knowledge-based tasks, identify where future training should concentrate.","Error patterns—missing implicit cues, conflating mental states across agents, and failing at abstract reasoning—give concrete targets for new benchmark tasks."],"supporting_citations":[{"why":"Supplies the SimToM perspective-taking prompting method that UniToMBench integrates.","marker":"Wilf et al., 2023"},{"why":"Supplies the TOMBENCH task battery and question set that UniToMBench extends and uses for comparison.","marker":"Chen et al., 2024"},{"why":"Documents the GPT-4o and GPT-4o Mini models whose results anchor the evaluation.","marker":"OpenAI et al., 2024"},{"why":"Introduces the Llama model family used for one of the evaluated models.","marker":"Touvron et al., 2023"},{"why":"Supplies the Chain-of-Thought prompting approach tested alongside SimToM in the experimental setup.","marker":"Wei et al., 2023"}],"fun_headline_variants":["Perspective-taking benchmark sharpens LLM Theory of Mind","UniToMBench: 1,025 scenarios test and boost LLM social cognition","New benchmark reveals GPT-4o's ToM peaks and knowledge gaps","Merging SimToM and TOMBENCH yields stronger ToM evaluation","LLMs get better at irony but still miss faux pas in new test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 1,025 hand-written scenarios are valid, unbiased measures of Theory of Mind, with unambiguous right answers and no way to solve them by story recall or surface wording alone.","fun_headline_variants_meta":{"raw":{"variants":["Perspective-taking benchmark sharpens LLM Theory of Mind","UniToMBench: 1,025 scenarios test and boost LLM social cognition","New benchmark reveals GPT-4o's ToM peaks and knowledge gaps","Merging SimToM and TOMBENCH yields stronger ToM evaluation","LLMs get better at irony but still miss faux pas in new test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000176,"raw_usage":{"total_tokens":1273,"prompt_tokens":911,"completion_tokens":362,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":262}},"tokens_in":527,"tokens_out":362,"duration_ms":4474,"temperature":1.0,"reasoning_tokens":262,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:47:10.771918+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Present the same 1,025 multiple-choice scenarios to human adults and to models after every mental-state verb has been replaced with a neutral verb, and after swapping which character knows which fact; if adult scores are at ceiling and model accuracy is unchanged by the swaps, the items measure literal comprehension, not Theory of Mind.","supporting_citations":[],"review_version":1}